Visual quotation for information provided in response to multimodal queries
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2024-05-08
- Publication Date
- 2026-08-04
AI Technical Summary
【0008】 本開示の様々な実施形態のこれらおよび他の特徴、態様、および利点は、以下の説明および添付の特許請求の範囲を参照すると、よりよく理解されよう。本明細書に組み込まれ本明細書の一部を構成する添付図面は、本開示の例示的な実施形態を示し、この説明と一緒に、関連する原理を説明するために役立つ。
Smart Images

Figure 0007900438000001 
Figure 0007900438000002 
Figure 0007900438000003
Abstract
Description
Technical Field
[0001] This disclosure generally relates to providing and presenting information for multimodal queries. More particularly, this disclosure relates to generating visual citations for information retrieved or derived in response to multimodal queries.
Background Art
[0002] In modern society, text-based search services are everywhere, but users often struggle to devise text-based queries in various situations. For example, users often find it difficult to describe objects that are unfamiliar to them. In another example, a user may not be able to correctly represent an intention (such as the subject of an intended query, etc.) in text. To facilitate a more efficient and accurate dialogue between the user and the search service, multimodal queries have been proposed. A multimodal query is a query devised using multiple types or formats of data (such as text content, audio data, video data, image data, etc.). For example, a user may provide a search service with a multimodal query that includes an image and an associated text prompt (such as an image of a bird and a text query "What kind of bird is this?"). The search service can use various multimodal query processing techniques to search for search results such as images and associated text content, and can be presented to the user so as to show some parts of the text content as being associated with a specific result image.
Summary of the Invention
Means for Solving the Problems
[0003] Aspects and advantages of embodiments of this disclosure will be described in part in the following description, or will be apparent from the description, or will be learned through the implementation of the embodiments.
[0004] One exemplary aspect of the present disclosure relates to a method implemented by a computer. The method implemented by a computer includes the step of searching for a result image based on the similarity between a query image and a result image by a computing system comprising one or more processor devices. The method implemented by a computer includes the step of obtaining a first text unit by the computing system, the first text unit comprising at least a portion of the text content of a source document containing the result image. The method implemented by a computer includes the step of determining a second text unit by the computing system in response to a prompt associated with the query image, the second text unit comprising one or more of (a) at least a portion of the first text unit, or (b) text derived from the first text unit. The method implemented by a computer includes the step of providing the second text unit and the result image by the computing system for display within an interface.
[0005] Another exemplary aspect of the present disclosure relates to a computing system. The computing system includes one or more processors. The computing system includes one or more non-temporary computer-readable media that store together a first set of instructions, which, when executed by one or more processors, cause the computing system to perform an operation. The operation includes retrieving a query image and associated prompts from a user computing device. The operation includes processing the query image with a machine learning-type embedding model to obtain a query image embedding. The operation includes searching for a result image based on the similarity between the query image embedding and the result image embedding. The operation includes identifying a source document for the result image, the source document including the result image and the text content associated with the result image. The operation includes determining a first text unit from the source document that includes at least a portion of the text content associated with the result image. The operation includes processing a first text unit and prompt with a machine learning language model to obtain a language output including a second text unit, the second text unit including one or more of (a) at least a portion of the first text unit, or (b) text derived from the first text unit. The operation also includes providing the second text unit and the resulting image for display within the interface of an application run by a user computing device.
[0006] Another exemplary aspect of the present disclosure relates to one or more non-temporary computer-readable media that collectively store a first set of instructions, which, when executed by one or more processors of a computing system, cause the computing system to perform an operation. The operation includes retrieving a plurality of result images based on the similarity between an intermediate representation of a query image and each of a plurality of intermediate representations associated with each of a plurality of result images. The operation includes identifying a plurality of source documents, each of which includes one of the plurality of result images and the text content associated with that result image. The operation includes determining a plurality of first text units for each of the plurality of result images, each first text unit including at least a portion of the text content associated with the result image from one or more source documents containing the result images. The operation includes processing a set of text inputs in a machine learning language model to obtain a language output including a second text unit, the set of text inputs including (a) two or more first text units associated with two or more of the plurality of result images, and (b) a prompt associated with a query image. The operation includes providing the user computing device with a second text unit and two or more resulting images for display within the user computing device interface.
[0007] Other aspects of this disclosure cover a variety of systems, apparatus, non-temporary computer-readable media, user interfaces, and electronic devices.
[0008] These and other features, aspects, and advantages of the various embodiments of this disclosure will be better understood by referring to the following description and the appended claims. The appended drawings incorporated herein and constituting part thereof illustrate exemplary embodiments of this disclosure and, together with this description, are useful for illustrating the relevant principles.
[0009] A detailed discussion of embodiments intended for those skilled in the art is provided herein, and this specification refers to the accompanying drawings. [Brief explanation of the drawing]
[0010] [Figure 1] This is a block diagram illustrating an exemplary visual search system in several implementations of the present disclosure. [Figure 2] This disclosure presents several implementations of data flow diagrams for providing information and accompanying visual citations in response to visual queries. [Figure 3] This is a flowchart illustrating exemplary methods for performing prompt response and corresponding visual quotation generation according to several implementations of the present disclosure. [Figure 4] This figure shows exemplary interfaces for a user computing device for displaying text content and corresponding interface elements, according to several implementations of the present disclosure. [Figure 5A] This figure shows an exemplary interface for a user computing device for displaying text content and corresponding interface elements, according to several other implementations of the present disclosure. [Figure 5B] This figure shows exemplary interfaces for a user computing device that appear following the interface in Figure 5A in response to the receipt of user input, according to several implementations of the present disclosure. [Figure 6A] This is a data flow diagram for the dynamic refinement of visual search information in response to user feedback in a first time period T1, based on several implementations of the present disclosure. [Figure 6B] This is a data flow diagram for dynamic refinement of visual search information in response to user feedback during a second time period T2, based on several implementations of the present disclosure. [Figure 7A]This figure shows exemplary interfaces for a user computing device for collecting user feedback on derived text content and corresponding result images, according to several implementations of the present disclosure. [Figure 7B] This figure shows exemplary interfaces for a user computing device for displaying visually refined search information based on user feedback, according to several implementations of the present disclosure. [Figure 8] This is a flowchart illustrating exemplary methods for providing visual search information derived from documents containing images retrieved based on visual similarity with query images, according to several implementations of the present disclosure. [Figure 9] This is a flowchart illustrating exemplary methods for refining visual search information based on user feedback, based on several implementations of the present disclosure. [Figure 10] This is a flowchart illustrating exemplary methods for collecting user feedback to refine visual search information, based on several implementations of the present disclosure. [Figure 11A] This is a block diagram of an exemplary computing system that performs a visual or multimodal search service in several implementations of the present disclosure. [Figure 11B] This is a block diagram of an exemplary computing system that performs visual search operations and / or refines visual search information, according to several implementations of the present disclosure. [Modes for carrying out the invention]
[0011] The repeated reference numbers across multiple diagrams identify the same features in various implementations.
[0012] In general, this disclosure concerns presenting information retrieved in response to multimodal queries to a user. More specifically, this disclosure concerns generating visual citations that visually identify the source of information provided and retrieved in response to queries such as visual queries or multimodal queries. A multimodal query is a query devised using various types of data (e.g., text content, audio data, video data, image data, etc.). In response to a multimodal query, a search system may retrieve and / or derive information using various multimodal query processing techniques.
[0013] As an example, suppose a user provides a multimodal query consisting of a query image of a bird and a corresponding prompt such as "What kind of bird is this?". The visual search system should search for result images that are visually similar to the query image. Based on the assumption that the source of the visually similar result image (e.g., a document containing images and text content) is likely to contain information related to the query image, the visual search system can extract information from the source of the result image. The visual search system can then derive text content from the extracted information based on the prompt (e.g., a summary of the information). For example, the visual search system may process the text content and prompt with a machine learning language model to generate a language output that includes the text content.
[0014] A visual search system can provide text content and visually similar images for display to the user on the interface of the user computing device. The interface may include attribute elements about the resulting images. These attribute elements may include a representation of the resulting image (e.g., a thumbnail) and information identifying the source of the image. As in the previous example, if one resulting image shows the same species of bird as indicated by the query image, and the source of the resulting image is a website, the attribute elements may include a thumbnail of the resulting image and information identifying the website (e.g., the website's title, URL, etc.). In this way, the user can quickly verify the accuracy of the text content based on the visual similarity between the resulting image and the query image, or by navigating to the source of the resulting image. For example, if the resulting image shows a bird that is clearly not the same species as the bird indicated by the query image, the user can quickly determine that the corresponding text content provided to the user is relatively likely to be inaccurate.
[0015] In some implementations, the user can select attribute elements to indicate that the corresponding result image is inaccurate, and therefore any information derived from the source of the result image is likely to be inaccurate. Based on the user's selection, the visual search system can derive text content. Following the previous example, the visual search system can search for four result images, each showing a bird. The visual search system can extract information from the sources of the four result images and, using a machine learning language model, process the extracted information according to prompts provided by the user to obtain a language output containing text content. The visual search system can provide the text content and the four attribute elements to the user's associated user computing device.
[0016] For example, assume that one of the four result images shows a bird that is clearly a different species from the birds shown in the query image and the other three result images. The user can select the attribute element that includes the result image (e.g., via a touch screen device), and the user computing device can indicate the selection of the result image to the visual search system. Previously, the visual search system may have generated text content provided to the user by processing a prompt and a corpus of information extracted from the sources of the four result images with a machine learning-based language model. Thus, in response to the selection of the result image, the visual search system may remove any information extracted from the source of the result image from the corpus of information and then process the remaining information with a machine learning-based language model to generate a second language output that includes different text content. This text content can be provided to the user computing device. In this way, the visual search system can repeatedly enhance the results based on user feedback regarding the visual citation.
[0017] Aspects of the present disclosure provide several technical effects and benefits. As one exemplary technical effect and benefit, a search service that can provide a direct response to a query is much more desirable to a user than a service that only provides a list of relevant documents, because the list of documents still requires the user to spend a significant amount of time and energy to conduct further investigations. However, most search services that are capable of providing a response to a user query do not provide the user with a function to verify the accuracy of the response. Without the function to verify the response, many users may reject using such search services.
[0018] However, the implementation forms of the present disclosure enable the provision of visual references for quickly and efficiently indicating the accuracy of responses to users. More specifically, by deriving a response to a query from information associated with a result image that is visually similar to the query image, the user can quickly determine the accuracy of the response based on what is shown in the result image. In this way, the implementation forms of the present disclosure can provide a response to the query and also enable the user to quickly verify the accuracy of the response.
[0019] It should be noted that, as long as described in this specification, "text unit", "text content", and "text" can be used interchangeably. Generally, each of the above terms can refer to one or more alphanumeric units. For example, text content, text unit, and text can refer to individual paragraphs, words, single numbers, alphanumeric sequences, lines of program code or instructions, machine language, machine-readable code, etc.
[0020] Furthermore, it should be noted that any text, text content, and / or text unit referred to in this specification can be derived from audio data, image data, audiovisual data, etc. For example, the "document" further defined in this specification may be a news article scanned and saved as an image. Text can be extracted from such an image using conventional optical character recognition techniques. Therefore, an image showing text can be called text even if intermediate processing techniques are used to extract the text from the image. This is also applicable to audio and audiovisual media such as recordings of conversations, videos, podcasts, music, video games with dialogues, etc. More generally, it will be understood by those skilled in the art that any uttered voice, text narration, or other medium from which text can be derived can generally be referred to as "text" throughout this specification.
[0021] Here, referring to the drawings, the exemplary embodiments of the present disclosure will be discussed in more detail.
[0022] Figure 1 shows a block diagram of an exemplary visual search system 100 in several implementations of the present disclosure. More specifically, the user computing device 102 may include an input device 104 and a communication module 106. The input device 104 is or may otherwise include a device that can directly or indirectly receive input from a user (e.g., a microphone, camera, touchscreen, physical button, infrared camera, mouse, keyboard, etc.). The communication module 106 is or may otherwise include hardware and / or software collectively configured to communicate with the visual search computing system 108 via a network 110. For example, the communication module 106 may include a device that facilitates wireless connectivity to the network.
[0023] The user computing device 102 can acquire a query image 112 and an associated prompt 114. The query image 112 may be a selected image that serves as a query to the visual search computing system 108. For example, a user of the user computing device 102 can acquire the query image 112 using the input device 104 of the user computing device 102. Alternatively, the user may acquire the query image 112 in some other way (for example, by taking a screen capture, downloading an image, or creating an image using an image creation tool).
[0024] The prompt 114 associated with the query image 112 may include text content provided by the user. For example, the user may directly prompt for the text content of prompt 114 using a keyboard or some other input method included in the input device 104. Alternatively, the user may provide prompt 114 indirectly. For example, the user may produce a spoken utterance, which the user computing device 102 may capture in the input device 104. The user computing device 102 may then process the spoken utterance using speech recognition techniques (e.g., a machine learning text-to-speech model) to generate prompt 114.
[0025] In some implementations, the query image 112 is provided without a corresponding prompt 114, and the user computing device 102 may determine an appropriate prompt associated with the query image 112. For example, if the query image 112 shows a bird as the subject in the image, the user computing device 102 may select an appropriate prompt 114 to provide with the query image 112 (e.g., "Identify this object," "Describe this," "Tell me more," etc.). Alternatively, in some implementations, the user computing device 102 may modify the prompt 114 provided by the user of the user computing device 102. For example, the user computing device 102 may modify the prompt 114 to add contextual information (e.g., time, geolocation, user information, information describing the preceding query image and / or the prompt provided by the user).
[0026] The user computing device 102 can provide a visual search request 116 to the visual search computing system 108 via the network 110. The visual search request 116 may include a query image 112 and a prompt 114.
[0027] The visual search computing system 108 may include a visual search module 118. The visual search module 118 can process a visual search request 116 to obtain text content 120 and a result image 122. The text content 120 may respond to a prompt 114 and a query image 112. For example, if the query image 112 shows an animal and the prompt 114 is "What is this animal?", the text content 120 may provide a response to the prompt (e.g., the species of the animal, or, if it is a known animal, its name). The result image 122 may be an image visually similar to the query image 112. These result images 122 are included in the document from which the text content 120 was derived. The visual search module 118 can extract at least a portion of the text contained in one or more documents. In some implementations, the visual search module 118 may employ various processing techniques to identify portions of text within a document that are likely to be related to a result image contained within the document.
[0028] More specifically, the visual search module 118 can retrieve a result image 122 based on the similarity between the query image 112 and the result image 122. For example, the visual search module 118 may include a machine learning model that can be used to identify images that are visually similar to the query image 112. The visual search module 118 can retrieve text from a document containing the result image. The document may be any type or format of source material containing the result image, such as a website, academic journal, book, newspaper, article, social media post, transcript, or blog, as described herein. In some implementations, the visual search module 118 may select text content 120 from the text extracted from the document. As an addition or alternative, in some implementations, the visual search module 118 may derive text content 120 from the text extracted from the document. For example, the visual search module 118 may include, or otherwise have access to, a machine learning model, such as a large language model, which can process the text and prompts extracted from the document to obtain text content 120.
[0029] The visual search module 118 may provide interface data 124 to the user computing device 102 via the network 110. The interface data 124 may include text content 120 and a result image 122. For example, the interface data 124 may include instructions for highlighting the text content 120 and for including a thumbnail representation of the result image 122, thereby allowing the user of the user computing device 102 to easily verify the accuracy of the result image 122 and, accordingly, the accuracy of the text content 120. In this way, the visual search computing system 108 can provide a response to a query while facilitating the user's rapid and accurate verification of the response.
[0030] Figure 2 shows a data flow diagram 200 for providing information and accompanying visual citations in response to a visual query, according to several implementations of the present disclosure. More specifically, a visual search computing system 202 (e.g., a physical server computing system, a cloud computing system, a virtualized and / or physical computing node in a network (e.g., an edge computing node)) may include a visual search module that can retrieve a query image 206 and a prompt 208 from a user computing device. For example, a user computing device 203 can provide a visual search request to the visual search computing system 202 over the network.
[0031] In some implementations, the query image 206 may be received from the visual search computing system 202 without an associated prompt. In such situations, the visual search module 204 may determine that it is likely to generate a prompt associated with the query image 206. For example, the visual search module 204 may include a machine learning semantic image model trained to generate a semantic description of the query image 206. In some implementations, the visual search module 204 can use the semantic description of the query image 206 as a prompt 208. Alternatively, in some implementations, the visual search module 204 may process the semantic description of the query image 206 with another machine learning model, such as a large-scale language model, to generate the prompt 208.
[0032] In some implementations, the visual search module 204 may modify the prompt 208 received from the user computing device 203. For example, the visual search module 204 may modify the prompt 208 to add contextual information (e.g., the time, the geolocation of the user computing device 203, stored user information associated with the user of the user computing device 203, information describing a previous query image and / or a prompt provided by the user computing device 203).
[0033] The visual search module 204 may include an image evaluation module 210. The image evaluation module 210 can perform various processing techniques to identify a result image 212 that is visually similar to the query image 206. For example, the image evaluation module 210 may include a machine learning visual search model 214 that is trained to identify images that are visually similar to the query image 206 from a corpus of stored image data. For example, in some implementations, the visual search model 214 may be a machine learning coding model, such as an embedding model, which can be used to generate an intermediate representation (e.g., an embedding) of the query image 206. The image evaluation module 210 may include or be able to access an image search space 215. The image search space 215 may include intermediate representations for multiple stored images. For example, the image search space 215 may be an embedding space containing embeddings generated for images stored in a data store (e.g., a database) that stores a large number of images and indexes them to facilitate the visual search service. The image evaluation module 210 should select the result image 212 that has the closest embedding to the embedding generated for the query image 206 in the embedding space.
[0034] Note that this example shows a single result image 212 solely to more clearly illustrate an exemplary implementation of the disclosure. However, such an implementation is not limited to obtaining a single result image 212. Rather, the result image 212 may be any number of result images obtained based on the similarity between the result image and the query image 206.
[0035] As previously mentioned, the visual search module 204 can index a large number of images to facilitate the visual search service. The visual search module 204 can also index information indicating the source documents that contain the result images in the document indexing information 216, or otherwise associate them with them. Documents may be any type or format of source material, including the result images, such as websites, journals, books, newspapers, articles, social media posts, transcripts, blogs, etc., as described herein. A result image may be “associated” with a document if the result image is generated, created, or hosted by the same entity as the document, for example, the document itself. A result image may be associated with a document if, for example, the result image is used as the cover image of the document, is derived from the document (e.g., the output of a generative model), and the document is a frame of a video from which the document is transcribed. A result image may be “contained” with a document if the result image is currently in the document, or was in the document when the result image and / or the document were indexed.
[0036] The resulting image 212 may be included in document 220. Document 220 may include the resulting image 212 and text content 222. In some implementations, document indexing information 216 may include the resulting image 212, or otherwise include the associated document 220, or include text content extracted from document 220. As an addition or alternative, in some implementations, document indexing information 216 may describe the source location of document 220 (e.g., a file location on a network, a website URL, an FTP address, etc.). As an addition or alternative, in some implementations, document indexing information 216 may include a compressed version of document 220.
[0037] As described with respect to Result Image 212, this example shows a single document 220 solely to more clearly illustrate an exemplary implementation of the disclosure. However, such an implementation is not limited to obtaining a single document 220. Rather, in some implementations, multiple documents 220 may each contain multiple Result Images 212 (for example, five documents for five Result Images). Additional or alternative, in some implementations, a single document 220 may contain multiple Result Images 212. Additional or alternative, in some implementations, multiple documents 220 may each contain an instance of a single Result Image 212.
[0038] As a concrete example, suppose the query image 206 shows a high-speed vessel, and the image evaluation module 210 selects a result image 212 showing the same high-speed vessel from a different angle. If the selected result image 212 is hosted on or originally hosted on a website (i.e., a document) for high-speed vessel enthusiasts, the visual search module 204 may store a link to the website, an archived version of the website, or text content extracted from the website in the document indexing information 216. More generally, the visual search module 204 may store information in the document indexing information 216 indicating the association between a document and its corresponding result image.
[0039] The visual search module 204 may include a document content selection module 218. The document content selection module 218 can search for a document 220 containing a result image 212. Once a document 220 is found, the document content selection module 218 can extract a first text unit 224 from the text content 222 of the document 220. The first text unit 224 may contain part or all of the text content 222 of the document 220. In some implementations, the document content selection module 218 may employ various processing techniques to identify portions of text within the text 224 that are likely to be related to the result image 212 contained in the document 220. For example, if the document 220 is an online article and the result image 212 is in the middle of the online article, the document content selection module 218 may heuristically select text before and after the result image (e.g., paragraphs, sentences or word count, column, etc.) to include in the first text unit 224. Alternatively, in some implementations, the document content selection module 218 can extract all the text contained within the document for inclusion in the first text unit 224.
[0040] The visual search module 204 may include a text determination module 226. The text determination module 226 can determine a second text unit 228 based on a first text unit 224 and a prompt 208. In some implementations, the text determination module 226 can use a machine learning language model 230 to determine the second text unit 228. For example, the text determination module 226 can process the first text unit 224 and the prompt 208 to obtain the second text unit 228. In some implementations, the machine learning language model 230 may be a large-scale language model trained on a large corpus of training data to perform multiple generative tasks. Furthermore, in some implementations, the machine learning language model 230 may have undergone additional training iterations to harmonize or optimize the model for a specific implementation of the language task relating to the generation of the second text unit 228.
[0041] The visual search module 204 may include an interface data generation module 232. The interface data generation module 232 can generate interface data 234 and send the interface data 234 to the user computing device 203. The interface data 234 may include a second text unit 228 and a result image 212. The interface data 234 can indicate how the second text unit 228 and the result image 212 will be displayed within the interface of the user computing device 203. For example, the interface data 234 may indicate how to display the second text unit 228 and the result image 212 within the interface of an application run by the user computing device 203 (e.g., a visual search application). The display of the second text unit 228 and the result image 212 within the interface of an application run by the user computing device 203 will be discussed in more detail with respect to Figures 4, 5A, and 5B.
[0042] In some implementations, the interface data generation module 232 can generate attribute information 236. Attribute information 236 is information that identifies the document 220, or otherwise may include such information. For example, if the document 220 is a news article, the attribute information 236 may be the title of the news article and the name of the publishing news organization. In another example, if the document 220 is a scholarly paper, the attribute information 236 may be the title of the scholarly paper, the first author, a list of authors, and reference citations. In yet another example, if the document 220 is a website or some other form of document accessible to the user computing device 203, the attribute information 236 may include a link (e.g., a URL) that facilitates access to the document 220.
[0043] Figure 3 is a flowchart of an exemplary method 300 for carrying out a response to a prompt and the generation of a corresponding visual quotation, according to an exemplary embodiment of the present disclosure. Figure 3 shows the actions performed in a specific order for illustrative and discussion purposes, but the methods of the present disclosure are not limited to the order or arrangement shown. Various actions of method 300 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.
[0044] In 302, the computing system can retrieve a result image based on the similarity between the query image and the result image. In some implementations, to retrieve a result image, the computing system can process the query image with a machine learning-based visual search model to obtain an intermediate representation of the query image, and then select a result image based on the similarity between the intermediate representation of the query image and the intermediate representation of the result image. For example, the machine learning-based visual search model may be a machine learning-based embedding model that generates an embedding of the query image for an image embedding space. The computing system can retrieve a result image by evaluating the embedding space, which includes the embedding of the result image and several other image embeddings. The embedding of the result image may be an embedding of a query image embedding within the embedding space.
[0045] In some implementations, the computing system can retrieve query images from the user computing device prior to searching for result images. In some implementations, the interface includes a user interface for an application run by the user computing device. For example, the user computing device can run a visual search application associated with a visual search service provided by the computing system. The visual search application can facilitate the capture of query images and prompts for transmission to the computing system.
[0046] In some implementations, obtaining a query image may involve obtaining the query image and the associated prompt from the user computing device. Furthermore, in some implementations, the prompt can be modified by the computing system. For example, the prompt may be modified to include instructions that facilitate processing of the prompt by a large-scale language model.
[0047] In some implementations, retrieving a result image further includes providing the result image to the user computing device for display within an interface, and receiving a prompt associated with the query image from the user computing device in response to the provision of the result image. For example, a computing system can receive a query image, perform a visual search, and retrieve the result image. The computing system can then provide the result image for display within the interface of an application run by the user computing device. In response, the user of the user computing device can enter a query into the user computing device, which can be provided to the computing system.
[0048] In 304, the computing system may obtain a first text unit, which includes at least a portion of the text content of the source document, including the resulting image. In some implementations, the document includes one or more web pages of a website, an article, a newspaper, a book, or a transcript. For example, if the source document is a website article, the computing system may obtain a first text unit that includes the text content of the article, the title of the article, other text content related to the website hosting the article, and so on.
[0049] In 306, the computing system can determine a second text unit in response to a prompt associated with a query image, the second text unit comprising one or more of (a) at least a portion of the first text unit, and (b) text derived from the first text unit. In some implementations, determining a second text unit in response to a prompt associated with a query image comprises processing the second text unit and the prompt associated with the query image with a machine learning language model to obtain a language output containing the second text unit. In some implementations, the second text unit comprises a subset of the first text unit. In some implementations, the second text unit comprises text derived from the first text unit, the text derived from the first text unit describing a summary of the first text unit.
[0050] In some implementations, prior to determining a second text unit, the computing system can generate a prompt associated with the query image, based at least partially on the query image. For example, the computing system can process the query image with a machine learning model, such as a semantic image analysis model, to generate a semantic output that describes the query image. The computing system can then use this semantic output as a prompt.
[0051] In 308, the computing system may provide a result image and a second text unit for display within the interface. For example, the computing system may send the result image and the second text unit to the user computing device that provided the query image and prompt. In some implementations, providing a result image and a second text unit involves providing data that describes an interface element containing the second text unit and an attribute element containing attribute information identifying the result image and the document containing the result image, for display within the interface. In some implementations, the document may be a web page, and the attribute information may include the address of the web page. Alternatively, in some implementations, the document may be a journal, and the attribute information may include a citation indicating the location of the result image within the journal.
[0052] Figure 4 shows exemplary interfaces 400 for a user computing device for displaying text content and corresponding interface elements in several implementations of the present disclosure. Figure 4 will be discussed in conjunction with Figure 2. More specifically, a visual search request 402 may include the query image 206 and prompt 208 in Figure 2. The visual search request 402 may be provided to a visual search computing system 202 as described with respect to Figure 2. The visual search computing system 202 may process the visual search request 402 to obtain interface data 234. The interface data 234 may include a second text unit 228 and a result image 212.
[0053] Specifically, interface data 234 may indicate how the resulting image 212 and the second text unit 228 are displayed within the interface 400 of the user computing device 203. For example, the user computing device 203 may be running a visual search application, or it may already be running an application integrated into the operating system of the user computing device 203. The application can display interface 400 on the display device of the user computing device 203.
[0054] As shown in the illustrated example, the query image 206 may represent a specific breed of dog, such as a beagle. The prompt 208 may be a question such as, "What breed of dog is this?". The visual search computing system 202 can process the visual search request 402 to obtain a second text unit 228 and a result image 212, which may be included in the interface data 234. As illustrated, the second text unit 228 may contain a response to the query imposed by the prompt 208, such as, "Answer: Beagle". Similarly, the result image 212 can be retrieved by visual similarity between the query image 206 and the result image 212.
[0055] In some implementations, interface data 234 may describe how the second text unit 228 and the resulting image 212 are presented. For example, interface data 234 may indicate that the second text unit 228 should be presented within the main interface element 404 in such a way that it is emphasized. Interface data 234 may further indicate that the main interface element 404 should contain an attribute element 405.
[0056] Attribute element 405 may include the resulting image 212 and corresponding attribute information 236. Attribute information 236 may identify the document containing the resulting image 212 (for example, document 220 in Figure 2) from which the second text unit 228 was extracted or derived. If the document is a website or otherwise accessible by the user computing device 203, attribute information 236 may also provide a link to access the document. In this way, attribute element 405 can function as a “visual reference” to allow the user of the user computing device 203 to easily verify the accuracy of the second text unit 228.
[0057] In some implementations, the interface data may include several other result images 212 in addition to the result image 212 included in the document from which the second text unit 228 is derived. The interface data 234 may indicate instructions for displaying the other result images 212 within result image elements 406A, 406B, 406C, and 406D (generally, result image element 406). In some implementations, the result images 212 included within result image element 406 may be result images that are less similar to the query image 206 than the result images included within the main interface element 404. For example, the visual search computing system 202 may select five result images 212 to include in the interface data 234. The result image 212 most similar to the query image 206 (for example, the image in the embedding space with the closest embedding to the embedding of the query image 206) may be indicated for inclusion in the main interface element 404. Result image element 406 may include four other result images 212.
[0058] Similar to the main interface element 404, the result image element 406 may include attribute information 236 that identifies the document containing the result image 212 of the result image element 406. As shown in the illustrated example, each result image element 406 may include a link to the website document containing the result image 212 contained within each result image element 406.
[0059] Figure 5A shows an exemplary interface 500A of a user computing device for displaying text content and corresponding interface elements, according to some other implementations of the present disclosure. Figure 5 will be discussed in conjunction with Figures 2 and 4. A visual search request 502 may be provided to the visual search computing system 202, as described with respect to Figure 2. The visual search request may include a query image 206 and a prompt 208. The visual search computing system 202 may process the visual search request 502 to obtain interface data 234. The interface data 234 may include a second text unit 228 and a result image 212.
[0060] As shown in the illustrated example, the query image 206 may represent a specific type of passenger jet. The prompt 208 may or may not function as a query; it may be a statement such as "Is it a good plane?". The visual search computing system 202 can process the visual search request 502 to obtain a second text unit 228, a result image 212, and attribute information 236, which may be included in interface data 234. As illustrated, the second text unit 228 may contain a response to the query imposed by the prompt 208. Similarly, the result image 212 can be retrieved by visual similarity between the query image 206 and the result image 212.
[0061] Interface 500A is similar to interface 400 in Figure 4, except that interface 500A may display interface elements in a different format than the format in which interface elements are displayed in interface 400 in Figure 4. For example, in Figure 4, the main interface element 404 includes text content from the second text unit 228 and attribute element 405 in a format that provides a clear response to a query imposed by the user. However, unlike the main interface element 404, the main interface element 504 in Figure 5A includes a first portion 228A of text content that includes an excerpt from a first document that provides more contextual information about the query imposed by the user in prompt 208. Furthermore, the main interface element 504 includes a highlighting element 506 that highlights or emphasizes information that is expected to function as a response to the query imposed by prompt 208. The first document can be identified by attribute element 505 in the same way as described with respect to attribute element 405 in Figure 4.
[0062] Specifically, when processing prompt 208, the visual search computing system 202 can determine whether to generate interface data 234 for interface elements that contain a direct response to the query, such as interface element 404 in Figure 4, or interface elements that contain contextual information that can assist the user, such as interface element 504 in Figure 5A. This determination may be based on the degree of certainty associated with the information retrieved in response to prompt 208, the semantic understanding of prompt 208, and so on. In the illustrated example, since "Is it a good airplane?" is a relatively subjective question, the visual search computing system 202 may determine, based on its semantic understanding of prompt 208, to generate the illustrated interface data 234. The determination of the type, style, and format of interface elements to be included in the interface data 234 will be discussed in more detail with respect to Figures 6A to 7B.
[0063] In some implementations, interface data 234 may contain information for display within multiple interface elements. In other words, interface data 234 may contain or be used to generate multiple interface elements containing different text content. As shown in the illustrated example, interface data 234 may contain information for inclusion within a primary interface element 504 and a second interface element 508. The second text unit 228 contained within interface data 234 may contain first text content from a first document and second text content from a second document. The first text content may be provided for inclusion in the primary interface element 504, and the second text content may be provided for inclusion in the second interface element 508.
[0064] The visual search computing system 202 can determine whether to generate interface data 234 containing information to be included in multiple interface elements. Similar to the determination of the format for the main interface element 504, this determination can be made based on the semantic understanding of the prompt 208, the quantity, quality, and / or semantic understanding of the text retrieved in response to the prompt 208, etc. Furthermore, in some implementations, the visual search computing system 202 may determine the order in which interface elements 504 and 508 will be presented to the user.
[0065] Assume that the user of user computing device 203 did not find the information presented in the main interface element 504 sufficient. The user may provide input 510 to user computing device 203 instructing it to display a second interface element 508. As shown in the illustrated example, the user may provide a "swipe" touch input 510 to move the second interface element 508 from a position where most of the element is hidden to a position where the entire element is visible.
[0066] For example, Figure 5B shows an exemplary interface 500B of a user computing device that appears following interface 500A in Figure 5A in response to the receipt of user input, according to some other implementations of the present disclosure. Moving to Figure 5B, interface 500B appears in response to the receipt of user input 510. As shown, in interface 500B, the primary interface element 504 is shifted to a fully concealed position, while the second interface element 508 is shifted to a fully visible position. Like the primary interface element 504, the second interface element may include a second emphasis element 512 that highlights, emphasizes, or otherwise indicates a portion of the second text content 228B that is expected to be particularly relevant to prompt 208.
[0067] In some implementations, the interface 500B of the user computing device 203 may include an information request element 514 that the user can select to indicate a request for additional information. For example, suppose the visual search computing system 202 determines that there is a relatively good chance that the information contained in the second text unit 228 is sufficient for the prompt 208. Rather than continuing to search for information to include in a third, fourth, or fifth interface element, the visual search computing system 202 may decide to include only the first text content 228A and the second text content 228B in the interface data 234 in order to reduce the consumption of computing resources (e.g., computation cycles, memory usage, power, storage, bandwidth, network resources, etc.), reduce latency, and increase efficiency.
[0068] However, if the user determines that the information contained in the primary interface element 504 and the second interface element 508 is insufficient, the user may select an information request element 514. If the user selects the information request element 514, the user computing device 203 may send a request to the visual search computing system 202. In response, the visual search computing system 202 may generate additional interface data to be included in the third interface element (or more). In this way, the visual search computing system 202 can facilitate iterative retrieval of information in response to prompt 208 while eliminating unnecessary use of computing resources.
[0069] Figure 6A is a data flow diagram for the dynamic refinement of visual search information in response to user feedback in a first time period T1, according to several implementations of the present disclosure. In particular, the visual search computing system 602 (e.g., the visual search computing system 202 in Figure 2) may include a visual search module 604 (e.g., the visual search module 204 in Figure 2). In the first time period T1, the visual search computing system 602 can acquire a query image 606 and a prompt 608, which can then be processed by the visual search module 604.
[0070] More specifically, in a first time T1, the visual search module 604 can process the query image 606 with an image evaluation module 610, such as the image evaluation module 210 in Figure 2, as described in the previous drawings, to obtain a result image 612. The result image 612 may include a first result image 612A, a second result image 612B, and a third result image 612C. The visual search module 604 can obtain a text unit 614 from the document associated with the result image 612. Specifically, the visual search module 604 can use the document content selection module 616 to obtain information from the document containing the result image 612 based on document indexing information 618. The document indexing information 618 may store information indicating the documents that the result image was in when it was indexed by the visual search computing system.
[0071] Assuming that the resulting image 612A was included in two separate documents 618A and 618B when indexed by the visual search computing system 602, as shown in the illustrated example, the document indexing information 618 may store the text content contained in documents 618A and 618B at the time of indexing, or it may store information indicating the location from which documents 618A and 618B can be accessed (e.g., URLs, download links, file locations, etc.). For example, if documents 618A and 618B are published academic journal articles, the visual search computing system 602 may store the text content directly contained in the documents, since the text content is relatively unlikely to change over time. Conversely, if documents 618A and 618B are both website pages, the visual search computing system 602 may store the URLs from which documents 618A and 618B can be accessed, since the information contained in website pages is relatively likely to be updated or repeated over time.
[0072] Continuing from the previous example, text unit 614 may include a first text unit 614A, a second text unit 614B, and a third text unit 614C. Each text unit may contain text content contained within the document containing the resulting image 612. For example, since the resulting image 612A is contained within documents 618A and 618B, text unit 614A corresponding to the resulting image 612A may contain text content from both documents 618A and 618B. Text unit 614B corresponding to the resulting image 612B may contain text content from document 618C containing the resulting image 612B. Text unit 614C may contain text content from document 618D containing the resulting image 612C.
[0073] The visual search computing system 602 can obtain a derived text unit 622 by processing text unit 614 and prompt 608 in text decision module 620, as described with respect to text decision module 226 in Figure 2. More specifically, the text decision module can obtain a derived text unit 622 by processing a set of text inputs including (a) text unit 614 and (b) prompt 608. For example, text decision module 620 may include a large-scale language model 621. The large-scale language model 621 may be a model trained on a large and diverse corpus of data for performing various types of language tasks. The large-scale language model 621 can process a set of text inputs to generate a derived text unit 622. The set of text inputs may include text unit 614 and prompt 608.
[0074] In some implementations, the derived text unit 622 may be the language output from a machine learning language model included in the text decision module 620. Therefore, the derived text unit 622 may be a generative language output that was generated from text unit 614 but contains some text content that is not included in it. Alternatively, the derived text unit 622 may be the language output that contains some (or all) of the text content of text unit 614.
[0075] The visual search computing system 602 may provide the result image 612 and derived text unit 622 to the user computing device 624 for display within the interface of the user computing device 624. Furthermore, in some implementations, the visual search computing system 602 may provide attribute information 626 to the user computing device 624 along with the result image 612 and derived text unit 622. For example, the attribute information 626 may include information stored in document indexing information 618 that identifies and / or provides access to documents 618A-618D.
[0076] In some implementations, in response to receiving the result image 612, the derived text unit 622, and attribute information 626, the user computing device 624 may provide the visual search computing system 602 with result image selection information 628. The result image selection information 628 may be information generated in response to user input collected by the user computing device 624, which selects one of the result images 612 within an interface to indicate that the selected result image 612 is inaccurate.
[0077] For example, moving to Figure 7A, Figure 7A shows an exemplary interface 700A for a user computing device for collecting user feedback on derived text content and corresponding result images, according to several implementations of the present disclosure. Figure 7A is discussed in relation to Figure 6A. In particular, assume that the query image 606 is an image of a passenger jet and the prompt 608 is the query "What is the maximum distance?". In response, the visual search computing system 602 can generate attribute information 626, derived text units 622, and result images 612, which can be provided for display on the interface 700A of the user computing device 624.
[0078] The user computing device 624 can display this information on interface 700A. Interface 700A may include an interface element 702 containing a derived text unit 622. As shown in the illustrated example, the derived text unit 622 may contain information regarding the maximum distance of a passenger jet aircraft, as shown in the query image 606 retrieved in response to prompt 608. Here, the derived text unit 622 is information related to the maximum distance of a passenger jet aircraft, summarized from multiple source documents.
[0079] Furthermore, interface 700A may include selectable attribute elements 704A, 704B, and 704C (generally, selectable attribute element 704). Selectable attribute element 704 is an interface element that includes a result image and attribute information that identifies the document containing the result image. In particular, the document identified by selectable attribute element 704 is the document from which the derived text unit 622 was derived. Based on the assumption that the text content of a document is closely related to the images contained within the document, the user can quickly and efficiently evaluate the relevance of the document used to derive the derived text unit 622 by viewing the result image contained in the attribute element associated with the document. To indicate that the result image (and therefore the document containing the result image) is irrelevant, the user can select the selectable attribute element 704 that contains the result image.
[0080] For example, a selectable attribute element 704A includes a result image 612A and attribute information 626 indicating the identity of the document containing the result image 612A (e.g., document 618A). Since the result image 612A included in the selectable attribute element 704A is a strict visual match with the query image 606, the user is unlikely to select the selectable attribute element 704A. However, the result image 612B included in the selectable attribute element 704B is clearly not visually similar to the query image 606, because the query image 606 shows a passenger jet, while the result image 612B shows a fighter jet. This visual mismatch may lead the user to provide input 706 that selects the attribute element 704B.
[0081] As shown in Figure 7A, the text content contained in document 618C (i.e., the “source” of result image 612B), which contains result image 612B, relates to fighter jets, not passenger jets, and is therefore irrelevant to prompt 608 and query image 606. Since the derived text unit 622 is generated based at least partially on the text content from document 618C, it is relatively likely that the derived text unit 622 is at least partially inaccurate. This is demonstrated by the summarized information contained in interface element 702, which includes information related to fighter jets (e.g., “The F-37 has a VTOL configuration and can take off relatively easily from aircraft carriers,” “US allies have purchased more than 200 F-37 aircraft,” etc.).
[0082] By providing an input 706 for selecting attribute element 704B, the user can indicate to the user computing device 624, and therefore to the visual search computing system 602, that the document containing the result image 612B is not related to the prompt 608 and therefore should not be used to generate the derived text unit 622. In response to receiving input 706, the user computing device 624 can generate result image selection information 628 and provide it to the visual search computing system 602.
[0083] Moving on to Figure 6B, which is a data flow diagram for the dynamic refinement of visual search information in response to user feedback in a second time period T2, according to several implementations of the present disclosure. Specifically, the visual search computing system 602 may receive result image selection information 628. The result image selection information 628 may indicate that the result image 612B is not visually similar to the query image 606. In response, at time T2, the visual search computing system 602 may generate a second derived text unit 630, which is generated based on each of the previous text units 614, excluding the text unit 614B extracted from the document 618B that contained the result image 612B.
[0084] The visual search module 604 may receive result image selection information 628 indicating the selection of attribute element 704B containing the result image 612B. In response, the visual search module 604 may identify each text unit previously used to generate the derived text unit 622. The visual search module 604 may then delete any text units obtained from document 618C (for example, the document that contained the result image 612B) that served as the source document for the result image 612B.
[0085] For example, to generate the derived text unit 622, the visual search module may have processed a first set of text inputs that included text units 614A, 614B, 614C, and prompt 608. In response to the result image selection information, the visual search module 604 may determine a second set of text inputs that includes each of the text units from the first set of text inputs, except for text unit 614B associated with the result image 612B contained within the selectable attribute element 704B indicated by the result image selection information 628. Here, the second set of text inputs may include text units 614A, 614C, and prompt 608.
[0086] If so, the visual search module 604 can process a second set of text inputs with the large language model 621 to generate a second derived text unit 630. Since the second derived text unit 630 is not based on information indicated as inaccurate by the user, it can be assumed that the second derived text unit 630 contains more accurate information than the information contained in the derived text unit 622. In this way, the visual search computing system 602 can dynamically and iteratively refine the visual search information (e.g., derived text units, attribute information, result images, etc.) in response to user feedback. The second text unit 630 may be provided to the user computing device 624 for display within the interface of the user computing device 624.
[0087] In some implementations, the visual search module 604 can generate a second attribute information 632. The second attribute information 632 may include all the information contained in the attribute information 626, except for the attribute information related to document 618C. Alternatively, in some implementations, the second attribute information 632 may include instructions to prevent the display of information related to document 618C on the interface of the user computing device 624. Additionally or alternatively, in some implementations, the visual search module 604 may retransmit a result image 612 other than the result image contained in the selectable attribute element 704B.
[0088] For example, moving to Figure 7B, which shows an exemplary interface 700B for a user computing device for displaying refined visual search information based on user feedback, according to several implementations of the present disclosure. In particular, upon receiving result image selection information 628 at time T2, the visual search computing system 602 can generate second attribute information 632 and a second derived text unit 630 and provide them to the user computing device 624 for display on interface 700B.
[0089] Interface 700B may include an interface element 708 containing a second derived text unit 630. As shown in the figure, the second derived text unit 630 is not generated based on information contained in document 618C, so the second derived text unit 630 does not contain errors associated with the contents of document 618C (for example, information about fighter aircraft). Furthermore, in response to input 706 selecting the selectable attribute element 704B, the selectable attribute element 704B is removed from interface 700B. Alternatively, an additional selectable attribute element may be displayed in place of attribute element 704B. In this way, the user computing device 624 can communicate with the visual search computing system 602 to refine the visual search information based on user feedback.
[0090] Figure 8 shows a flowchart of an exemplary method 800 for providing visual search information derived from a document containing images retrieved based on visual similarity to a query image, according to some exemplary embodiments of the present disclosure. Figure 8 shows steps performed in a particular order for illustrative and explanatory purposes, but the methods of the present disclosure are not limited to the order or arrangement shown. Various steps of method 800 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.
[0091] In 802, the computing system may retrieve multiple result images based on the similarity between an intermediate representation of a query image and each of the intermediate representations associated with each of the multiple result images. In some implementations, retrieving multiple result images may involve processing the query image with a machine learning-based visual search model to obtain an intermediate representation of the query image (e.g., an embedding, encoding, latent representation, etc.). The computing system may retrieve the result images based on the degree of similarity between the intermediate representation of the query image and the intermediate representations of the multiple result images.
[0092] In some implementations, processing a query image with a machine learning-based visual search model may involve processing the query image with a machine learning-based embedding model to obtain a query image embedding for the query image. The computing system can then search for multiple result images based on the distance between the query image embedding and the embeddings of multiple result images within the embedding space.
[0093] Prior to processing the query image, the operation includes obtaining the query image from the user computing device. For example, a user can use their computing device to capture an image showing an unfamiliar object. To learn more about that object, the user may use a visual search service by providing the computing system with the image and associated prompts (e.g., "What is this object?"). Alternatively, in some implementations, the computing system may receive the image and associated prompts from an automated service or software program. For example, an indexing service may provide the computing system with an image that has associated prompts corresponding to an indexing task (e.g., "Which key keywords should be associated with this image?"). Additionally or alternatively, in some implementations, the user computing device may automatically capture the image, generate prompts, and send the image and prompts to the computing system. For example, the user computing device may be a wearable augmented reality (AR) / virtual reality (VR) device. The user computing device can capture an image of an object and send the image to the computing system along with automatically generated prompts (e.g., "Identify this object and provide relevant summary information"). The user computing device may then display such information in an AR / VR context.
[0094] In 804, a computing system may identify multiple source documents. Each of the multiple source documents may include one of several resulting images and the text content associated with that resulting image. In some implementations, identifying multiple source documents may further include obtaining attribute information for each of the source documents. The attribute information may include (a) identifying information for the source document (e.g., title, citation, numerical identifier such as a Digital Object Identifier (DOI)), and / or (b) information describing a location from which the source document can be accessed (e.g., file path, link to download or purchase an application, URL, hotlink, API call to a library or other information repository that may hold a physical copy of the document, etc.).
[0095] In 806, the computing system may determine multiple first text units for multiple resulting images. Each first text unit may contain at least a portion of the text content associated with the resulting image from one or more source documents containing the resulting image. For example, suppose the first document containing the first resulting image is an online article for a popular blog. In some cases, if the first resulting image is one of many images contained in the article, only the text content closest to the first resulting image in the article is relatively likely to be relevant to the first resulting image, and therefore the computing system may decide to include the text content close to the first resulting image in the article in the first text unit. Alternatively, if the article is relatively small and contains only a few paragraphs, or only the first resulting image, the computing system may decide to include all of the article's text content in the first text unit.
[0096] Therefore, it should be understood that the computing system may use any prior art to determine which parts of the text content from the source document should be included in the first text unit. In some implementations, the computing system may process the text content of the source document with a machine learning model, such as a classification model, to predict the relevance of different parts of the text content to the resulting image. As an addition or alternative, in some implementations, the computing system may use a heuristic method to select the text content to be included in the first text unit. For example, the computing system may use a schema based on the following rules: IF doc_type == article; THEN search for statements X-5 to X+5, where X is the location of the image within the document; IF doc_length <= 1000 words; THEN Search all words.
[0097] In 808, a computing system may process a set of text inputs with a machine learning language model to obtain a language output. The language output may include a second text unit. The set of text inputs may include (a) two or more first text units, each associated with two or more of a set of result images, and (b) prompts associated with a query image.
[0098] In 810, the computing system may provide the user computing device with a second text unit and two or more resulting images for display within the user computing device's interface. In some implementations, providing the user computing device with a second text unit and two or more resulting images for display within the user computing device's interface includes providing the user computing device with interface data. The interface data may include instructions for generating (a) an interface element containing the second text unit and (b) two or more selectable attribute elements, each associated with one or more of the two or more resulting images. Each selectable attribute element may include the associated resulting image or some representation of the resulting image, such as a thumbnail. The selectable attribute elements may also include attribute information about one or more source documents containing the associated resulting images.
[0099] In some implementations, the computing system may receive data from the user computing device indicating the user's selection of a first selectable attribute element from two or more selectable attribute elements. The first selectable attribute element may be associated with a first result image from two or more result images. The computing system may identify a first text unit from two or more first text units, which include at least a portion of the text content from a source document containing the first result image. The computing system may remove the first text unit from the set of text inputs to obtain a second set of text inputs. The computing system may process the second set of text inputs with a machine learning language model to obtain a second language output containing a refined second text unit. The computing system may provide the refined second text unit to the user computing device.
[0100] In some implementations, removing a first text unit from a set of text inputs to obtain a second set of text inputs further includes removing information associated with the source document containing the first resulting image from the attribute information to obtain refined attribute information. Providing the refined second text unit to the user computing device may further include providing the refined attribute information to the user computing device.
[0101] In some implementations, the language output may further include predictive information that predicts a portion of a second text unit as most relevant to the prompt. The interface data may further include instructions for generating emphasis elements that highlight that portion of the second text unit.
[0102] In some implementations, providing a second text unit and two or more resulting images to a user computing device for display within the user computing device's interface may include providing interface data to the user computing device. The interface data may include a first interface element, a second interface element, and instructions for generating first and second attribute elements. The first interface element may include a first portion of the second text unit. The first portion of the second text unit may be associated with a first resulting image among two or more resulting images. For example, if the second text unit is a summary of a first document containing a first resulting image and a second document containing a second resulting image, the first portion of the second text unit may be the portion summarizing the first document. Similarly, the second interface element may include a second portion of the second text unit. The second portion of the second text unit may be associated with a second resulting image among two or more resulting images. The first selectable attribute element may include a thumbnail of the first result image, the result image itself, or an image derived from the result image, and may include attribute information about the source document containing the first result image. The second selectable attribute element may include a second result image (or a thumbnail or an image derived therefrom) and attribute information about the source document containing the second result image.
[0103] Figure 9 shows a flowchart of an exemplary method 900 for refining visual search information based on user feedback, according to an exemplary embodiment of the present disclosure. While Figure 9 shows steps performed in a specific order for illustrative and explanatory purposes, the methods of the present disclosure are not limited to the order or arrangement shown. Various steps of method 900 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.
[0104] In 902, the computing system may retrieve two or more result images based on the similarity between an intermediate representation of a query image and intermediate representations of two or more result images. For example, the intermediate representation may be an image embedding, and the computing system may retrieve two or more result images based on the distance between the image embedding and the image embeddings for two or more result images in the embedding space.
[0105] In 904, the computing system may process a set of text inputs with a machine learning language model to obtain a language output containing text content. The set of text inputs may include text content from a source document containing two or more resulting images, and prompts associated with a query image.
[0106] In some implementations, processing a set of text inputs with a machine learning language model may involve obtaining attribute information for each of the source documents. This attribute information may include (a) identifying information that identifies the source document, and / or (b) information that describes the location from which the source document can be accessed.
[0107] In 906, the computing system may provide the user computing device with language output and two or more resulting images for display within the user computing device's interface. For example, if the user computing device is running a visual search application associated with a visual search service provided by the computing system, the computing system may provide the language output and resulting images for display within the interface of the visual search application. In some implementations, the computing system may also provide attribute information.
[0108] In some implementations, providing language output and two or more resulting images to a user computing device for display within the user computing device's interface may include providing interface data to the user computing device. The interface data may include instructions for generating (a) an interface element containing the language output and (b) two or more selectable attribute elements, each associated with one or more resulting images. Each attribute element may include a thumbnail of the associated resulting image and attribute information about one or more source documents containing the associated resulting image.
[0109] In 908, the computing system may receive information from the user computing device describing an instruction by the user of the user computing device that a first result image among two or more result images does not visually resemble the query image. In some implementations, receiving information describing an instruction by the user of the user computing device that a first result image among two or more result images does not visually resemble the query image may include receiving data indicating the user of the user computing device's selection of a first selectable attribute element among two or more selectable attribute elements. The first selectable attribute element may be associated with the first result image among two or more result images.
[0110] In 910, the computing system may remove text content associated with the source document containing the first result image from the set of text inputs. In some implementations, removing text content associated with the source document containing the first result image from the set of text inputs further includes removing information associated with the source document containing the first result image from the attribute information to obtain refined attribute information. In some implementations, providing the refined language output to the user computing device further includes providing the refined attribute information to the user computing device.
[0111] In 912, a computing system may process a set of text inputs with a machine learning language model to obtain refined language output.
[0112] In 914, the computing system may provide the user computing device with refined language output for display within the interface of the user computing device.
[0113] Figure 10 shows a flowchart of an exemplary method 1000 for conducting user feedback collection for refining visual search information, according to an exemplary embodiment of the present disclosure. Figure 10 shows the steps to be performed in a particular order for illustrative and explanatory purposes, but the methods of the present disclosure are not limited to the order or arrangement shown. Various steps of method 1000 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.
[0114] In 1002, the user computing device may acquire a query image. In some implementations, acquiring a query image involves acquiring an input indicating a request to acquire an image using an image acquisition device associated with the user computing device. In response to acquiring the input, the user computing device may acquire the query image using an image acquisition device associated with the user computing device.
[0115] In 1004, the user computing device may obtain text data describing a prompt. In some implementations, obtaining text data describing a prompt may include obtaining spoken utterances from the user by an audio capture device associated with the user computing device. The user computing device may determine text data describing a prompt based at least partially on the spoken utterances. For example, the user computing device may process the spoken utterances with a machine learning speech recognition model to obtain text data.
[0116] In 1006, the user computing device may provide the computing system with a query image and text data describing a prompt. For example, the computing system may be a system associated with a visual search service, such as a multimodal search service, which provides information in response to a multimodal query including an image and an associated prompt.
[0117] In 1008, the user computing device may, in response to providing a query image and a prompt, receive from the computing system (a) two or more result images and (b) language output from a machine learning language model. The language output is generated based on the prompt and the text content from a source document containing the two or more result images.
[0118] In 1010, the user computing device may display, within the interface of an application executed by the user computing device, (a) an interface element containing language output, and (b) two or more selectable attribute elements, each associated with two or more result images. Each selectable attribute element includes a thumbnail of the associated result image and attribute information identifying the source document containing the associated result image.
[0119] In 1012, the user computing device may receive input from the user via an input device associated with the user computing device, in which the user selects a first selectable attribute element from two or more selectable attribute elements.
[0120] In some implementations, each selectable attribute element may include a first selectable portion and a second selectable portion. The user computing device receives input from the user via an input device associated with the user computing device, selecting the first selectable portion of the first selectable attribute element from among two or more selectable attribute elements. In response to receiving input to the first selectable portion of the first selectable attribute element, the user computing device may provide the computing system with information indicating the selection of the first selectable attribute element.
[0121] Alternatively, in some implementations, a user computing device may receive input from the user via an input device associated with the user computing device, selecting a second selectable portion of a first selectable attribute element from two or more selectable attribute elements. In response to receiving input selecting a second selectable portion of a first selectable attribute element, the user computing device may trigger the display of a source document identified by the attribute information contained in the first selectable attribute element. For example, if the source document is a website, the user computing device may run a web browser application and navigate to the website. In another example, if the source document is a PDF, the user computing device may run a PDF reader application and open the PDF.
[0122] In 1014, the user computing device may, in response to receiving input, provide the computing system with information indicating the selection of a first selectable attribute element.
[0123] In 1016, the user computing device may receive refined language output from the computing system in response to having provided information. The refined language output may be generated based on a prompt and text content from a source document, including two or more result images other than a first result image associated with a first selectable attribute element.
[0124] Figure 11A shows a block diagram of an exemplary computing system 1100 that performs a visual or multimodal search service according to an exemplary embodiment of the present disclosure. The system 1100 includes a user computing system 1102, a server computing system 1130, and / or a third-party computing system 1150, which are communicably coupled over a network 1180.
[0125] The user computing system 1102 may include any type of computing device, such as a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0126] The user computing system 1102 includes one or more processors 1112 and memory 1114. The one or more processors 1112 may be any suitable processing device (e.g., a processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be one processor or multiple processors operably connected. The memory 1114 may include one or more non-temporary computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 1114 can store data 1116 and instructions 1118 executed by the processors 1112 to cause the user computing system 1102 to perform operations.
[0127] In some implementations, the user computing system 1102 may store or include one or more machine learning models 1120. For example, the machine learning models 1120 may be a variety of machine learning models, including neural networks (e.g., deep neural networks) or other types of machine learning models including nonlinear and / or linear models, or may otherwise include such machine learning models. Neural networks may include feedforward neural networks, recurrent neural networks (e.g., long-short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks.
[0128] In some implementations, one or more machine learning models 1120 may be received from a server computing system 1130 via a network 1180, stored in user computing device memory 1114, and then used or otherwise implemented by one or more processors 1112. In some implementations, the user computing system 1102 may implement multiple parallel instances of a single machine learning model 1120 (for example, to perform parallel machine learning model processing across multiple instances of input data and / or detected features).
[0129] More specifically, one or more machine learning models 1120 may include one or more detection models, one or more classification models, one or more segmentation models, one or more expansion models, one or more generative models, one or more natural language processing models, one or more optical character recognition models, and / or one or more other machine learning models. One or more machine learning models 1120 may include one or more transformer models. One or more machine learning models 1120 may include one or more neural radiative field models, one or more diffusion models, and / or one or more autoregressive language models.
[0130] One or more machine learning models 1120 can be used to detect one or more object features. Detected object features can be classified and / or embedded. Classification and / or embedding may then be used to perform a search to determine one or more search results. Alternatively and / or additionally, one or more detected features may be used to determine that an indicator (e.g., a user interface element indicating a detected feature) should be provided to indicate that the feature has been detected. The user may then select an indicator to perform feature classification, embedding, and / or search. In some implementations, classification, embedding, and / or search may be performed before an indicator is selected.
[0131] In some implementations, one or more machine learning models 1120 can process image data, text data, audio data, and / or latent coding data to generate output data which may include image data, text data, audio data, and / or latent coding data. One or more machine learning models 1120 can perform optical character recognition, natural language processing, image classification, object classification, text classification, audio classification, contextual determination, action prediction, image correction, image magnification, text magnification, sentiment analysis, object detection, error detection, repair, video stabilization, audio correction, audio magnification, and / or data segmentation (e.g., mask-based segmentation).
[0132] As an addition or alternative, one or more machine learning models 1140 may be included in, or otherwise stored and implemented by, a server computing system 1130 that communicates with a user computing system 1102 according to a client-server relationship. For example, a machine learning model 1140 may be implemented by the server computing system 1130 as part of a web service (e.g., a viewfinder service, a visual search service, an image processing service, an ambient computing service, and / or an overlay application service). Thus, one or more models 1120 may be stored and implemented in the user computing system 1102, and / or one or more models 1140 may be stored and implemented in the server computing system 1130.
[0133] The user computing system 1102 may also include one or more user input components 1122 that receive user input. For example, a user input component 1122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). A touch-sensitive component may be useful for implementing a virtual keyboard. Other exemplary user input components include a microphone, a conventional keyboard, or other means by which a user can provide user input.
[0134] In some implementations, a user computing system may store and / or provide one or more user interfaces that may be associated with one or more applications. One or more user interfaces may be configured to receive input and / or provide data for display (e.g., image data, text data, audio data, one or more user interface elements, augmented reality experiences, virtual reality experiences, and / or other data for display). A user interface may be associated with one or more other computing systems (e.g., server computing system 1130 and / or third-party computing system 1150). A user interface may include a viewfinder interface, a search interface, a generative model interface, a social media interface, a media content gallery interface, and the like.
[0135] The user computing device 1102 may include and / or receive data from one or more sensors 1126. One or more sensors 1126 may be housed in a housing component that houses one or more processors 1112, memory 1114, and / or one or more hardware components that can store and / or execute one or more software packets. One or more sensors 1126 may include one or more image sensors (e.g., cameras), one or more LiDAR sensors, one or more audio sensors (e.g., microphones), one or more inertial sensors (e.g., inertial measurement units), one or more biosensors (e.g., heart rate sensors, pulse sensors, retinal sensors, and / or fingerprint sensors), one or more infrared sensors, one or more location sensors (e.g., GPS), one or more touch sensors (e.g., conductive touch sensors and / or mechanical touch sensors), and / or one or more other sensors. One or more sensors can be used to acquire data associated with the user's environment (e.g., images of the user's environment, recordings of the environment, and / or the user's location).
[0136] The user computing system 1102 includes and / or is part of a user computing device 1104. The user computing device 1104 may include a mobile computing device (e.g., a smartphone or tablet), a desktop computer, a laptop computer, a smart wearable, and / or a smart device. Additionally and / or alternatively, the user computing system may acquire and / or generate data from one or more user computing devices 1104. For example, a smartphone camera may be used to capture image data describing the environment, and / or an overlay application on the user computing device 1104 may be used to track and / or process data provided to the user. Similarly, one or more sensors associated with a smart wearable may be used to acquire data about the user and / or the user's environment (e.g., image data may be acquired by a camera housed in the user's smart glasses). Additionally and / or alternatively, data may be acquired and uploaded from other user devices that may be specialized for data acquisition or generation.
[0137] The server computing system 1130 includes one or more processors 1132 and memory 1134. The one or more processors 1132 may be any suitable processing device (e.g., a processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be one processor or multiple processors operably connected. The memory 1134 may include one or more non-temporary computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof. The memory 1134 can store data 1136 and instructions 1138 executed by the processors 1132 to cause the server computing system 1130 to perform operations.
[0138] In some implementations, the server computing system 1130 includes one or more server computing devices, or is otherwise implemented by server computing devices. In cases where the server computing system 1130 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or any combination thereof.
[0139] As described above, the server computing system 1130 may store or otherwise include one or more machine learning models 1140. For example, the models 1140 may be various machine learning models or otherwise include them. Exemplary machine learning models include neural networks or other multilayer nonlinear models. Exemplary neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks.
[0140] In addition and / or alternatively, the server computing system 1130 may include and / or be communicably connected to a search engine 1142 that can be used to crawl one or more databases (and / or resources). The search engine 1142 can process data from the user computing system 1102, the server computing system 1130, and / or the third-party computing system 150 to determine one or more search results associated with the input data. The search engine 1142 may perform term-based search, label-based search, Boolean-based search, image search, embedding-based search (e.g., nearest neighbor search), multimodal search, and / or one or more other search techniques.
[0141] The server computing system 1130 can store and / or provide one or more user interfaces 1144 for acquiring input data and / or providing output data to one or more users. One or more user interfaces 1144 may include one or more user interface elements, which may include input fields, navigation tools, content chips, selectable tiles, widgets, data display carousels, dynamic animations, information popups, image zoom, text-to-speech, voice-to-text, augmented reality, virtual reality, feedback loops, and / or other interface elements.
[0142] The user computing system 1102 and / or the server computing system 1130 can train models 1120 and / or 1140 through interaction with a third-party computing system 1150, which is communicably coupled via a network 1180. The third-party computing system 1150 may be separate from the server computing system 1130 or may be part of the server computing system 1130. Alternatively and / or additionally, the third-party computing system 1150 may be associated with one or more web resources, one or more web platforms, one or more other users, and / or one or more contexts.
[0143] The third-party computing system 1150 may include one or more processors 1152 and memory 1154. The one or more processors 1152 may be any suitable processing device (e.g., a processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be one processor or multiple processors operably connected. The memory 1154 may include one or more non-temporary computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 1154 can store data 1156 and instructions 1158 executed by the processors 1152 to cause the third-party computing system 1150 to perform operations. In some implementations, the third-party computing system 1150 includes one or more server computing devices, or is otherwise implemented by server computing devices.
[0144] Network 1180 may be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or any combination thereof, and may include any number of wired or wireless links. Generally, communication over Network 1180 may be carried over any type of wired and / or wireless connection using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encoding or formatting (e.g., HTML, XML), and / or protection methods (e.g., VPN, Secure HTTP, SSL).
[0145] The machine learning models described herein may be used in a variety of tasks, applications, and / or use cases.
[0146] In some implementations, the input to the machine learning model of this disclosure may be image data. The machine learning model may process the image data to generate an output. For example, the machine learning model may process the image data to generate an image recognition output (e.g., recognition of image data, latent embedding of image data, encoded representation of image data, hash of image data, etc.). Another example is that the machine learning model may process the image data to generate an image segmentation output. Another example is that the machine learning model may process the image data to generate an image classification output. Another example is that the machine learning model may process the image data to generate an image data modification output (e.g., modification of image data, etc.). Another example is that the machine learning model may process the image data to generate an encoded image data output (e.g., encoded and / or compressed representation of image data, etc.). Another example is that the machine learning model may process the image data to generate an upscaled image data output. Another example is that the machine learning model may process the image data to generate a prediction output.
[0147] In some implementations, the input to the machine learning model of this disclosure may be text or natural language data. The machine learning model may process the text or natural language data to produce an output. For example, the machine learning model may process natural language data to produce a language coding output. As another example, the machine learning model may process text or natural language data to produce a latent text embedding output. As yet another example, the machine learning model may process text or natural language data to produce a transformation output. As yet another example, the machine learning model may process text or natural language data to produce a classification output. As yet another example, the machine learning model may process text or natural language data to produce a text segmentation output. As yet another example, the machine learning model may process text or natural language data to produce a semantic intent output. As yet another example, the machine learning model may process text or natural language data to produce an upscaled text or natural language output (e.g., text or natural language data that is of higher quality than the input text or natural language). As yet another example, the machine learning model may process text or natural language data to produce a predictive output.
[0148] In some implementations, the input to the machine learning model of this disclosure may be audio data. The machine learning model may process the audio data to produce an output. For example, the machine learning model may process the audio data to produce a speech recognition output. As another example, the machine learning model may process the audio data to produce a speech translation output. As yet another example, the machine learning model may process the audio data to produce a latent embedding output. As yet another example, the machine learning model may process the audio data to produce an encoded audio output (e.g., an encoded and / or compressed representation of the audio data). As yet another example, the machine learning model may process the audio data to produce an upscaled audio output (e.g., audio data of higher quality than the input audio data). As yet another example, the machine learning model may process the audio data to produce a text representation output (e.g., a text representation of the input audio data). As yet another example, the machine learning model may process the audio data to produce a prediction output.
[0149] In some implementations, the input to the machine learning model of this disclosure may be sensor data. The machine learning model may process the sensor data to generate an output. For example, the machine learning model may process the sensor data to generate a recognition output. As another example, the machine learning model may process the sensor data to generate a prediction output. As yet another example, the machine learning model may process the sensor data to generate a classification output. As yet another example, the machine learning model may process the sensor data to generate a segmentation output. As yet another example, the machine learning model may process the sensor data to generate a segmentation output. As yet another example, the machine learning model may process the sensor data to generate a visualization output. As yet another example, the machine learning model may process the sensor data to generate a diagnostic output. As yet another example, the machine learning model may process the sensor data to generate a detection output.
[0150] In some cases, the input includes visual data, and the task is a computer vision task. In some cases, the input includes pixel data for one or more images, and the task is an image processing task. For example, the image processing task may be image classification, where the output is a set of scores, each corresponding to a different object class, and representing the likelihood that one or more images depict an object belonging to that object class. The image processing task may be object detection, where the image processing output identifies one or more regions in one or more images, and for each region, the likelihood that the region depicts the object of interest. As another example, the image processing task may be image segmentation, where the image processing output defines the likelihood for each pixel in one or more images for each category in a given set of categories. For example, the set of categories may be foreground and background. As yet another example, the set of categories may be object classes. As yet another example, the image processing task may be depth estimation, where the image processing output defines the depth value for each pixel in one or more images. As another example, the image processing task may be motion estimation, where the network input includes multiple images, and the image processing output defines the motion of the scene, shown in the pixels between the images in the network input, for each pixel of one of the input images.
[0151] The user computing system 1102 may include several applications (for example, applications 1 to N). Each application may include its own machine learning libraries and machine learning models. For example, each application may include a machine learning model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and so on. In some implementations, each application may communicate with a central intelligence layer (and the models stored therein) using an API (for example, a common API across all applications).
[0152] The central intelligence layer may contain several machine learning models. For example, each machine learning model (e.g., Model) can be dedicated to a specific application and managed by the central intelligence layer. In other implementations, two or more applications may share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model (e.g., Single Model) to all applications. In some implementations, the central intelligence layer is contained within the operating system of the computing system 1100, or otherwise implemented by the operating system.
[0153] The central intelligence layer can communicate with the central device data layer. The central device data layer may be a centralized repository of data for the computing system 1100. The central device data layer can communicate with several other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0154] Figure 11B shows a block diagram of an exemplary computing system 1250 that performs a visual search operation and / or refinement of visual search information according to an exemplary embodiment of the present disclosure. In particular, the exemplary computing system 1250 may include one or more computing devices 1252 that can be used to acquire and / or generate one or more datasets that can be processed by a sensor processing system 1260 and / or output determination system 1280, which can provide feedback to the user about features in one or more acquired datasets. One or more datasets may include image data, text data, audio data, multimodal data, latent coding data, etc. One or more datasets may be acquired by one or more sensors associated with one or more computing devices 1252 (for example, one or more sensors in computing device 1252). Additionally and / or alternatively, one or more datasets may be stored data and / or retrieved data (for example, data retrieved from a web resource). For example, images, text, and / or other content items can interact with the user. The interacted-with content items can then be used to generate one or more decisions.
[0155] One or more computing devices 1252 can acquire and / or generate one or more datasets based on image acquisition, sensor tracking, data storage retrieval, content download (e.g., downloading images or other content items from web resources via the internet), and / or by one or more other techniques. One or more datasets can be processed by a sensor processing system 1260. The sensor processing system 1260 may implement one or more processing techniques using one or more machine learning models, one or more search engines, and / or one or more other processing techniques. One or more processing techniques may be implemented in any combination and / or individually. One or more processing techniques may be implemented sequentially and / or in parallel. In particular, one or more datasets can be processed by a context determination block 1262, which may determine the context associated with one or more content items. The context determination block 1262 may determine a specific context associated with a user by identifying and / or processing metadata, user profile data (e.g., preferences, user search history, user browsing history, user purchase history, and / or user input data), previous interaction data, global trend data, location data, time data, and / or other data. The context may be associated with an event, determined trend, specific action, specific type of data, specific environment, and / or another context associated with the user and / or searched or retrieved data.
[0156] The sensor processing system 1260 may include an image preprocessing block 1264. The image preprocessing block 1264 may be used to adjust one or more values of the acquired and / or received image to prepare the image for processing by one or more machine learning models and / or one or more search engines 1274. The image preprocessing block 1264 may resize the image, adjust saturation values, adjust resolution, remove and / or add metadata, and / or perform one or more other operations.
[0157] In some implementations, the sensor processing system 1260 may include one or more machine learning models, which may include a detection model 1266, a segmentation model 1268, a classification model 1270, an embedding model 1272, and / or one or more other machine learning models. For example, the sensor processing system 1260 may include one or more detection models 1266 that can be used to detect specific features in a processed dataset. In particular, one or more images may be processed by one or more detection models 1266 so that they generate one or more bounding boxes associated with detected features in one or more images.
[0158] As an addition and / or alternative, one or more segmentation models 1268 can be used to segment one or more portions of a dataset from one or more datasets. For example, one or more segmentation models 1268 can use one or more segmentation masks (e.g., one or more manually generated and / or generated based on one or more bounding boxes) to segment a portion of an image, a portion of an audio file, and / or a portion of text. Segmentation may include separating one or more detected objects from an image and / or removing one or more detected objects.
[0159] One or more classification models 1270 can be used to process image data, text data, audio data, latent coding data, multimodal data, and / or other data to generate one or more classifications. One or more classification models 1270 may include one or more image classification models, one or more object classification models, one or more text classification models, one or more audio classification models, and / or one or more other classification models. One or more classification models 1270 can process data to determine one or more classifications.
[0160] In some implementations, data may be processed by one or more embedding models 1272 to generate one or more embeddings. For example, one or more images may be processed by one or more embedding models 1272 to generate one or more image embeddings in the embedding space. One or more image embeddings may be associated with one or more image features of one or more images. In some implementations, one or more embedding models 1272 may be configured to process multimodal data to generate multimodal embeddings. One or more embeddings can be used for classification, search, and / or training of embedding space distributions.
[0161] The sensor processing system 1260 may include one or more search engines 1274 that can be used to perform one or more searches. One or more search engines 1274 may crawl one or more databases (e.g., one or more local databases, one or more global databases, one or more private databases, one or more public databases, one or more specialized databases, and / or one or more general databases) to determine one or more search results. One or more search engines 1274 may perform feature matching, text-based search, embedding-based search (e.g., k-nearest neighbor search), metadata-based search, multimodal search, web resource search, image search, text search, and / or application search.
[0162] As an addition and / or alternative, the sensor processing system 1260 may include one or more multimodal processing blocks 1276 which can be used to assist in the processing of multimodal data. One or more multimodal processing blocks 1276 may include generating multimodal queries and / or multimodal embeddings to be processed by one or more machine learning models and / or one or more search engines 1274.
[0163] The output of the sensor processing system 1260 may then be processed by the output determination system 1280 to determine one or more outputs to be provided to the user. The output determination system 1280 may include heuristic-based determination, machine learning model-based determination, user selection-based determination, and / or context-based determination.
[0164] The output determination system 1280 may determine how and / or where one or more search results should be provided in the search result interface 1282. Additionally and / or alternatively, the output determination system 1280 may determine how and / or where one or more machine learning model outputs should be provided in the machine learning model output interface 1284. In some implementations, one or more search results and / or one or more machine learning model outputs may be provided for display by one or more user interface elements. One or more user interface elements may be overlaid on the displayed data. For example, one or more detection indicators may be overlaid on detected objects in the viewfinder. One or more user interface elements may be selectable for performing one or more additional searches and / or one or more additional machine learning model processes. In some implementations, user interface elements may be provided as specialized user interface elements for a particular application and / or uniformly provided across different applications. One or more user interface elements may include pop-up displays, interface overlays, interface tiles and / or chips, carousel interfaces, audio feedback, animations, interactive widgets, and / or other user interface elements.
[0165] As an addition and / or alternative, data associated with the output of the sensor processing system 1260 can be used to generate and / or provide an augmented reality experience and / or virtual reality experience 1286. For example, one or more acquired datasets may be processed to generate one or more augmented reality rendering assets and / or one or more virtual reality rendering assets, which can then be used to provide the augmented reality experience and / or virtual reality experience 1286 to the user. The augmented reality experience may render information associated with the environment into its respective environment. As an alternative and / or addition, objects associated with the processed datasets may be rendered into the user environment and / or virtual environment. Rendering of dataset generation may include training one or more neural radiative field models to learn three-dimensional representations for one or more objects.
[0166] In some implementations, one or more action prompts 1288 may be determined based on the output of the sensor processing system 1260. For example, a search prompt, a purchase prompt, a generation prompt, a reservation prompt, a call prompt, a redirect prompt, and / or one or more other prompts may be determined to be associated with the output of the sensor processing system 1260. One or more action prompts 1288 may then be provided to the user via one or more selectable user interface elements. In response to the selection of one or more selectable user interface elements, each action of each action prompt may be performed (for example, a search may be performed, a purchase application programming interface may be used, and / or another application may be opened).
[0167] In some implementations, one or more datasets and / or the output of a sensor processing system 1260 may be processed by one or more generative models 1290 to generate model-generated content items, which may then be provided to the user. Generation may be prompted based on user selection and / or performed automatically (for example, automatically based on one or more conditions, which may be associated with thresholds of unidentified search results).
[0168] The output determination system 1280 may process one or more datasets and / or the output of the sensor processing system 1260 in the data augmentation block 1292 to generate augmented data. For example, one or more images can be processed in the data augmentation block 1292 to generate one or more augmented images. Data augmentation may include data correction, data cropping, deletion of one or more features, addition of one or more features, resolution adjustment, illumination adjustment, saturation adjustment, and / or other augmentation.
[0169] In some implementations, one or more datasets and / or outputs of the sensor processing system 1260 may be stored based on the judgment of the data storage block 1294.
[0170] The output of the output determination system 1280 may then be provided to the user by one or more output components of the user computing device 1252. For example, one or more user interface elements associated with one or more outputs may be provided for display by the visual display of the user computing device 1252.
[0171] The process may be performed repeatedly and / or sequentially. One or more user inputs to the provided user interface elements may condition and / or influence the subsequent processing loop.
[0172] The technologies described herein refer to servers, databases, software applications, and other computer-based systems, as well as actions performed and information transmitted to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functionalities among their components. For example, the processes described herein may be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0173] This subject matter has been described in detail with respect to various specific exemplary embodiments, each example provided for illustrative purposes only and not as an limitation of the disclosure. Those skilled in the art, upon understanding the above, will readily be able to generate modifications, variations, and equivalents to such embodiments. Therefore, this disclosure does not preclude the inclusion of such modifications, variations, and / or additions to the subject matter, so as will be readily apparent to those skilled in the art. For example, a feature illustrated or described as part of one embodiment may also be used in conjunction with another embodiment to generate further embodiments. Thus, this disclosure is intended to include such alternatives, variations, and equivalents.
[0174] Embodiment The following describes some embodiments of the present disclosure. However, it should be noted that the following embodiments are not an exhaustive list of all embodiments of the present disclosure. Rather, the following embodiments are provided to illustrate various scenarios in which embodiments of the present disclosure may be used.
[0175] Embodiment 1: A method implemented by a computer, - A computing system comprising one or more processor devices searches for a result image based on the similarity between a query image and a result image, - A step of obtaining a first text unit by a computing system, wherein the first text unit includes at least a portion of the text content of a source document including a result image, - A step in which a computing system determines a second text unit in response to a prompt associated with a query image, wherein the second text unit is: ○(a) At least a portion of the first text unit, or ○(b) A step including one or more of the texts derived from the first text unit, - A method comprising the steps of providing a second text unit and a resulting image for display within an interface by a computing system.
[0176] Embodiment 2: The method of Embodiment 1, wherein the step of searching for a result image includes the steps of: processing a query image with a machine learning-type visual search model using a computing system to obtain an intermediate representation of the query image; and searching for a result image using a computing system based on the degree of similarity between the intermediate representation of the query image and the intermediate representation of the result image.
[0177] Embodiment 3: A computer-based method of Embodiment 2, wherein the step of processing a query image with a machine learning visual search model includes the step of a computing system processing the query image with a machine learning embedding model to obtain a query image embedding for the query image, and the step of searching for a result image includes the step of a computing system searching for a result image based on the distance between the query image embedding and the result image embedding within the embedding space.
[0178] Embodiment 4: A method performed by a computer of Embodiment 1, the method comprising the step of obtaining a query image from a user computing device by a computing system, prior to processing the query image.
[0179] Embodiment 5: The method implemented by the computer of Embodiment 4, wherein the interface includes a user interface for an application run by a user computing device.
[0180] Embodiment 6: A method performed by a computer of Embodiment 4, wherein the step of acquiring a query image includes the step of the computing system acquiring the query image and the prompt associated with the query image from a user computing device.
[0181] Embodiment 7: A method performed by a computer of Embodiment 4, wherein the step of searching for a result image further includes the steps of providing the result image to a user computing device for display within an interface by the computing system, and receiving a prompt associated with the query image from the user computing device in response to the provision of the result image.
[0182] Embodiment 8: A method performed by a computer according to Embodiment 1, wherein the step of determining a second text unit in response to a prompt associated with a query image includes the step of a computing system processing the second text unit and the prompt associated with the query image with a machine learning language model to obtain a language output containing the second text unit.
[0183] Embodiment 9: The computer-based method of Embodiment 8, wherein the second text unit includes a subset of the first text unit.
[0184] Embodiment 10: The computer-based method of Embodiment 8, wherein the second text unit includes text derived from the first text unit, and the text derived from the first text unit describes a summary of the first text unit.
[0185] Embodiment 11: The source document is, - One or more web pages of a website, - article, - Newspaper, - Book, or - A computer-based method of Embodiment 1, including a transcript.
[0186] Embodiment 12: A method performed by a computer according to Embodiment 1, wherein the step of providing a second text unit and a resulting image further includes the step of providing attribute information for display in an interface by a computing system, which (a) identifies a source document and / or (b) indicates a location accessible from the source document.
[0187] Embodiment 13: A method implemented by a computer according to Embodiment 12, wherein the source document includes a web page and the attribute information includes the address of the web page.
[0188] Embodiment 14: A computer-based method of Embodiment 12, wherein the source document includes a journal, and the attribute information includes a citation indicating the location of the resulting image within the journal.
[0189] Embodiment 15: A method performed by a computer of Embodiment 1, wherein, prior to determining a second text unit, the method includes the step of a computing system generating a prompt associated with the query image, at least in part, based on the query image.
[0190] Embodiment 16: A computer-based method of Embodiment 15, wherein the step of generating a prompt associated with a query image includes the steps of: the computing system processing the query image with a machine learning model to generate a semantic output describing the image; and the computing system generating a prompt based at least partially on the semantic output.
[0191] Embodiment 17: A computing system, - One or more processors, - When executed by one or more processors, it comprises one or more non-temporary computer-readable media that collectively store a first set of instructions that cause a computing system to perform an action, and the action is, ○ Obtain the query image and associated prompt from the user computing device, ○Processing query images with a machine learning-based embedding model to obtain query image embeddings, ○ Searching for result images based on the similarity between the query image embedding and the result image embedding, ○ Identifying the source document for the resulting image, wherein the source document includes the resulting image and the text content associated with the resulting image. ○ Determine a first text unit from the source document that includes at least a portion of the text content associated with the resulting image, ○Processing the first text unit and prompt with a machine learning language model to obtain a language output including a second text unit, wherein the second text unit is: (a) at least a portion of the first text unit, or (b) containing one or more of the texts derived from the first text unit, A computing system that includes providing a second text unit and a resulting image for display within the interface of an application run by a user computing device.
[0192] Embodiment 18: The operation is, - Receiving information from the user computing device indicating a request for additional information, - Searching for a second result image based on the similarity between the query image embedding and the second result image embedding, - Identifying a first source document and a second source document for the resulting image, wherein each of the first and second source documents includes the resulting image and text content associated with the resulting image, and the text content associated with the resulting image in the first source document is different from the text content associated with the resulting image in the second source document. - Determining an additional first text unit from one or more of the first or second source documents, which includes at least a portion of the text content associated with the resulting image, - Processing an additional first text unit and prompt with a machine learning language model to obtain a second language output including an additional second text unit, the additional second text unit is: ○(a) At least a portion of the additional first text unit, or (b) containing one or more of the text derived from the additional first text unit, - A computing system according to embodiment 17, further comprising providing an additional second text unit and a second resulting image for display within the interface of an application run by a user computing device.
[0193] Embodiment 19: The computing system of Embodiment 17, further comprising providing a second text unit and a resulting image, which also includes providing attribute information identifying the source document for display within the interface of an application run by a user computing device.
[0194] Embodiment 20: One or more non-temporary computer-readable media for storing a first set of instructions together, wherein when an instruction is executed by one or more processors, it causes one or more processors to perform an action, and the action is - Searching for result images based on the similarity between the query image and the result image, - To obtain a first text unit, wherein the first text unit includes at least a portion of the text content of the source document containing the resulting image. - Determining a second text unit in response to a prompt associated with the query image, wherein the second text unit is: ○(a) At least a portion of the first text unit, or ○(b) containing one or more of the texts derived from the first text unit, - One or more non-temporary computer-readable media, including providing a second text unit and a resulting image for display within the interface.
[0195] Embodiment 21: A computing system, - One or more processors, - When executed by one or more processors, it comprises one or more non-temporary computer-readable media that collectively store a first set of instructions that cause a computing system to perform an action, and the action is, ○ Searching for multiple result images based on the similarity between the intermediate representation of the query image and each of the multiple intermediate representations associated with each of the multiple result images, ○ Identifying multiple source documents, where each of the multiple source documents includes one of the multiple result images, and the text content associated with that result image. ○ Determining multiple first text units for multiple result images, wherein each first text unit includes at least a portion of the text content associated with the result image from one or more source documents containing the result image. ○ Processing a set of text inputs with a machine learning language model to obtain a language output that includes a second text unit, wherein the set of text inputs is: (a) Two or more first text units associated with two or more of the multiple result images, and (b) The query includes a prompt associated with the image, A computing system comprising providing a second text unit and two or more resulting images to a user computing device for display within the interface of the user computing device.
[0196] Embodiment 22: A computing system of Embodiment 21, wherein searching for multiple result images includes processing a query image with a machine learning-type visual search model to obtain an intermediate representation of the query image, and searching for result images based on the degree of similarity between the intermediate representation of the query image and the intermediate representations of multiple result images.
[0197] Embodiment 23: The computing system of Embodiment 22, wherein processing a query image with a machine learning visual search model includes processing the query image with a machine learning embedding model to obtain a query image embedding for the query image, and searching for multiple result images includes searching for multiple result images based on the distance between the query image embedding and the embeddings of the multiple result images within the embedding space.
[0198] Embodiment 24: The computing system of Embodiment 21, wherein, prior to processing the query image, the operation includes obtaining the query image from a user computing device.
[0199] Embodiment 25: The computing system of Embodiment 21, wherein acquiring a query image includes acquiring a query image and a prompt associated with the query image from a user computing device.
[0200] Embodiment 26: The computing system of Embodiment 21, wherein identifying a plurality of source documents further comprises obtaining attribute information, and for each of the plurality of source documents, the attribute information includes (a) identification information that identifies the source document, and / or (b) information that describes a location from which the source document can be accessed.
[0201] Embodiment 27: Providing a user computing device with a second text unit and two or more resulting images for display within the interface of the user computing device comprises providing the user computing device with interface data, the computing system of Embodiment 26, wherein the interface data comprises instructions for generating (a) an interface element comprising the second text unit and (b) two or more selectable attribute elements each associated with two or more resulting images, each selectable attribute element comprising a thumbnail of the associated resulting image and attribute information about one or more source documents containing the associated resulting image.
[0202] Embodiment 28: The operation is, - Receiving data from a user computing device indicating the selection of a first selectable attribute element among two or more selectable attribute elements by the user of the user computing device, wherein the first selectable attribute element is associated with a first result image among two or more result images. - Identifying a first text unit among two or more first text units that include at least a portion of the text content from a source document containing a first result image, - Remove the first text unit from the set of text inputs to obtain the second set of text inputs, - Process a second set of text inputs with a machine learning language model to obtain a second language output containing refined second text units, A computing system according to embodiment 27, further comprising providing a refined second text unit to a user computing device.
[0203] Embodiment 29: The computing system of Embodiment 28, wherein removing a first text unit from a set of text inputs to obtain a second set of text inputs further comprises removing information associated with a source document containing a first result image from the attribute information to obtain refined attribute information, and providing the refined second text unit to a user computing device further comprises providing the refined attribute information to a user computing device.
[0204] Embodiment 30: The computing system of Embodiment 27, wherein the language output further includes predictive information that predicts a portion of a second text unit as most relevant to the prompt, and the interface data further includes instructions for generating emphasis elements that highlight that portion of the second text unit.
[0205] Embodiment 31: Providing a second text unit and two or more resulting images to a user computing device for display within the interface of the user computing device includes providing interface data to the user computing device, the interface data being: - A first interface element comprising a first portion of a second text unit, wherein the first portion of the second text unit is associated with a first result image among two or more result images, - A second interface element comprising a second portion of a second text unit, wherein the second portion of the second text unit is associated with a second result image among two or more result images, A computing system of Embodiment 26, comprising instructions for generating a first selectable attribute element and a second selectable attribute element, wherein the first selectable attribute element includes a thumbnail of a first result image and attribute information about a source document containing the first result image, and the second selectable attribute element includes a thumbnail of a second result image and attribute information about a source document containing the second result image.
[0206] Embodiment 32: The computing system of Embodiment 21, wherein the second text unit includes summaries of two or more first text units.
[0207] Embodiment 33: A method implemented by a computer, - A computing system including one or more computing devices searches for multiple result images based on the similarity between an intermediate representation of a query image and each of the multiple intermediate representations associated with each of the multiple result images. - A step of a computing system identifying multiple source documents, each of which includes a result image from among multiple result images and text content associated with that result image. - A step of determining, by a computing system, a plurality of first text units for a plurality of resulting images, wherein each first text unit includes at least a portion of the text content associated with the resulting image from one or more source documents containing the resulting image. - A computing system processes a set of text inputs with a machine learning language model to obtain a language output containing a second text unit, wherein the set of text inputs is: ○(a) Two or more first text units associated with two or more of the multiple result images, and ○(b) A step including a prompt associated with the query image, - A method comprising the steps of providing a second text unit and two or more resulting images to a user computing device for display within the interface of the user computing device by a computing system.
[0208] Embodiment 34: A computer-based method of Embodiment 33, wherein the step of searching for multiple result images includes the steps of: processing a query image with a machine learning-type visual search model using a computing system to obtain an intermediate representation of the query image; and searching for result images using a computing system based on the degree of similarity between the intermediate representation of the query image and the intermediate representations of multiple result images.
[0209] Embodiment 35: A computer-based method of Embodiment 34, wherein the step of processing a query image with a machine learning visual search model includes the step of a computing system processing the query image with a machine learning embedding model to obtain a query image embedding for the query image, and the step of searching for multiple result images includes the step of a computing system searching for multiple result images based on the distance between the query image embedding and the embeddings of the multiple result images within the embedding space.
[0210] Embodiment 36: A method performed by a computer according to Embodiment 33, which includes the step of obtaining a query image from a user computing device by a computing system prior to processing the query image.
[0211] Embodiment 37: A method performed by a computer according to Embodiment 33, wherein the step of acquiring a query image includes the step of a computing system acquiring a query image and a prompt associated with the query image from a user computing device.
[0212] Embodiment 38: The computer-based method of Embodiment 33, wherein the step of identifying a plurality of source documents further includes the step of obtaining attribute information by a computing system, wherein for each of the plurality of source documents, the attribute information includes (a) identification information that identifies the source document, and / or (b) information that describes a location from which the source document can be accessed.
[0213] Embodiment 39: A method performed by a computer of Embodiment 38, wherein the step of providing a user computing device to a second text unit and two or more resulting images for display within the interface of the user computing device includes the step of providing interface data to the user computing device by a computing system, the interface data including instructions for generating (a) an interface element comprising the second text unit and (b) two or more selectable attribute elements each associated with two or more resulting images, each attribute element comprising a thumbnail of the associated resulting image and attribute information about one or more source documents comprising the associated resulting image.
[0214] Embodiment 40: One or more non-temporary computer-readable media for storing a first set of instructions together, wherein when an instruction is executed by one or more processors, it causes one or more processors to perform an action, and the action is - Obtain the query image and associated prompt from the user computing device, - Processing query images with a machine learning-based embedding model to obtain query image embeddings, - Searching for result images based on the similarity between the query image embedding and the result image embedding, - Identifying the source document for the resulting image, wherein the source document includes the resulting image and the text content associated with the resulting image. - Determine a first text unit from the source document that includes at least a portion of the text content associated with the resulting image, - Processing a first text unit and prompt with a machine learning language model to obtain a language output including a second text unit, wherein the second text unit is: ○(a) At least a portion of the first text unit, or ○(b) containing one or more of the texts derived from the first text unit, - One or more non-temporary computer-readable media, including providing a second text unit and a resulting image for display within the interface of an application run by a user computing device.
[0215] Embodiment 41: A method carried out by a computer, - A computing system comprising one or more computing devices searches for two or more result images based on the similarity between an intermediate representation of a query image and intermediate representations of two or more result images. - A computing system processes a set of text inputs with a machine learning language model to obtain a language output containing text content, wherein the set of text inputs includes text content from a source document containing two or more result images, and prompts associated with a query image. - The computing system provides language output and two or more resulting images to the user computing device for display within the interface of the user computing device. - The computing system receives information from the user computing device describing instructions from the user of the user computing device that the first of two or more result images does not visually resemble the query image. - A computing system removes text content associated with a source document containing a first result image from a set of text inputs. - A computing system processes a set of text inputs with a machine learning language model to obtain refined language output. - A method comprising the steps of a computing system providing a user computing device with refined language output for display within the interface of the user computing device.
[0216] Embodiment 42: A computer-based method of Embodiment 41, wherein the step of searching for two or more result images includes the steps of: a computing system processing a query image with a machine learning-type visual search model to obtain an intermediate representation of the query image; and a computing system searching for result images based on the degree of similarity between the intermediate representation of the query image and the intermediate representations of two or more result images.
[0217] Embodiment 43: A method performed by a computer according to Embodiment 42, the method comprising the step of obtaining a query image from a user computing device by a computing system prior to processing the query image.
[0218] Embodiment 44: The computer-based method of Embodiment 41, wherein the step of processing a set of text inputs with a machine learning language model further includes the step of obtaining attribute information, wherein for each of the source documents, the attribute information includes (a) identification information that identifies the source document, and / or (b) information that describes a location from which the source document can be accessed.
[0219] Embodiment 45: A method performed by a computer of Embodiment 44, wherein the step of providing language output and two or more resulting images further includes the step of providing attribute information to a user computing device for display within the interface of the user computing device by the computing system.
[0220] Embodiment 46: A method performed by a computer of Embodiment 44, wherein the step of providing language output and two or more resulting images to a user computing device for display within the interface of the user computing device includes the step of providing interface data to the user computing device, the interface data including instructions for generating (a) an interface element including language output and (b) two or more selectable attribute elements associated with each of the two or more resulting images, each attribute element including a thumbnail of the associated resulting image and attribute information about one or more source documents containing the associated resulting image.
[0221] Embodiment 47: A method performed by a computer according to Embodiment 46, comprising the step of receiving information describing an instruction by a user of a user computing device that a first result image among two or more result images does not visually resemble a query image, the step of receiving data from the user computing device indicating a selection by the user of the user computing device of a first selectable attribute element among two or more selectable attribute elements, wherein the first selectable attribute element is associated with a first result image among two or more result images.
[0222] Embodiment 48: A method performed by a computer of Embodiment 47, wherein the step of removing text content associated with a source document containing a first result image from a set of text inputs further includes the step of removing information associated with the source document containing the first result image from attribute information to obtain refined attribute information, and the step of providing the refined language output to a user computing device further includes the step of providing the refined attribute information to a user computing device.
[0223] Embodiment 49: A method implemented by a computer of Embodiment 46, wherein the language output further includes predictive information that predicts which portion of the language output is most relevant to the prompt, and the interface data further includes instructions for generating emphasis elements that highlight the portion of the language output.
[0224] Embodiment 50: A method implemented by a computer, - A user computing device equipped with one or more processors acquires a query image, - The user computing device obtains text data that describes the prompt, - The user computing device provides a query image and text data describing the prompt to a computing system associated with the visual search service. - Steps include: - In response to providing a query image and a prompt, the user computing device receives from the computing system (a) two or more result images and (b) language output from a machine learning language model, wherein the language output is generated based on the prompt and text content from a source document containing the two or more result images; - Within the interface of the application executed by the user computing device, ○(a) Interface elements including language output, and ○(b) A method comprising the step of displaying two or more selectable attribute elements, each associated with two or more result images, wherein each selectable attribute element includes a thumbnail of the associated result image and attribute information identifying the source document containing the associated result image.
[0225] Embodiment 51: A computer-based method of Embodiment 50, wherein each selectable attribute element includes a first selectable portion and a second selectable portion.
[0226] Embodiment 52: A method performed by a computer of Embodiment 51, further comprising the step of receiving input from a user via an input device associated with the user computing device, for selecting a first selectable portion of a first selectable attribute element among two or more selectable attribute elements.
[0227] Embodiment 53: A computer-based method of Embodiment 52, further comprising the steps of: providing a computing system with information indicating a selection of a first selectable attribute element by a user computing device in response to receiving input to a first selectable portion of a first selectable attribute element; and receiving a refined language output from the computing system by the user computing device in response to having provided the information, wherein the refined language output is generated based on text content from a source document including a prompt and two or more result images other than a first result image associated with the first selectable attribute element.
[0228] Embodiment 54: A method implemented by a computer of Embodiment 53, further comprising the steps of a user computing device displaying, within the interface of an application executed by the user computing device, (a) an interface element including refined language output, and (b) one or more selectable attribute elements, wherein the one or more selectable attribute elements each include two or more selectable attribute elements other than a first selectable attribute element.
[0229] Embodiment 55: A method performed by a computer of Embodiment 52, further comprising the steps of: receiving input from a user via an input device associated with the user computing device, in which the user computing device selects a second selectable portion of a first selectable attribute element among two or more selectable attribute elements; and, in response to receiving the input of selecting a second selectable portion of a first selectable attribute element, causing the user computing device to display a source document identified by the attribute information contained in the first selectable attribute element.
[0230] Embodiment 56: Each of the source documents is, - One or more web pages of a website, - article, - Newspaper, - Book, or - A method carried out by a computer according to Embodiment 50, including a transcript.
[0231] Embodiment 57: A method performed by a computer of Embodiment 50, wherein the step of obtaining text data describing a prompt includes the steps of: obtaining spoken utterances from a user via an audio capture device associated with the user computing device by the user computing device; and determining text data describing a prompt by the user computing device, at least in part, based on the spoken utterances.
[0232] Embodiment 58: A method performed by a computer of Embodiment 50, wherein the step of acquiring a query image includes the steps of: the user computing device acquiring an input indicating a request to acquire an image using an image acquisition device associated with the user computing device; and, in response to acquiring the input, the user computing device acquiring the query image using an image acquisition device associated with the user computing device.
[0233] Embodiment 59: A user computing device, - One or more processors, - comprising one or more non-temporary computer-readable media that collectively store a first set of instructions, wherein when the instructions are executed by one or more processors, they cause a user computing device to perform an action, and the action is ○ Obtain the query image, ○ Obtain text data that describes the prompt, ○Providing query images and text data describing prompts to the computing system associated with the visual search service, ○ In response to providing a query image and a prompt, the computing system receives (a) two or more result images and (b) language output from a machine learning language model, wherein the language output is generated based on the prompt and text content from a source document containing the two or more result images. ○Within the interface of an application executed by a user computing device, • Interface elements including language output, and Displaying two or more selectable attribute elements, each associated with two or more result images, wherein each selectable attribute element includes a thumbnail of the associated result image and attribute information identifying the source document containing the associated result image. ○ Receiving input from the user via an input device associated with the user computing device, selecting the first selectable attribute element from two or more selectable attribute elements, ○In response to receiving input, provide the computing system with information indicating the selection of a first selectable attribute element, A user computing device that, in response to providing information, receives refined language output from a computing system, the refined language output being generated based on a prompt and text content from a source document, which includes two or more result images other than a first result image associated with a first selectable attribute element.
[0234] Embodiment 60: One or more non-temporary computer-readable media for storing a first set of instructions together, wherein when the instructions are executed by one or more processors of a user computing device, the user computing device performs an action, and the action is - Obtaining query images, - Obtain text data describing the prompt, - Provide the query image and text data describing the prompt to the computing system associated with the visual search service, - In response to providing a query image and a prompt, the system receives from a computing system two or more result images and language output from a machine learning language model, wherein the language output is generated based on the prompt and text content from a source document containing the two or more result images. - Within the interface of an application run by a user computing device, ○ Interface elements including language output, and ○One or more non-temporary computer-readable media, which include displaying two or more selectable attribute elements, each associated with two or more result images, wherein each selectable attribute element includes a thumbnail of the associated result image and attribute information identifying the source document containing the associated result image. [Explanation of Symbols]
[0235] 100 Visual Search Systems 102 User Computing Devices 104 Input Devices 106 Communication Module 108 Visual Search Computing System 110 Network 112 Query Images 114 Prompt 116 Visual Search Request 118 Visual Search Module 120 Text Contents 122 Result Images 124 Interface Data 202 Visual Search Computing System 203 User computing device 204 Visual search module 206 Query image 208 Prompt 210 Image evaluation module 212 Result image 214 Machine learning-based visual search model, visual search model 215 Image search space 216 Document indexing information 218 Document content selection module 220 Document 222 Text content 224 First text unit 226 Text judgment module 228 Second text unit 230 Machine learning-based language model 234 Interface data 236 Attribute information 232 Interface data generation module 400 Interface 402 Visual search request 404 Main interface element 405 Attribute element 406 Result image element 500A Interface 500B Interface 502 Visual search request 504 Main interface element 505 Attribute element 506 Emphasis element 508 Second interface element 510 Input 514 Information request element 602 Visual search computing system 604 Visual search module 606 Query image 608 Prompt 610 Image evaluation module 612 Result Images 614 Text Units 616 Document Content Selection Module 618 Document Indexing Information 620 Text Recognition Module 621 Large-scale language models 622 Derived text units 624 User Computing Devices 626 Attribute information 628 Result Image Selection Information 630 Second derived text unit 632 Second attribute information 700A Interface 700B Interface 702 Interface element 704 Selectable attribute elements 706 inputs 1100 Computing Systems, Systems 1180 Network 1102 User Computing System 1104 User Computing Devices 1112 processors 1114 memory 1116 Data 1118 command 1120 Machine learning models, models 1122 User Input Components 1126 Sensor 1130 Server Computing System 1132 processors 1134 memory 1136 data 1138 command 1140 Machine learning models, models 1142 Search Engine 1144 User Interface 1150 Third-Party Computing Systems 1152 processors 1154 Memory 1156 Data 1158 Command 1250 Computing System 1252 Computing Device, User Computing Device 1260 Sensor Processing System 1262 Context Judgment Block 1264 Image Preprocessing Block 1266 Detection Model 1268 Segmentation Model 1270 Classification Model 1272 Embedding Model 1274 Search Engine 1276 Multimodal Processing Block 1280 Output Judgment System 1282 Search Result Interface 1284 Machine Learning Model Output Interface 1286 Augmented Reality Experience and / or Virtual Reality Experience 1288 Action Prompt 1290 Generative Model 1292 Data Augmentation Block 1294 Data Storage Block
Claims
1. A method carried out by a computing system comprising one or more processor devices, The computing system performs the steps of searching for the result image based on the similarity between the query image and the result image, A step of obtaining a first text unit using the computing system, wherein the first text unit includes at least a portion of the text content of the source document including the resulting image. The computing system processes the first text unit and the prompt associated with the query image using a machine learning language model to obtain a language output including a second text unit, wherein the second text unit is text content describing a response to the prompt. (a) at least a portion of the first text unit, or (b) A step comprising one or more of the texts derived from the first text unit, The computing system provides the second text unit and the resulting image for display within the interface. A method implemented by a computing system that includes [specific components / systems].
2. The search step for the resulting image is as follows: The computing system performs the steps of processing the query image with a machine learning-type visual search model to obtain an intermediate representation of the query image, The computing system performs the steps of searching for the result image based on the degree of similarity between the intermediate representation of the query image and the intermediate representation of the result image. A method carried out by the computing system according to claim 1, including the method described in claim 1.
3. The step of processing the query image with the machine learning visual search model includes the step of processing the query image with a machine learning embedding model by the computing system to obtain a query image embedding for the query image, The method, performed by the computing system according to claim 2, wherein the step of searching for the result image includes the step of searching for the result image based on the distance between the query image embedding and the result image embedding within the embedding space.
4. The method, as performed by the computing system according to claim 1, comprises the step of obtaining the query image from a user computing device by the computing system prior to processing the query image.
5. The method implemented by the computing system according to claim 4, wherein the interface includes a user interface for an application executed by the user computing device.
6. The method performed by the computing system according to claim 4, wherein the step of obtaining the query image includes the step of the computing system obtaining the query image and the prompt associated with the query image from the user computing device.
7. The search step for the resulting image is as follows: The steps include providing the user computing device with the result image for display within the interface using the computing system, In response to providing the aforementioned result image, the user computing device receives the prompt associated with the query image. A method carried out by the computing system according to claim 4, further comprising:
8. A method carried out by the computing system according to claim 1, wherein the second text unit includes a subset of the first text unit.
9. A method carried out by the computing system according to claim 1, wherein the second text unit includes text derived from the first text unit, the text derived from the first text unit describing a summary of the first text unit.
10. The aforementioned source document is One or more web pages of a website, article, newspaper, Books, or A method carried out by the computing system according to claim 1, including a transcript.
11. The method, performed by the computing system according to claim 1, wherein the step of providing the second text unit and the resulting image further includes the step of providing attribute information for display by the computing system within the interface, which (a) identifies the source document and / or (b) indicates a location accessible from the source document.
12. The method, carried out by the computing system according to claim 11, wherein the source document includes a web page, and the attribute information includes the address of the web page.
13. The method, carried out by the computing system according to claim 11, wherein the source document includes a journal, and the attribute information includes a reference indicating the location of the resulting image within the journal.
14. Prior to processing the first text unit and the prompt associated with the query image, the method comprises the step of the computing system generating the prompt associated with the query image based at least in part on the query image, as performed by the computing system according to claim 1.
15. The step of generating the prompt associated with the query image is: The computing system performs the steps of processing the query image with a machine learning model to generate a semantic output that describes the query image, The computing system provides the steps of generating the prompt based at least partially on the semantic output. A method carried out by the computing system according to claim 14, including the following:
16. A computing system, One or more processors, One or more non-temporary computer-readable storage media that store a first set of instructions together, and The first set of instructions, when executed by the one or more processors, causes the computing system to perform an operation, and the operation is Retrieving query images and associated prompts from the user's computing device, The process involves processing the query image with a machine learning-based embedding model to obtain the query image embedding, Based on the similarity between the embedded query image and the embedded result image, the result image is searched, Identifying the source document for the aforementioned result image, wherein the source document includes the aforementioned result image and the text content associated with the aforementioned result image. Determining a first text unit from the source document that includes at least a portion of the text content associated with the resulting image, The first text unit and the prompt are processed by a machine learning language model to obtain a language output including a second text unit, wherein the second text unit is text content describing a response to the prompt. (a) at least a portion of the first text unit, or (b) Text derived from the first text unit, This includes one or more of the following: The second text unit and the resulting image are provided for display within the interface of an application executed by the user computing device. A computing system that includes this.
17. The computing system according to claim 16, further comprising providing the second text unit and the resulting image for display within the interface of the application executed by the user computing device, wherein providing the second text unit and the resulting image further includes providing attribute information identifying the source document for display within the interface of the application executed by the user computing device.
18. One or more non-temporary computer-readable storage media storing a first set of instructions, wherein the first set of instructions, when executed by one or more processors of a computing system, causes the computing system to perform an operation, and the operation is Searching for the multiple result images based on the similarity between the intermediate representation of the query image and each of the multiple intermediate representations associated with each of the multiple result images, Identifying multiple source documents, wherein each of the multiple source documents includes one of the multiple result images and text content associated with that result image. Determining a plurality of first text units for each of the plurality of result images, wherein each first text unit includes at least a portion of the text content associated with the corresponding result image from one or more source documents containing the corresponding result image among the plurality of result images, Processing a set of text inputs with a machine learning language model to obtain a language output that includes a second text unit describing a response to a prompt associated with the query image, wherein the set of text inputs is: (a) Two or more first text units associated with two or more of the result images among the plurality of result images, and (b) The prompt associated with the query image, This includes, The second text unit and the two or more resulting images are provided to the user computing device for display within the interface of the user computing device. One or more non-temporary computer-readable storage media, including [the specified element].