Machine-learned models for multimodal image search and retrieval

The machine-learned query refinement model addresses inefficiencies in traditional visual search by incorporating textual refinements, enhancing search efficiency and reducing resource waste.

JP2025526787APending Publication Date: 2025-08-15GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025507648
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-12
Filing Date
2022-11-04
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional visual search models are unable to incorporate user-provided textual query refinements, leading to inefficient and resource-intensive re-capturing of query targets due to incorrect determination of user intent.

Method used

A machine-learned query refinement model processes query image embeddings and textual query refinements to generate refined image embeddings, utilizing a loss function to modify model parameters, enabling efficient retrieval of images based on textual refinements.

Benefits of technology

Facilitates faster and more efficient visual search by allowing users to refine their queries with text data, reducing unnecessary resource usage and improving search efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025526787000001_ABST
    Figure 2025526787000001_ABST
Patent Text Reader

Abstract

The disclosed systems and methods are directed to a computer-implemented method for machine-learned multimodal search refinement. The method includes obtaining a query image embedding for a query image and a textual query refinement associated with the query image. The method includes processing the query image embedding and the textual query refinement with a machine-learned query refinement model to obtain a refined query image embedding that incorporates the textual query refinement. The method includes evaluating a loss function that assesses a distance between the refined query image embedding and an embedding for a ground truth image in an image embedding space. The method includes modifying values of parameters of the machine-learned query refinement model based on the loss function.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Aspects of the present disclosure relate to image searching and retrieval, and more particularly, to machine-learned models that enable retrieval of images from a database in a faster or more efficient manner. [Background technology]

[0002] Recently, visual search capabilities have been offered as a feature across a wide variety of applications (e.g., virtual assistant applications, camera applications, etc.). Traditionally, to perform a visual search, a user first provides an image to a search service. These search services typically use machine learning techniques to process the image and identify visually similar or semantically similar images and / or information related to entities depicted in the image. However, because these traditional models are trained only to process image data, they are unable to incorporate user-provided textual query refinements. Summary of the Invention

[0003] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the description that follows, or may be learned from the description, or may be learned by practice of the embodiments.

[0004] One exemplary aspect of the present disclosure is directed to a computer-implemented method for machine-learned multimodal search refinement. The method includes obtaining, by a computing system including one or more computing devices, a query image embedding for a query image and a textual query refinement associated with the query image. The method includes processing, by the computing system, the query image embedding and the textual query refinement with a machine-learned query refinement model to obtain a refined query image embedding that incorporates the textual query refinement. The method includes evaluating, by the computing system, a loss function that evaluates a distance between the refined query image embedding and an embedding for a ground truth image in an image embedding space. The method includes modifying, by the computing system, one or more values of one or more parameters of the machine-learned query refinement model based at least in part on the loss function.

[0005] Another exemplary aspect of the present disclosure is directed to a computing system for machine-learned multimodal search refinement. The computing system includes one or more processors. The computing system includes a machine-learned query refinement model trained to refine a query image with textual query refinement. The computing system includes one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations include obtaining image embeddings for a query image provided by a user of a visual search application. The operations include obtaining textual query refinements for the query image from the user of the visual search application, the textual query refinements being responsive to providing one or more initial result images for the query image to the user of the visual search application. The operations include processing the image embeddings and the textual query refinements for the query image with the machine-learned query refinement model to obtain refined image embeddings that incorporate the textual query refinement. The operations include determining one or more refined result images based at least in part on the refined image embeddings that incorporate the textual query refinement.

[0006] Another example aspect of the present disclosure is directed to one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause a computing system to perform operations. The operations include obtaining image embeddings for a query image provided by a user of a visual search application. The operations include obtaining a text query refinement for the query image from the user of the visual search application, the text query refinement being responsive to providing one or more initial result images for the query image to the user of the visual search application. The operations include processing the image embeddings and the text query refinement for the query image with a machine-learned query refinement model to obtain refined image embeddings that incorporate the text query refinement. The operations include determining one or more refined result images based at least in part on the refined image embeddings that incorporate the text query refinement.

[0007] Other aspects of the present disclosure are directed to various systems, apparatus, non-transitory computer-readable media, user interfaces, and electronic devices.

[0008] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following detailed description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the detailed description, serve to explain associated principles.

[0009] Detailed descriptions of embodiments directed to those skilled in the art are set forth herein with reference to the accompanying drawings. [Brief explanation of the drawings]

[0010] [Figure 1A]1 illustrates a block diagram of an exemplary computing system for performing machine-learned query refinement, according to an exemplary embodiment of the present disclosure. [Figure 1B] 1 illustrates a block diagram of an exemplary computing device that performs machine-learned query refinement, according to an exemplary embodiment of the present disclosure. [Figure 1C] 1 illustrates a block diagram of an exemplary computing device for performing training of a machine-learned query refinement model, according to an exemplary embodiment of the present disclosure. [Figure 2] FIG. 1 illustrates a dataflow diagram for training a machine-learned query refinement model for visual search query refinement, according to some embodiments of the present disclosure. [Figure 3] FIG. 1 illustrates a dataflow diagram for refining a visual search query with a machine-learned query refinement model, according to some embodiments of the present disclosure. [Figure 4] FIG. 1 illustrates a dataflow diagram for processing inputs of a machine-learned query refinement model for visual search query refinement, according to some embodiments of the present disclosure. [Figure 5] FIG. 1 illustrates a communication flow diagram for providing visual search services with query refinement, according to some embodiments of the present disclosure. [Figure 6] FIG. 1 illustrates a flowchart diagram of an exemplary method for performing training of a machine-learned query refinement model, according to an exemplary embodiment of the present disclosure. [Figure 7] FIG. 1 illustrates a flowchart diagram of an exemplary method for performing query refinement with a machine-learned query refinement model, according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] Reference numbers repeated among the drawings are intended to identify like features in the various embodiments.

[0012] overview Aspects of the present disclosure are directed to the technical task of searching and retrieving images using multimodal search. More specifically, the present disclosure relates to a machine-learned model for text refinement of image queries to form multimodal queries. As an example, a query image embedding for a query image can be obtained using text query refinements associated with the query image. For example, the query image may depict a particular person, and the text query refinement may describe visual characteristics associated with the particular person (e.g., an item of clothing or a facial feature such as a beard) that differ from or are absent from the query image. The query image embedding and the text query refinement can be processed with a machine-learned query refinement model (e.g., a Transformer model, etc.) to obtain a refined query image embedding that incorporates the text query refinement. Following the foregoing example, the refined query image embedding can be an image embedding for a particular (or at least visually similar) person with the characteristics described in the text query refinement. A loss function may be evaluated that assesses the distance between the refined query image embedding and an embedding for a ground truth image in the image embedding space. One or more values of one or more parameters of the machine-learned query refinement model may be modified based at least in part on the loss function. In this way, the machine-learned query refinement model can be trained to refine image embeddings for the initial query image such that the refined image embeddings incorporate the text query refinement, thus enabling a fast and efficient mechanism for retrieving images and / or other information related to those images from a database.

[0013] Embodiments of the present disclosure provide several technical effects and advantages. As one example of a technical effect and advantage, the technical task of searching for and retrieving similar images may be performed in a faster and / or more efficient manner. For example, users of conventional visual search applications often need to re-capture images of query targets due to the visual search application's incorrect determination of the user's intent. As a result, this can lead to a frustrating user experience and unnecessary use of resources (e.g., power, computational cycles, memory, storage, bandwidth, etc.) to re-capture the query targets. However, embodiments of the present disclosure provide a machine-learned query refinement model that can be utilized by visual search applications to provide text refinement features, allowing users to quickly and efficiently refine their visual searches using text data, thus significantly improving search efficiency by eliminating unnecessary resource usage associated with re-capturing query targets. Additionally, as will be appreciated, embodiments of the present disclosure may facilitate visual-search-based retrieval of images that cannot otherwise be easily retrieved using only the query image.

[0014] Referring now to the figures, exemplary embodiments of the present disclosure will be described in more detail.

[0015] Exemplary Devices and Systems 1A illustrates a block diagram of an exemplary computing system 100 for performing machine-learned query refinement, according to an exemplary embodiment of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150, communicatively coupled via a network 180.

[0016] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0017] The user computing device 102 includes one or more processors 112 and memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple operatively connected processors. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0018] In some implementations, the user computing device 102 can store or include one or more machine-learned query refinement models 120. For example, the machine-learned query refinement models 120 can be or otherwise include various machine-learned models, such as neural networks (e.g., deep neural networks), or other types of machine-learned models, including nonlinear and / or linear models. The neural networks can include feed-forward neural networks, recurrent neural networks (e.g., long-short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Some exemplary machine-learned models can leverage attention mechanisms, such as self-attention. For example, some exemplary machine-learned models can include multi-head self-attention models (e.g., Transformer models). Exemplary machine-learned query refinement models 120 are described with reference to FIGS. 2-5.

[0019] In some implementations, one or more machine-learned query refinement models 120 may be received from a server computing system 130 over a network 180, may be stored in a user computing device memory 114, and may then be used or otherwise implemented by one or more processors 112. In some implementations, a user computing device 102 may implement multiple parallel instances of a single machine-learned query refinement model 120 (e.g., to perform parallel query refinement across multiple instances of the machine-learned query refinement model 120).

[0020] More specifically, the machine-learned query refinement model 120 can be trained and utilized to refine a query image provided for visual image search using textual query refinement. For example, the machine-learned query refinement model 120 can process a query image or a representation of the query image (e.g., a query image embedding) simultaneously with the textual query refinement or a representation of the query refinement (e.g., a token embedding for the textual query refinement). The machine-learned query refinement model 120 can then generate a refined query image embedding that incorporates the textual query refinement. For example, if the query image depicts a particular person and the textual query refinement describes an item of clothing, "hat," the refined image embedding may correspond to an image of that person (or a visually similar person) wearing the hat. This refined query image embedding can be utilized to retrieve images related to the image embedding within a certain distance of the refined image embedding in an image embedding space (e.g., an embedding space for an image search service). In this way, the machine-learned query refinement model 120 can be utilized to provide text refinement capabilities to visual search applications.

[0021] Additionally or alternatively, one or more machine-learned query refinement models 140 can be included in or otherwise stored and implemented by a server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine-learned query refinement models 140 can be implemented by the server computing system 140 as part of a web service (e.g., a query refinement service, a visual search service, etc.). Thus, one or more models 120 can be stored and implemented at the user computing device 102 and / or one or more models 140 can be stored and implemented at the server computing system 130.

[0022] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component may serve to implement a virtual keyboard. Other exemplary user input components include a microphone, a traditional keyboard, or other means by which a user can provide user input.

[0023] The server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple operatively connected processors. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 can store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to operate.

[0024] In some implementations, server computing system 130 includes or is otherwise implemented by one or more server computing devices. When server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0025] As described above, the server computing system 130 may store or otherwise include one or more machine-learned query refinement models 140. For example, the models 140 may be or otherwise include various machine-learned models. Exemplary machine-learned models include neural networks or other multi-layer nonlinear models. Exemplary neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some exemplary machine-learned models may utilize attention mechanisms such as self-attention. For example, some exemplary machine-learned models may include multi-head self-attention models (e.g., Transformer models). Exemplary models 140 are described with reference to FIGS. 2-5.

[0026] The user computing device 102 and / or the server computing system 130 can train the models 120 and / or 140 by interacting with a training computing system 150 that is communicatively coupled via a network 180. The training computing system 150 can be separate from the server computing system 130 or can be part of the server computing system 130.

[0027] The training computing system 150 includes one or more processors 152 and memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple operably connected processors. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is otherwise implemented by one or more server computing devices.

[0028] The training computing system 150 may include a model trainer 160 that trains the machine-learned models 120 and / or 140 stored on the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as, for example, backpropagation. For example, a loss function may be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent may be used to iteratively update the parameters over several training iterations.

[0029] In some implementations, performing backpropagation may include performing truncated backpropagation through time. Model trainer 160 may implement several generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the trained model.

[0030] In particular, model trainer 160 can train machine-learned query refinement model 120 and / or 140 based on a set of training data 162. Training data 162 can include various pairs of image data and associated text query refinements. For example, in some embodiments, training data 162 can include a corpus of image search data. The corpus of image search data can describe interactions with search results from users and can include search result images provided to users in response to a query and refined search result images provided to users in response to selection of query refinement elements provided to the user with the search result images. The query images, text query refinements, and ground truth images can be selected from the search result images, selectable query refinement elements, and refined search result images.

[0031] For example, a corpus of image search data may include multiple search result images provided to a user by an image search service in response to a text query (e.g., a text query for "car"). The corpus of image search data may indicate search result images with which the user most frequently interacts (e.g., images depicting black cars, etc.). The corpus of image search data may also include refined search result images provided to a user in response to selection of a query refinement element provided to the user along with the search result images. For example, an image search service may provide, in addition to search result images, several query refinement user interface elements (e.g., elements such as "truck," "red," "blue," "fast," and "van" for a text query such as "car") that are selectable to refine a user-provided text query. If a query refinement element is selected by the user, refined search result images may be provided to the selecting user in response to the text query and the selected query refinement element. For example, a user may provide an initial text query of "car" and then select a query refinement element for "blue." The refined search result images may each depict a blue car.

[0032] The corpus of image search data can indicate the refined search result images that users most frequently interact with after selecting an associated query refinement element. Query images, text query refinements, and ground truth images can be selected from the search result images, selectable query refinement elements, and refined search result images for inclusion in training data 162.

[0033] For example, multiple users may each provide a text query for "cars" to an image search application, and in response, the image search application may provide multiple search result images to the users. The corpus of image search data may indicate a first search result image of the multiple search result images as one with which the users are most interacting. Next, multiple users may each select the same query refinement element of "blue." The image search application may provide multiple refined search result images depicting blue cars, and the corpus of image search data may indicate a first refined search result image of the multiple refined search result images as one with which the multiple users are most interacting. The first search result image may be selected as a query image, the text content of the query refinement element (e.g., "blue") may be selected as a text query refinement, and the first refined search result image may be selected as a ground truth image. The query image, the text query refinement, and the ground truth image may collectively be included in training data 162 as training examples for training machine-learned query refinement model 120 / 140 by model trainer 160. In this manner, a corpus of image search data can be utilized to generate multiple training examples for inclusion in training data 162 for training machine-learned query refinement model 120.

[0034] In some implementations, if the user provides consent, training examples may be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 against user-specific data received from the user computing device 102. In some cases, this process may be referred to as personalizing the model.

[0035] Model trainer 160 includes computer logic utilized to provide desired functionality. Model trainer 160 can be implemented in hardware, firmware, and / or software controlling a general-purpose processor. For example, in some embodiments, model trainer 160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other embodiments, model trainer 160 includes one or more sets of computer-executable instructions stored on a tangible computer-readable storage medium, such as RAM, a hard disk, or optical or magnetic media.

[0036] Network 180 can be any type of communications network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communications over network 180 can occur over any type of wired and / or wireless connection, using a wide variety of communications protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or security schemes (e.g., VPN, Secure HTTP, SSL).

[0037] In some implementations, the input to the machine-learned model of the present disclosure can be image data. The machine-learned model can process the image data to generate an output. As an example, the machine-learned model can process the image data to generate an image recognition output (e.g., a recognition of the image data, a latent embedding of the image data, an encoded representation of the image data, a hash of the image data, etc.). As another example, the machine-learned model can process the image data to generate an image segmentation output. As another example, the machine-learned model can process the image data to generate an image classification output. As another example, the machine-learned model can process the image data to generate an image data modification output (e.g., a modification of the image data, etc.). As another example, the machine-learned model can process the image data to generate an encoded image data output (e.g., an encoded and / or compressed representation of the image data, etc.). As another example, the machine-learned model can process the image data to generate an upscaled image data output. As another example, the machine-learned model can process the image data to generate a predicted output.

[0038] In some implementations, the input to the machine-learned models of the present disclosure may be text or natural language data. The machine-learned model may process the text or natural language data to generate an output. As an example, the machine-learned model may process the natural language data to generate a language-encoded output. As another example, the machine-learned model may process the text or natural language data to generate a latent text embedding output. As another example, the machine-learned model may process the text or natural language data to generate a translation output. As another example, the machine-learned model may process the text or natural language data to generate a classification output. As another example, the machine-learned model may process the text or natural language data to generate a text segmentation output. As another example, the machine-learned model may process the text or natural language data to generate a semantic intent output. As another example, the machine-learned model may process the text or natural language data to generate an upscaled text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language). As another example, the machine-learned model may process the text or natural language data to generate a predicted output.

[0039] In some implementations, the input to a machine-learned model of the present disclosure can be audio data. The machine-learned model can process the audio data to generate an output. As an example, the machine-learned model can process the audio data to generate a speech recognition output. As another example, the machine-learned model can process the audio data to generate a speech translation output. As another example, the machine-learned model can process the audio data to generate a latent embedding output. As another example, the machine-learned model can process the audio data to generate an encoded audio output (e.g., an encoded and / or compressed representation of the audio data, etc.). As another example, the machine-learned model can process the audio data to generate an upscaled audio output (e.g., audio data of higher quality than the input audio data, etc.). As another example, the machine-learned model can process the audio data to generate a text representation output (e.g., a text representation of the input audio data, etc.). As another example, the machine-learned model can process the audio data to generate a predicted output.

[0040] In some implementations, input to the machine-learned models of the present disclosure can be latent-coded data (e.g., a latent space representation of the input, an image embedding for an image, an embedding (e.g., a token embedding) for textual content, etc.). The machine-learned model can process the latent-coded data to generate an output. As an example, the machine-learned model can process the latent-coded data to generate a recognition output. As another example, the machine-learned model can process the latent-coded data to generate a reconstruction output. As another example, the machine-learned model can process the latent-coded data to generate a search output. As another example, the machine-learned model can process the latent-coded data to generate a reclustered output. As another example, the machine-learned model can process the latent-coded data to generate a predicted output.

[0041] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data of one or more images and the task is an image processing task. For example, the image processing task can be image classification, and the output is a set of scores, each score corresponding to a different object class and representing the likelihood that one or more images depict an object belonging to that object class. The image processing task can be object detection, and the image processing output identifies one or more regions in one or more images and, for each region, the likelihood that the region depicts an object of interest. As another example, the image processing task can be image segmentation, and the image processing output specifies, for each pixel in one or more images, a respective likelihood for each category in a predetermined set of categories. For example, the category set can be foreground and background. As another example, the category set can be object classes. As another example, the image processing task can be depth estimation, and the image processing output specifies, for each pixel in one or more images, a respective depth value. As another example, the image processing task can be motion estimation, where the network input includes multiple images and the image processing output specifies, for each pixel of one of the input images, the motion of the scene depicted in pixels between images in the network input.

[0042] 1A illustrates one exemplary computing system that can be used to implement the present disclosure. Other computing systems can be used as well. For example, in some implementations, a user computing device 102 can include a model trainer 160 and a training dataset 162. In such implementations, the model 120 can be both trained and used locally on the user computing device 102. In some such implementations, the user computing device 102 can implement the model trainer 160, which personalizes the model 120 based on user-specific data.

[0043] 1B illustrates a block diagram of an exemplary computing device 10 that performs machine-learned query refinement according to an exemplary embodiment of the present disclosure. The computing device 10 can be a user computing device or a server computing device.

[0044] Computing device 10 includes several applications (e.g., applications 1-N). Each application includes its own machine learning library and machine-learned models. For example, each application may include a machine-learned model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.

[0045] 1B , each application can communicate with several other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0046] 1C illustrates a block diagram of an exemplary computing device 50 for performing training of a machine-learned query refinement model, according to an exemplary embodiment of the present disclosure. The computing device 50 may be a user computing device or a server computing device.

[0047] Computing device 50 includes several applications (e.g., applications 1-N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API across all applications).

[0048] The central intelligence layer includes several machine-learned models. For example, as shown in FIG. 1C , each machine-learned model can be provided for each application and managed by the central intelligence layer. In other embodiments, two or more applications can share a single machine-learned model. For example, in some embodiments, the central intelligence layer can provide a single model for all applications. In some embodiments, the central intelligence layer is included within or otherwise implemented by the operating system of computing device 50.

[0049] The central intelligence layer can communicate with a central device data layer, which can be a centralized repository of data for computing device 50. As shown in FIG. 1C , the central device data layer can communicate with several other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0050] Example Model Configuration FIG. 2 illustrates a dataflow diagram for training a machine-learned query refinement model for refining visual search queries, according to some embodiments of the present disclosure. Specifically, query image embeddings 202 for a query image and textual query refinements 204 associated with the query image may be obtained for processing by a machine-learned query refinement model 206. In some embodiments, the query image embeddings 202 may be image embeddings for a visual query image depicting one or more entities (e.g., objects, scenes, people, etc.). Each of the one or more entities may have one or more characteristics. For example, the entity may be an object such as a car. The characteristics may be make, model, color, location, pose (e.g., interior or exterior of the car), vehicle condition, etc. Specifically, the entity depicted in the query image may have a first characteristic. For example, the entity may be pants, and the first characteristic may be the color blue (e.g., the query image depicts blue pants). The textual query refinement may describe a second characteristic that is different from the second characteristic (e.g., the color red). The ground truth image may depict an entity (e.g., pants) having a second characteristic (e.g., pants that are red in color).

[0051] In some implementations, the query image embedding 202 may be an image embedding for a portion of the visual query image. For example, a user may provide input selecting a portion of the query image. The query image embedding 202 may be generated for the portion of the query image selected by the user. In other examples, the computing system that generates the query image embedding 202 may automatically select the portion of the query image for which to generate the query image embedding 202 based on the content of the selected portion of the image and the non-selected portion of the image. For example, the computing system may determine that the non-selected portion of the image does not include any objects of interest, while the selected portion of the image includes some objects of interest.

[0052] In some embodiments, the text query refinement may describe one of the characteristics of the entity. For example, the entity depicted by the query image may be an item of clothing. The characteristics of the clothing described by the text query refinement may be the color, brand, size, etc. of the clothing.

[0053] Note that the query image embedding 202 is an image embedding in the embedding space, but the query image embedding 202 may refer to any encoding or latent representation of the query image 202.

[0054] In some embodiments, the query image embeddings 202 and the text query refinements 204 can be obtained from a corpus of image search data. The corpus of image search data can describe user interactions with search results and can include search result images provided to the user in response to a query and refined search result images provided to the user in response to selection of a query refinement element provided to the user with the search result images. The query images for the query image embeddings 202, the text query refinements 204, and the ground truth images for the ground truth image embeddings 210 can be selected from the search result images, selectable query refinement elements, and refined search result images.

[0055] For example, multiple users may each provide a text query of "cars" to an image search application, and in response, the image search application may provide multiple search result images to the users. A corpus of image search data may indicate a first search result image of the multiple search result images as one with which the users are most interacting. The multiple users may then each select the same query refinement element of "blue." The image search application may provide multiple refined search result images depicting blue cars, and the corpus of image search data may indicate a first refined search result image of the multiple refined search result images as one with which the multiple users are most interacting. The first search result image may be selected as a query image for query image embedding 202, the text content of the query refinement element (e.g., "blue") may be selected as a text query refinement 204, and the first refined search result image may be selected as a ground truth image for ground truth image embedding 210.

[0056] The machine-learned query refinement model 206 can process the textual query refinement 204 and the query image embedding 202 to obtain refined image embeddings 208. The refined image embeddings 208 are image embeddings that incorporate the textual query refinement 204. For example, the query image embedding 202 may depict a blue dress. The textual query refinement 204 may include the word "red." The refined image embeddings 208 can correspond to image embeddings for images of red dresses. In other words, the refined image embeddings can be alternative representations of images of red dresses. As will be appreciated from this disclosure, the refined image embeddings can be used to obtain one or more images that are similar to the image represented by the refined image embedding based on low-level features of images in a corpus of images.

[0057] The loss function evaluator 212 can evaluate a loss function that evaluates the difference between the refined image embedding 208 and the ground truth embedding 210. Specifically, the loss function evaluates the distance between the refined image embedding 208 and the ground truth image embedding 210 in the image embedding space 214.

[0058] Based at least in part on the loss function, the modification determiner 216 may modify one or more values of one or more parameters of the machine-learned query refinement model 206 via parameter value modification 218.

[0059] 3 illustrates a dataflow diagram 300 for refining a visual search query with a machine-learned query refinement model, according to some embodiments of the present disclosure. Specifically, image embeddings may be obtained for a query image provided by a user of a visual search application. For example, the visual search application may be an application feature provided by an operating system. In other examples, the visual search application may be an application communicatively coupled to a second application accessed by the user. In yet other examples, the visual search application may include a machine-learned query refinement model 306 or may communicate with a service that provides the machine-learned query refinement model 306.

[0060] A text query refinement 304 for the query image may be obtained from a user of the visual search application. The text query refinement may be responsive to providing one or more initial result images for the query image to the user of the visual search application. The text query refinement 304 and the image embeddings 302 may be processed by a machine-learned query refinement model 306 to obtain refined image embeddings 308. The refined image embeddings may incorporate the text query refinement 304 as discussed with respect to the refined image embeddings 208 of FIG. 2 .

[0061] The result image determiner 310 may determine one or more refined result images 314 based at least in part on the refined image embedding 308. In some embodiments, determining the one or more refined result images includes determining one or more image embeddings within a threshold distance of the refined image embedding 308 in the image embedding space 312, and selecting one or more refined result images 314 corresponding to the one or more image embeddings within the threshold distance, respectively.

[0062] FIG. 4 illustrates a dataflow diagram 400 for processing inputs of a machine-learned query refinement model for refining a visual search query, according to some embodiments of the present disclosure. Specifically, a query image embedding 406 (e.g., query image embedding 202 of FIG. 2 , image embedding 302 of FIG. 3 , etc.) can be determined from a query image. In some embodiments, a machine-learned model, such as a machine-learned image encoding model 404, can process the query image 402 to generate the query image embedding 406. For example, the machine-learned image encoding model can be a model trained to generate image embeddings for an image embedding space. The image embedding space can be an image embedding space utilized for visual search applications. The query image embedding can be processed by the machine-learned query refinement model 206, as described with respect to FIG. 2 .

[0063] In some embodiments, the text query refinement 408 can be processed using a machine-learned model, such as a machine-learned text encoding model 410, to obtain a latent representation or encoding of the text query refinement 408. In some embodiments, the machine-learned text encoding model 410 can be trained to process the text query refinement to generate a plurality of token embeddings 412. Alternatively, in some embodiments, the machine-learned text encoding model 410 can process the text query refinement 408 to generate some other type of latent representation 410 of the text query refinement.

[0064] In some embodiments, the machine-learned image encoding model may be a sub-model or part of the machine-learned query refinement model 206. For example, while processing the query image 402 and the text query refinement 408 as described with respect to FIG. 2, the part of the machine-learned query refinement model 206 that includes the machine-learned image encoding model 404 may process the query image 402 to obtain the query image embedding 406, and then process the query image embedding 406 together with the token embedding 412 to obtain the refined image embedding as described with respect to FIG.

[0065] Similarly, in some embodiments, the machine-learned text encoding model may be a sub-model or part of the machine-learned query refinement model 206. For example, as described with respect to Figure 2, while processing the query image 402 and the text query refinement 408, the part of the machine-learned query refinement model 206 that includes the machine-learned text encoding model 410 may process the text query refinement 408 to obtain token embeddings 412, and then process the query image embeddings 406 together with the token embeddings 412 to obtain refined image embeddings.

[0066] FIG. 5 illustrates a communication flow diagram for providing a visual search service with query refinement according to some embodiments of the present disclosure. Specifically, in step 504, a computing system 500 (e.g., computing system 100 of FIG. 1 , etc.) may receive a query image from a user of a user device 502 (e.g., user device 102 of FIG. 1 , etc.). Note that the communicating entities 500 and 502 are depicted solely to more easily illustrate the flow of information and may represent any other type of computing device or system. For example, the computing system 500 and the user device 502 may be the user device 102 of FIG. 1 , and the information flow illustrated in FIG. 5 may represent or otherwise involve inter-process communication on the same device. Alternatively, the computing system 100 may be the server computing system 130 of FIG. 1 and may provide a service that facilitates a visual search application running on the user device 102 of FIG. 1 . Thus, it should be broadly understood that embodiments of the present disclosure may be implemented across any configuration of computing devices of the present disclosure.

[0067] Once the query images are received, in step 506, the computing system 500 may determine an image embedding for each query image, as described above with respect to Figure 4. In step 508, the image embeddings may be utilized to determine an initial result image. For example, the computing system 500 may determine the image embedding of the initial result image that is closest to the image embedding of the query image in image embedding space.

[0068] At step 510, the computing system 500 may provide the initial result images to the user device 502. In some embodiments, the computing system may provide the initial result images within an interface of a visual search application (e.g., executed by the user device 502, etc.).

[0069] At step 512, in response to providing the initial result image, computing system 500 may obtain a text query refinement in response to providing the initial result image at step 510. For example, the initial result image may be provided to a user of the visual search application within an interface of the visual search application. The visual search application may provide a display to a user of the application that provides text query refinement via a text input field presented within the interface. The user may provide the text query refinement via the text input field, and the text query refinement may be provided to computing system 500.

[0070] At step 514, the computing system can process the image embeddings and the text query refinement for the query image using the machine-learned query refinement model to obtain refined image embeddings that incorporate the text query refinement, as described above with respect to Figures 3 and 4.

[0071] In some embodiments, the computing system can also process information related to the image corresponding to the image embedding. For example, the image corresponding to the image embedding may be hosted on a website that includes a description of the image. The information related to the image may include a description or may be otherwise generated based on the description. In another example, the image corresponding to the image embedding may be processed with a semantic image processing model operable to generate a semantic output that describes the image. The information may include the semantic output. In yet another example, the image corresponding to the image embedding may be hosted on a website or application that allows users to post text content (e.g., tags, comments, etc.) associated with the image. This information may include or otherwise describe the user-posted text content. The machine-learned query refinement model can process text query refinement on the information, the image embedding, and the query image to obtain refined image embeddings.

[0072] At step 516, the computing system 500 may determine refined result images based on the refined image embeddings, as described above with respect to Figure 3. At step 518, the computing system 500 may provide the refined result images to the user device 502. In some embodiments, the refined result images may be provided to the user device 502 for display within the interface of a visual search application.

[0073] In some embodiments, at step 520, computing system 500 may obtain a second text query refinement in response to providing the refined result image. For example, the refined result image may be provided to a user of the visual search application within an interface of the visual search application. The visual search application may provide a user of the application with a display that provides the second text query refinement via a text input field presented within the interface. The user may provide the second text query refinement via the text input field and provide the second text query refinement to computing system 500.

[0074] In some embodiments, at step 522, the computing system 500 may process the second textual query refinement and the image embeddings of the query image with a machine-learned query refinement model to obtain second refined image embeddings that incorporate the second textual query refinement. Alternatively, in some embodiments, at step 22, the computing system 500 may process the refined image embeddings and the second textual query refinement with a machine-learned query refinement model to obtain second refined image embeddings that incorporate the textual query refinement and the second textual query refinement.

[0075] Exemplary Methods 6 illustrates a flowchart of an example method 600 for training a machine-learned query refinement model, according to an example embodiment of the present disclosure. While FIG. 6 depicts steps performed in a particular order for purposes of illustration and explanation, the method of the present disclosure is not limited to the particularly depicted order or arrangement. Various steps of method 600 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0076] At 602, the computing system may obtain query image embeddings and textual query refinements. Specifically, the computing system may obtain query image embeddings for the query image and textual query refinements associated with the query image. In some embodiments, obtaining the query image embeddings includes determining text embeddings for the textual query refinements, and processing the query image embeddings and textual query refinements with the machine-learned query refinement model includes processing the query image embeddings and text embeddings for the textual query refinements with the machine-learned query refinement model to obtain refined query image embeddings that incorporate the textual query refinements.

[0077] At 604, the computing system can process the query image embeddings and the text query refinement. Specifically, the computing system can process the query image embeddings and the text query refinement with a machine-learned query refinement model to obtain refined query image embeddings that incorporate the text query refinement.

[0078] At 606, the computing system may evaluate a loss function. Specifically, the computing system may evaluate a loss function that evaluates the distance between the refined query image embedding and an embedding for the ground truth image in the image embedding space. In some embodiments, before evaluating the loss function, the computing system obtains a corpus of image search data that includes search result images provided to a user in response to a query and refined search result images provided to a user in response to selection of a query refinement element that is provided to the user with the search result images. In some embodiments, the computing system selects a query image, a text query refinement, and a ground truth image from the search result images, the query refinement elements, and the refined search result images.

[0079] At 608, the computing system may modify one or more values of the one or more parameters of the machine-learned query refinement model based at least in part on the loss function.

[0080] In some embodiments, the computing system obtains a user query image and a text query refinement from a user for the user query image. The text query refinement is responsive to providing one or more initial result images to the user in response to the user query image. The computing system processes the user query image and the text query refinement for the user query image with a machine-learned query refinement model to obtain a refined image embedding of the user query image that incorporates the text query refinement.

[0081] In some embodiments, the computing system can obtain one or more refined result images in response to the refined image embedding of the user query image. In some embodiments, to obtain the refined result images, the computing system can select one or more image embeddings that are within a threshold distance of the refined image embedding of the user query image in the image embedding space. The one or more image embeddings can be associated with one or more refined result images, respectively. In some embodiments, the computing system can provide one or more refined result images. In some embodiments, to provide the refined result images, the computing system provides the one or more refined result images for display within an interface of a search application of the user's user device.

[0082] In some embodiments, the computing system may receive data indicating a user selection of at least one refined result image of the one or more result images. In some embodiments, the computing system may modify one or more values of one or more parameters of the machine-learned query refinement model based at least in part on the at least one refined result image.

[0083] 7 illustrates a flowchart diagram of an example method 700 for performing query refinement with a machine-learned query refinement model, according to an example embodiment of the present disclosure. While FIG. 7 depicts steps performed in a particular order for purposes of illustration and explanation, the methods of the present disclosure are not limited to the particularly depicted order or arrangement. Various steps of method 700 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0084] At 702, a computing system may obtain image embeddings for a query image provided by a user of the visual search application. In some embodiments, to obtain the image embeddings, the computing system obtains a query image from a user of the visual search application and determines image embeddings based at least in part on the query image, where the image embeddings represent the query image.

[0085] At 704, the computing system may obtain a text query refinement for the query image from a user of the visual search application. The text query refinement is responsive to providing one or more initial result images for the query image to the user of the visual search application. In some embodiments, the computing system determines one or more token embeddings that represent the text query refinement.

[0086] At 706, the computing system can process the image embeddings and the text query refinement for the query image using the machine-learned query refinement model to obtain refined image embeddings that incorporate the text query refinement.

[0087] At 708, the computing system may determine one or more refined result images based at least in part on the refined image embedding that incorporates the text query refinement. In some embodiments, to determine the one or more refined result images, the computing system determines one or more image embeddings within a threshold distance of the refined image embedding of the query image in the image embedding space and selects one or more refined result images that respectively correspond to the one or more image embeddings.

[0088] Additional Disclosures The technology described herein refers to servers, databases, software applications, and other computer-based systems, as well as actions performed on and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functionality among components. For example, the processes described herein can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0089] While the subject matter of the present disclosure has been described in detail with reference to various specific exemplary embodiments thereof, each example is provided by way of explanation and not by way of limitation of the present disclosure. Those skilled in the art, once they arrive at the foregoing understanding, will be able to readily produce modifications, variations, and equivalents to such embodiments. Accordingly, the disclosure of the subject matter does not exclude the inclusion of such modifications, variations, and / or additions to the subject matter as would be readily apparent to one skilled in the art. For example, features illustrated or described as part of one embodiment may be used with other embodiments to yield still other embodiments. Accordingly, the present disclosure is intended to cover such modifications, variations, and equivalents.

Claims

1. 1. A computing system for machine-learned multimodal search of images, comprising: one or more processors; a machine-learned query refinement model trained to refine image queries using text query refinement; one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations including: Obtaining image embeddings for a query image provided by a user of a visual search application; obtaining, from the user of the visual search application, a text query refinement for the query image, the text query refinement being responsive to providing one or more initial result images for the query image to the user of the visual search application; and processing the image embeddings and the text query refinement for the query image with the machine-learned query refinement model to obtain refined image embeddings that incorporate the text query refinement; determining one or more refined result images based at least in part on the refined image embeddings that incorporate the text query refinement; and one or more non-transitory computer-readable media, A computing system comprising:

2. Determining the one or more refined result images includes: determining one or more image embeddings within a threshold distance of the refined image embedding of the query image in an image embedding space; selecting the one or more refined result images corresponding respectively to the one or more image embeddings; The computing system of claim 1 , comprising:

3. Obtaining the image embedding for the query image includes: obtaining the query image from the user of the visual search application; determining the image embeddings based at least in part on the query image, the image embeddings representing the query image; and A computing system according to any one of claims 1 to 2, comprising:

4. obtaining the textual query refinement for the query image further includes determining one or more token embeddings representing the textual query refinement; processing the image embeddings and the text query refinement includes processing the image embeddings and the one or more token embeddings with the machine-learned query refinement model to obtain the refined image embeddings for the query image that incorporate the text query refinement. A computing system according to any one of claims 1 to 3.

5. The computing system of claim 1 , wherein the machine-learned query refinement model comprises a transformer model.

6. The operation is providing the one or more refined result images to a user device for display within an interface of the visual search application. The computing system of any one of claims 1 to 5, further comprising:

7. The operation is obtaining a second text query refinement for the query image in response to providing the one or more refined result images. The computing system of claim 6 further comprising:

8. 8. The computing system of claim 7, wherein the operations further include processing the second textual query refinement and the image embeddings of the query image with the machine-learned query refinement model to obtain a second refinement image embedding that incorporates the second textual query refinement.

9. 8. The computing system of claim 7, wherein the operations further include processing the refined image embeddings and the second textual query refinement with the machine-learned query refinement model to obtain second refined image embeddings that incorporate the textual query refinement and the second textual query refinement.

10. 1. A computer-implemented method comprising: obtaining, by a computing system including one or more computing devices, a query image embedding for a query image and a text query refinement associated with the query image; processing, by the computing system, the query image embeddings and the text query refinement with a machine-learned query refinement model to obtain refined query image embeddings that incorporate the text query refinement; evaluating, by the computing system, a loss function that measures the distance between the refined query image embedding and an embedding for a ground truth image in an image embedding space; modifying, by the computing system, one or more values of one or more parameters of the machine-learned query refinement model based at least in part on the loss function; A computer-implemented method comprising:

11. the query image depicts an entity having a first characteristic; the text query refinement describes a second characteristic about the entity that is different from the first characteristic; the ground truth image depicts the entity having the second characteristic.

11. The computer-implemented method of claim 10.

12. Obtaining the query image embedding and the text query refinement includes: determining, by the computing system, text embeddings for the text query refinement; further comprising processing the query image embeddings and the textual query refinement with the machine-learned query refinement model includes processing, by the computing system, the query image embeddings and the text embeddings for the textual query refinement with the machine-learned query refinement model to obtain refined query image embeddings that incorporate the textual query refinement. A computer-implemented method according to any one of claims 10 to 11.

13. The computer-implemented method of claim 12 , wherein the text embeddings for the text query refinement include multiple token embeddings.

14. Before evaluating the loss function, the method comprises: obtaining, by the computing system, a corpus of image search data including search result images that are provided to a user in response to a query and refined search result images that are provided to the user in response to selection of a query refinement element that is provided to the user along with the search result images; selecting, by the computing system, the query image, the text query refinement, and the ground truth image from the search result images, the query refinement elements, and the refined search result images; The computer-implemented method of any one of claims 10 to 13, comprising:

15. 15. The computer-implemented method of any one of claims 10 to 14, wherein the machine-learned query refinement model comprises a Transformer model.

16. The method comprises: obtaining, by the computing system from a user, a user query image and a text query refinement for the user query image, the text query refinement being responsive to providing one or more initial result images to the user in response to the user query image; processing, by the computing system, the user query image and the text query refinement for the user query image with the machine-learned query refinement model to obtain a refined image embedding of the user query image that incorporates the text query refinement; The computer-implemented method of any one of claims 10 to 15, further comprising:

17. The method comprises: obtaining, by the computing system, one or more refined result images in response to the refined image embedding of the user query image; providing, by the computing system, the one or more refined result images; 17. The computer-implemented method of claim 16, further comprising:

18. 20. The computer-implemented method of claim 17, wherein providing the one or more refined result images includes providing, by the computing system, the one or more refined result images for display within an interface of a search application of a user device of the user.

19. 20. The computer-implemented method of claim 17, wherein obtaining the one or more refined result images includes selecting, by the computing system, one or more image embeddings that are within a threshold distance of the refined image embedding of the user query image in the image embedding space, the one or more image embeddings being associated with the one or more refined result images, respectively.

20. The method comprises: receiving, by the computing system, data indicating a selection by the user of at least one refined result image of the one or more refined result images; modifying, by the computing system, one or more values of the one or more parameters of the machine-learned query refinement model based at least in part on the at least one refined result image; 20. The computer-implemented method of claim 18, further comprising:

21. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations, the operations including: Obtaining image embeddings for a query image provided by a user of a visual search application; obtaining, from the user of the visual search application, a text query refinement for the query image, the text query refinement being responsive to providing one or more initial result images for the query image to the user of the visual search application; and processing the image embeddings and the text query refinement for the query image with a machine-learned query refinement model to obtain refined image embeddings that incorporate the text query refinement; determining one or more refined result images based at least in part on the refined image embeddings that incorporate the text query refinement; and 1. One or more non-transitory computer-readable media, including:

Citation Information

Patent Citations

  • Join embedding for item association

    JP2013519138A

  • Text-conditioned image search based on transformation, aggregation, and composition of visio-linguistic features

    US20220245391A1