Visual search decisions for text-to-image substitution
The method and system address the challenge of capturing visual intent in search queries by replacing text with images, enhancing search accuracy and efficiency through multimodal search queries.
Patent Information
- Application Number
- JP2023178743
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-10-18
- Filing Date
- 2023-10-17
- Publication Date
- 2025-08-28
- Estimated Expiration
- 2043-10-17
AI Technical Summary
Existing search queries often fail to capture the visual intent of users, leading to inadequate search results when descriptive terms are used, as they may not fully convey the user's intended visual features.
A computer-implemented method and system that determines visual intent in search queries, providing an image selection interface to replace text with images, leveraging visual descriptors and machine learning models to enhance multimodal search queries.
Enhances search accuracy by allowing users to express visual intent through images, reducing the need for additional searches and improving computational efficiency by providing comprehensive multimodal search results.
Smart Images

Figure 0007730873000001 
Figure 0007730873000002 
Figure 0007730873000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to replacing text with images based on determined visual intent. More particularly, the present disclosure relates to processing text strings, determining visual intent, and providing an interface for image insertion. [Background technology]
[0002] A search query may include a text input to search for a specific item and / or specific knowledge. For example, a user may want to know the score of a particular sports game. Or, a user may want to learn more about a historical figure or find the contact address of a business.
[0003] Additionally, users may utilize search queries to search for specific objects to purchase, search for specific locations, etc. Search queries for specific objects or locations may include descriptive terms that may narrow the search results obtained, but may not capture the details that the user is trying to provide. Summary of the Invention [Means for solving the problem]
[0004] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the description that follows, or may be learned from the description, or may be learned by practice of the embodiments.
[0005] One exemplary aspect of the present disclosure is directed to a computer-implemented method for multimodal search. The method can include obtaining, by a computing system including one or more processors, a search query. The search query can include one or more words. The method can include determining, by the computing system, that the one or more words comprise visual intent. In some implementations, the visual intent can be associated with one or more visual features. The method can include providing, by the computing system, an image selection interface for display. The image selection interface can include a plurality of images for selection. In some implementations, the image selection interface can be provided for display based on a determination that the one or more words comprise visual intent. The method can include obtaining, by the computing system, selection data. The selection data can describe a selection of images. The method can include providing, by the computing system, an image for display in place of the one or more words. In some implementations, the method can include determining, by the computing system, one or more search results associated with the image and providing, by the computing system, the one or more search results as output.
[0006] In some implementations, providing the image selection interface for display can include providing, by the computing system, a user interface element. The user interface element can be descriptive of text replacement options. Providing the image selection interface for display can include obtaining, by the computing system, first input data. The first input data can describe a first selection of the text replacement options. Providing the image selection interface for display can include providing, by the computing system, the image selection interface for display based on the first input data.
[0007] In some implementations, the one or more search results can be provided via a search result page. The search result page can include a query box that displays the image. The search result page can include a search result panel for displaying information associated with the one or more search results. In some implementations, the search query can include one or more additional words. The one or more search results can be determined at least in part based on the one or more additional words. In some implementations, obtaining the search query can include obtaining the search query via a query box of a search interface. The one or more search results can include one or more image search results. In some implementations, the one or more search results can include one or more product search results that describe products associated with one or more visual features of the image.
[0008] Another exemplary aspect of the present disclosure is directed to a computing system for text-to-image substitution. The system can include one or more processors and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations can include obtaining text data. The text data can describe a plurality of text characters. The operations can include processing the text data to determine if a subset of the plurality of text characters includes a visually descriptive term. In some implementations, the visually descriptive term can be associated with one or more visual features. The operations can include providing an image selection interface for display. The image selection interface can include a plurality of images for selection. In some implementations, the plurality of images can be obtained based at least in part on the visually descriptive term. The operations can include obtaining selection data. The selection data can describe a selection of the image. The operations can include providing the image for display in place of the subset of the plurality of text characters.
[0009] In some implementations, providing the image selection interface for display can include providing an indicator for display. The indicator can describe text replacement options for replacing the visually descriptive term with the image data. Providing the image selection interface for display can include obtaining first input data. The first input data can describe a first selection of the text replacement options. Providing the image selection interface for display can include providing the image selection interface for display based on the first input data. In some implementations, the indicator can include a subset of the plurality of text characters displayed in one or more colors different from the remaining characters of the plurality of text characters.
[0010] In some implementations, the plurality of text characters can include a subset of the plurality of text characters and a second subset. The operations can include processing the image and the second subset to determine a plurality of search results. The plurality of search results can be determined based on the image and the second subset. The operations can include providing the plurality of search results in a search result page interface. In some implementations, the plurality of images can be obtained by querying a search engine with the subset of the plurality of text characters and receiving the plurality of images. The plurality of images can be obtained by determining that image data in a user-specific image database is associated with one or more visual features. The image data associated with the one or more visual features can include the plurality of images.
[0011] In some implementations, providing an image selection interface for display can include providing an image search option, a user image database option, and an image capture option. The image search option can include conducting a query on a network of computing systems using a subset of the plurality of text characters. The user image database option can include retrieving images from a user image database. The image capture option can include utilizing one or more image sensors of the user device. In some implementations, the visually descriptive terms can be determined based on historical search data. The historical search data can describe a plurality of terms previously utilized to retrieve one or more image search results. In some implementations, the visually descriptive terms can be determined based on processing the text data with a semantic understanding model.
[0012] Another example aspect of the present disclosure is directed to one or more non-transitory computer-readable media collectively storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations. The operations can include obtaining a plurality of words. The plurality of words can include one or more specific words and one or more additional words. The operations can include determining that one or more specific words of the plurality of words comprise a visual intent. In some implementations, the visual intent can be associated with one or more visual features. The operations can include providing the plurality of words for display with an indicator identifying the one or more specific words. The operations can include determining a plurality of images associated with the one or more specific words. The plurality of images can be associated with the visual intent. The operations can include providing the plurality of images to a user interface panel. In some implementations, the user interface panel can include a plurality of interactive user interface elements associated with the plurality of images. The operations can include obtaining a selection of a specific image of the plurality of images and providing the specific image for output without the one or more additional words and the one or more specific words.
[0013] In some implementations, the operations may include processing the output to generate a translated output. The translated output may be generated based at least in part on the particular image. The operations may include providing the output to a search engine and receiving a plurality of search results. In some implementations, the plurality of search results may be associated with one or more additional words and the particular image.
[0014] Other aspects of the present disclosure are directed to various systems, apparatus, non-transitory computer-readable media, user interfaces, and electronic devices.
[0015] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the associated principles.
[0016] Detailed descriptions of embodiments, directed to those skilled in the art, are set forth herein with reference to the accompanying drawings. [Brief explanation of the drawings]
[0017] [Figure 1A] FIG. 1 is a block diagram of an exemplary computing system for implementing text-to-image determination, according to an exemplary embodiment of the present disclosure. [Figure 1B] FIG. 1 is a block diagram of an exemplary computing device for implementing text-to-image determination, according to an exemplary embodiment of the present disclosure. [Figure 1C] FIG. 1 is a block diagram of an exemplary computing device for implementing text-to-image determination, according to an exemplary embodiment of the present disclosure. [Figure 2A] 10A-10C illustrate examples of exemplary query indicators according to exemplary embodiments of the present disclosure. [Figure 2B] 1A-1C illustrate examples of exemplary image selection interfaces, according to exemplary embodiments of the present disclosure. [Figure 2C] 1A-1C illustrate examples of exemplary image selection interfaces, according to exemplary embodiments of the present disclosure. [Figure 2D] 1A-1C illustrate examples of exemplary image selection interfaces, according to exemplary embodiments of the present disclosure. [Figure 3] FIG. 2 is a block diagram of an exemplary search interface according to an exemplary embodiment of the present disclosure. [Figure 4] 1A-1C illustrate examples of exemplary image selection interfaces, according to exemplary embodiments of the present disclosure. [Figure 5]FIG. 1 is a block diagram illustrating an exemplary text-to-image substitution system, according to an exemplary embodiment of the present disclosure. [Figure 6] FIG. 2 is a flowchart diagram of an exemplary method for performing text-to-image substitution, according to an exemplary embodiment of the present disclosure. [Figure 7] FIG. 1 is a flowchart diagram of an exemplary method for performing a multimodal search, according to an exemplary embodiment of the present disclosure. [Figure 8] FIG. 2 is a flowchart diagram of an exemplary method for performing text-to-image substitution, according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0018] Reference numbers that are repeated among the drawings are intended to identify like features in various implementations.
[0019] overview In general, the present disclosure is directed to systems and methods for expanding character strings by replacing text with visual tokens (e.g., images and / or videos). In particular, the systems and methods disclosed herein can leverage the determination of visual descriptors to prompt a user to replace text data with visual data to provide a multimodal output. For example, the systems and methods can be utilized to expand a search query to obtain a multimodal search query that can leverage both text data and image data to query a database. In some implementations, the systems and methods can include acquiring text data. The text data can describe a plurality of text characters. The systems and methods can include processing the text data to determine if a subset of the plurality of text characters includes a visually descriptive term. The visually descriptive term can be associated with one or more visual features. An indicator can be provided for display. The indicator can describe text replacement options for replacing the visually descriptive term with image data. The systems and methods can include acquiring first input data. In some implementations, the first input data can describe a first selection of the text replacement options. An image selection interface can be provided for display. The image selection interface can include a plurality of images for selection. The systems and methods can include obtaining second input data. In some implementations, the second input data can describe a second selection of images. The images can be provided for display in place of a subset of the plurality of text characters.
[0020] The systems and methods may acquire text data. The text data may describe a plurality of text characters. The plurality of text characters may describe one or more words. The plurality of characters may be acquired via one or more inputs to a user interface. Alternatively and / or additionally, the text data may be generated by processing audio data associated with a verbal utterance.
[0021] The text data can be processed to determine whether a subset of the plurality of text characters includes a visually descriptive term. The visually descriptive term can be associated with one or more visual features. In some implementations, the visually descriptive term can be determined based on historical search data. The historical search data can describe a plurality of terms utilized to retrieve one or more image search results. In some implementations, the visually descriptive term can be determined based on processing the text data with a semantic understanding model. The visually descriptive term can be determined based on historical click data. The historical selection data can be global selection data, user-specific historical selection data, region-specific historical selection data, and / or context-specific historical selection data. In some implementations, the historical selection data can describe how often the image search tab is selected when a particular term is entered.
[0022] The present systems and methods can provide an indicator for display. The indicator can describe text replacement options for replacing visually descriptive terms with image data. The indicator can include a subset of a plurality of text characters displayed in one or more colors different from the remaining characters of the plurality of text characters. In some implementations, the indicator can include a pop-up user interface element. The indicator can include highlighting one or more words, underlining one or more words, circling one or more words, and / or flashing one or more words.
[0023] First input data can then be obtained. The first input data can describe a first selection of a text replacement option. The first input data can describe audio input (e.g., a voice command), touch input (e.g., input to a touchscreen), keyboard input, and / or mouse input. The first input data can include a selection of an indicator.
[0024] An image selection interface can then be provided for display. The image selection interface can include a plurality of images for selection. The plurality of images can be obtained by determining that the image data in the user-specific image database includes a plurality of images. In some implementations, the plurality of images can be associated with one or more visual features. In some implementations, the plurality of images can be obtained based on one or more visually descriptive terms. In some implementations, the image selection interface can be provided immediately after determining the visually descriptive terms. Alternatively and / or additionally, the image selection interface can be provided in response to receiving the first input data.
[0025] In some implementations, the plurality of images can be obtained by querying a search engine using a subset of the plurality of text characters and receiving the plurality of images. The query utilized in querying the search engine can include visually descriptive terms. Additionally and / or alternatively, one or more contexts can be obtained and / or determined. The one or more contexts can then be utilized to refine the search. The one or more contexts can include user-specific information (e.g., user location, application history, user search history, user purchase history, user preferences, and / or user profile). In some implementations, the one or more contexts can include time of day, day of the week, time of year, global trends, and / or past selection of images when particular visually descriptive terms are used.
[0026] Additionally and / or alternatively, providing an image selection interface for display may include providing an image search option, a user image database option, and an image capture option. The image search option may include querying the web (e.g., a network of computing systems) using a subset of the plurality of text characters. The user image database option may include retrieving images from a user image database. The image capture option may include utilizing one or more image sensors of the user device. The user image database may be associated with one or more user profiles and may also be associated with one or more image gallery applications. In some implementations, the user image database option enables selection of locally stored data. Alternatively and / or additionally, the user image database option may enable a user to select images stored associated with the user in one or more image storage applications, which may include cloud storage, server storage, and / or local storage.
[0027] The present systems and methods can obtain second input data (e.g., selection data). The second input data can describe a second selection of an image. The second input data can describe audio input (e.g., a voice command), touch input (e.g., input to a touchscreen), keyboard input, and / or mouse input. The first input data can include a selection of a selection icon, a selection of a thumbnail, and / or a drop-and-drag selection.
[0028] The image can then be provided for display in place of the subset of text characters. For example, the subset of text characters can be deleted and the image can be added in place of the subset of text characters before deletion.
[0029] In some implementations, the plurality of text characters may include a subset of the plurality of text characters and a second subset. The systems and methods may include processing the image and the second subset to determine a plurality of search results. In some implementations, the plurality of search results may be determined based on the image and the second subset. The plurality of search results may then be provided in a search result page interface.
[0030] The present systems and methods can be utilized for multimodal search. In particular, one or more words in a query string can be replaced with an image to generate a more comprehensive search query. For example, the present systems and methods can include obtaining a search query. The search query can include one or more words. The one or more words can be determined to include visual intent. In some implementations, the visual intent can be associated with one or more visual features. The present systems and methods can include providing an image selection interface for display. The image selection interface can include a plurality of images for selection. In some implementations, the image selection interface can be provided for display based on the determination of one or more words that include visual intent. The present systems and methods can include obtaining selection data. The selection data can describe a selection of the image. The image can then be provided for display as a substitute for the one or more words. Additionally and / or alternatively, the present systems and methods can include determining one or more search results associated with the image and providing the one or more search results as output.
[0031] The system and method can obtain a search query. The search query can include one or more words. In some implementations, obtaining the search query can include obtaining the search query through a query box of a search interface. The search interface can be provided by a web platform, a mobile application, and / or a desktop application. The search query can include Boolean terms, syntax, and / or natural language constructs.
[0032] One or more words may be determined to comprise a visual intent. The visual intent may be associated with one or more visual features. The visual intent may be based on one or more words associated with a color, pattern, design, object, and / or visual feature. The association may be based on one or more words that are visual descriptors, one or more words associated with a label for a particular visual feature, and / or one or more words associated with previous image search queries. Words that describe a color, pattern, shape, and / or other visual descriptors may be determined to comprise a visual intent.
[0033] The systems and methods can provide user interface elements. In some implementations, the user interface elements can describe text replacement options. The user interface elements can be indicators that the systems and methods have determined that one or more words are associated with a visual intent. The user interface elements can include visual effects. The user interface elements can include pop-up elements, drop-down menus, changes to the display of one or more words, and / or the appearance of an icon.
[0034] The system and method can then obtain first input data. The first input data can describe a first selection of a text replacement option. The first input data can include sensor data. The first input data can describe an interaction with a user interface element (e.g., a tap input, a gesture input, and / or a lack of input due to a threshold time elapsed without input being obtained).
[0035] An image selection interface can then be provided for display. The image selection interface can include multiple images for selection. The image selection interface can include one or more different tabs for viewing and selecting images from different databases and / or different media or types. The image selection interface can include one or more panels for providing different types of media content items and / or media content items from different sources.
[0036] The systems and methods can then obtain second input data (e.g., selection data). The second input data (e.g., selection data) can describe a selection of an image. The second input data can include sensor data. The second input data can describe an interaction with the image selection interface (e.g., a tap input, a gesture input, and / or a lack of input due to a threshold time lapse without any input being obtained).
[0037] The image may then be provided for display in place of one or more words. For example, a preview and / or thumbnail of the image may be provided for display in a query box of a search interface.
[0038] The systems and methods can include determining one or more search results associated with the image. In some implementations, the one or more search results can be provided via a search results page. The search results page can include a query box that displays the image. Additionally and / or alternatively, the search results page can include a search results panel for displaying information associated with the one or more search results. The search query can include one or more additional words. In some implementations, the one or more search results can be determined based at least in part on the one or more additional words. The one or more search results can include one or more image search results. Additionally and / or alternatively, the one or more search results can include one or more product search results that describe products associated with one or more visual features of the image.
[0039] One or more search results may be provided as output. One or more search results may be provided for display in a search results page interface. The search results may be provided in different panels based on the type of search result, the source of the search results, and / or the classification of the search results.
[0040] The system and method may include obtaining a plurality of words. The plurality of words may include one or more specific words and one or more additional words. The system and method may include determining that one or more specific words of the plurality of words include visual intent. An indicator identifying the one or more specific words may be provided for display of the plurality of words. The system and method may include determining a plurality of images associated with the one or more specific words. The plurality of images may be provided to a user interface panel. The system and method may include obtaining a selection of a specific image of the plurality of images and providing the specific image for output without the one or more additional words and the one or more specific words.
[0041] The systems and methods can include obtaining a plurality of words. The plurality of words can include one or more specific words and one or more additional words. The one or more specific words can include visually descriptive terms. The one or more additional words can be complementary to the one or more specific words and / or can target different descriptive aspects of the search query or phrase.
[0042] The systems and methods can then include determining that one or more particular words of the plurality of words contain visual intent. The determination can be based on processing the plurality of words with one or more machine learning models to generate one or more outputs. The one or more machine learning models can include one or more detection models, one or more segmentation models, one or more classification models, and / or one or more expansion models. In some implementations, the one or more machine learning models can include one or more natural language processing models. The one or more machine learning models can include one or more transformer models. In some implementations, the determination can be based on historical search data.
[0043] A multi-word indicator identifying one or more particular words may be provided for display. The indicator may be a visual indicator describing one or more possible actions that can be performed based on the identified one or more particular words. The indicator may include an explanation, may include a change in text color, may include highlighting, and / or may include a pop-up element.
[0044] A plurality of images associated with the one or more particular words can then be determined. This determination can be based on querying a database using the one or more particular words. The database can be a local database stored on the user's device and / or a database accessed via a network connection. The one or more images can be cropped to isolate a particular portion of the image associated with the one or more particular words.
[0045] The multiple images can then be provided for display in a user interface panel, which may be a pop-up panel and / or may replace a portion of the originally displayed interface.
[0046] A selection of a particular image of the plurality of images can be obtained. In some implementations, the particular image can be a cropped image from an image database. The cropped image can be generated by processing an uncropped image with one or more machine learning models to detect relevant portions of the image and segmenting the relevant portions from the uncropped image.
[0047] The one or more additional words and the particular image may be provided as output without the one or more particular words. The particular image may be positioned where the one or more particular words were previously displayed. In some implementations, a thumbnail and / or preview may be provided for display in place of the complete particular image.
[0048] In some implementations, the systems and methods can include processing the output to generate a translation output. The translation output can be generated based at least in part on the particular image.
[0049] Alternatively and / or additionally, the systems and methods may include providing the output to a search engine and receiving a plurality of search results, which may be associated with one or more additional words and a particular image.
[0050] Users may be accustomed to expressing the visual portion of a question in text. However, parts of a question may be better expressed using an image. For example, a user may be inspired by a dress they saw on social media. However, the user may instead desire a sock pattern. To search for socks with a specific pattern, a user may enter the query "colorful floral socks," but "colorful floral" may lose fidelity to the intent. A more accurate search might consider replacing "colorful floral" with an actual image the user saw.
[0051] The systems and methods disclosed herein can detect strings of characters that appear to have visual intent and can highlight that portion of the string. When the user taps the highlight, the systems and methods can trigger a visual search tool, providing the user with an easy way to exchange strings of characters for image tokens.
[0052] The systems and methods of the present disclosure provide many technical effects and advantages. As one example, the systems and methods can provide a text-to-image replacement interface. In particular, the systems and methods disclosed herein can utilize an interactive user interface to determine candidate images to provide to a user for selection to replace one or more words.
[0053] Another technical advantage of the systems and methods of the present disclosure is that they can leverage visual intent determination to determine when and to what extent a text-to-image replacement interface can be provided. For example, the systems and methods can determine that one or more words are associated with a visual intent. The systems and methods can determine that an indicator is provided that enables a user to open a text-to-image replacement interface to replace the one or more words with one or more images.
[0054] Another example of technical effects and advantages relates to increased computational efficiency and improved functionality of computing systems. For example, the systems and methods disclosed herein can leverage text-to-image substitution to provide more comprehensive multi-modal search queries that can reduce the use of additional searches and browsing additional search result pages, thereby saving time and computational power.
[0055] Referring now to the drawings, exemplary embodiments of the present disclosure will be described in more detail.
[0056] Exemplary Devices and Systems 1A illustrates a block diagram of an exemplary computing system 100 for implementing text-to-image determination in accordance with an exemplary embodiment of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.
[0057] The user computing device 102 may be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0058] The user computing device 102 includes one or more processors 112 and memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple operatively coupled processors. The memory 114 may include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 may store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.
[0059] In some implementations, the user computing device 102 can store or include one or more visual intention decision models 120. For example, the visual intention decision model 120 can be or otherwise include various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models including nonlinear and / or linear models. The neural networks can include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Examples of the visual intention decision model 120 are described with reference to FIGS. 2A-5.
[0060] In some implementations, one or more visual intention decision models 120 may be received from the server computing system 130 over the network 180, stored in the user computing device memory 114, and then used or implemented by one or more processors 112. In some implementations, the user computing device 102 may implement multiple parallel instances of a single visual intention decision model 120 (e.g., to perform parallel visual intention decision across multiple instances of a text string).
[0061] More specifically, the visual intent determination model 120 can process one or more words to determine whether the one or more words are associated with a visual intent. The visual intent determination model 120 can include one or more classification models, one or more segmentation models, and / or one or more detection models. The visual intent determination model 120 can include a natural language model. In some implementations, the visual intent determination model 120 can generate a semantic understanding output that describes the semantic understanding of the text string.
[0062] Additionally or alternatively, one or more visual intention decision models 140 may be included in or otherwise stored and implemented on a server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the visual intention decision model 140 may be implemented by the server computing system 130 as part of a web service (e.g., a text-to-image substitution service). Thus, one or more models 120 may be stored and implemented on the user computing device 102 and / or one or more models 140 may be stored and implemented on the server computing system 130.
[0063] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) that senses the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component functions to implement a virtual keyboard. Other examples of user input components include a microphone, a traditional keyboard, or other means by which a user can provide user input.
[0064] The server computing device 130 includes one or more processors 132 and memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple operatively coupled processors. The memory 134 may include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 may store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing device 130 to perform operations.
[0065] In some implementations, server computing system 130 includes or is implemented by one or more server computing devices. When server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a serial computing architecture, a parallel computing architecture, or a combination thereof.
[0066] As described above, the server computing system 130 can store or otherwise include one or more machine-learned visual intent-decision models 140. For example, the models 140 can be or include various machine-learned models. Examples of machine-learned models include neural networks or other multi-layer nonlinear models. Examples of neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Exemplary models 140 are described with reference to FIGS. 2A-5.
[0067] The user computing device 102 and / or the server computing system 130 can train the models 120 and / or 140 through interaction with a training computing system 150 that is communicatively coupled via a network 180. The training computing system 150 may be separate from or part of the server computing system 130.
[0068] The training computing system 150 includes one or more processors 152 and memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple operably coupled processors. The memory 154 may include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 may store data 156 and instructions 158 executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is implemented by one or more server computing devices.
[0069] The training computing system 150 may include a model trainer 160 that trains the machine learning models 120 and / or 140 stored on the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as, for example, backpropagation of errors. For example, a loss function may be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent may be used to iteratively update the parameters over multiple training iterations.
[0070] In some implementations, performing backpropagation of errors may include performing truncated backpropagation over time. The model trainer 160 can implement many generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.
[0071] In particular, model trainer 160 can train visual intent decision models 120 and / or 140 based on a set of training data 162. Training data 162 can include, for example, training words and phrases, ground truth labels, historical search queries, historical selection data associated with query refinements, large-scale linguistic datasets, and / or ground truth semantic intent mappings.
[0072] In some implementations, if the user consents, the training examples may be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 based on user-specific data received from the user computing device 102. In some cases, this process may be referred to as personalizing the model.
[0073] Model trainer 160 includes computer logic utilized to provide desired functionality. Model trainer 160 can be implemented in hardware, firmware, and / or software controlling a general-purpose processor. For example, in some implementations, model trainer 160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, model trainer 160 includes one or more sets of computer-executable instructions stored on a tangible computer-readable storage medium, such as RAM, a hard disk, an optical medium, or a magnetic medium.
[0074] Network 180 may be any type of communications network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. In general, communications on network 180 may be conducted over any type of wired and / or wireless connections using a variety of communications protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).
[0075] The machine learning models described herein may be used in a variety of tasks, applications, and / or use cases.
[0076] In some implementations, the input to the machine learning model of the present disclosure may be image data. The machine learning model may process the image data to generate an output. As an example, the machine learning model may process the image data to generate an image recognition output (e.g., a recognition of the image data, a latent embedding of the image data, an encoded representation of the image data, a hash of the image data, etc.). As another example, the machine learning model may process the image data to generate an image segmentation output. As another example, the machine learning model may process the image data to generate an image classification output. As another example, the machine learning model may process the image data to generate an image data modification output (e.g., a modification of the image data, etc.). As another example, the machine learning model may process the image data to generate an encoded image data output (e.g., an encoded and / or compressed representation of the image data, etc.). As another example, the machine learning model may process the image data to generate an upscaled image data output. As another example, the machine learning model may process the image data to generate a prediction output.
[0077] In some implementations, the input to the machine learning model of the present disclosure may be text or natural language data. The machine learning model may process the text or natural language data to generate an output. As an example, the machine learning model may process the natural language data to generate a language encoding output. As another example, the machine learning model may process the text or natural language data to generate a potential text embedding output. As another example, the machine learning model may process the text or natural language data to generate a translation output. As another example, the machine learning model may process the text or natural language data to generate a classification output. As another example, the machine learning model may process the text or natural language data to generate a text segmentation output. As another example, the machine learning model may process the text or natural language data to generate a semantic intent output. As another example, the machine learning model may process the text or natural language data to generate an upscaled text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language). As another example, the machine learning model may process the text or natural language data to generate a predicted output.
[0078] In some implementations, the input to the machine learning model of the present disclosure may be speech data. The machine learning model may process the speech data to generate an output. As an example, the machine learning model may process the speech data to generate a speech recognition output. As another example, the machine learning model may process the speech data to generate a speech translation output. As another example, the machine learning model may process the speech data to generate a potential embedded output. As another example, the machine learning model may process the speech data to generate an encoded speech output (e.g., an encoded and / or compressed representation of the speech data, etc.). As another example, the machine learning model may process the speech data to generate an upscaled speech output (e.g., speech data of higher quality than the input speech data, etc.). As another example, the machine learning model may process the speech data to generate a text representation output (e.g., a text representation of the input speech data, etc.). As another example, the machine learning model may process the speech data to generate a predicted output.
[0079] In some implementations, the input to the machine learning model of the present disclosure may be latent encoding data (e.g., a latent spatial representation of the input, etc.). The machine learning model may process the latent encoding data to generate an output. As an example, the machine learning model may process the latent encoding data to generate a recognition output. As another example, the machine learning model may process the latent encoding data to generate a reconstruction output. As another example, the machine learning model may process the latent encoding data to generate a search output. As another example, the machine learning model may process the latent encoding data to generate a reclustering output. As another example, the machine learning model may process the latent encoding data to generate a prediction output.
[0080] In some implementations, input to a machine learning model of the present disclosure may be statistical data. The machine learning model may process the statistical data to generate an output. As an example, the machine learning model may process the statistical data to generate a recognition output. As another example, the machine learning model may process the statistical data to generate a prediction output. As another example, the machine learning model may process the statistical data to generate a classification output. As another example, the machine learning model may process the statistical data to generate a segmentation output. As another example, the machine learning model may process the statistical data to generate a segmentation output. As another example, the machine learning model may process the statistical data to generate a visualization output. As another example, the machine learning model may process the statistical data to generate a diagnostic output.
[0081] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data of one or more images and the task is an image processing task. For example, the image processing task is image classification and the output is a set of scores, each score corresponding to a different object class and representing the likelihood that one or more images depict an object belonging to that object class. The image processing task may be object detection and the image processing output identifies one or more regions in one or more images and, for each region, the likelihood that the region represents an object of interest. As another example, the image processing task may be image segmentation and the image processing output defines, for each pixel in one or more images, a respective likelihood for each category in a predetermined set of categories. For example, the set of categories may be foreground and background. As another example, the set of categories may be object classes. As another example, the image processing task may be depth estimation and the image processing output defines, for each pixel in one or more images, a respective depth value. As another example, the image processing task may be motion estimation, where the network input includes multiple images and the image processing output defines, for each pixel of one of the input images, the motion of the scene depicted at that pixel between images in the network input.
[0082] In some cases, the input includes audio data representing verbal speech and the task is a speech recognition task. The output may comprise text output mapped to the verbal speech. In some cases, the task comprises encryption or decryption of input data. In some cases, the task comprises a microprocessor performance task such as branch prediction or memory address translation.
[0083] 1A illustrates one exemplary computing system that can be used to implement the present disclosure. Other computing systems can also be used. For example, in some implementations, a user computing device 102 can include a model trainer 160 and a training dataset 162. In such implementations, the model 120 can be trained and used locally on the user computing device 102. In some such implementations, the user computing device 102 can implement the model trainer 160 to personalize the model 120 based on user-specific data.
[0084] 1B illustrates a block diagram of an exemplary computing device 10 implemented in accordance with an exemplary embodiment of the present disclosure. Computing device 10 may be a user computing device or a server computing device.
[0085] The computing device 10 includes multiple applications (e.g., applications 1 through N). Each application includes its own machine learning library and machine learning model. For example, each application may include a machine learning model. Examples of applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.
[0086] 1B , each application can communicate with numerous other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.
[0087] 1C illustrates a block diagram of an exemplary computing device 50 for implementation in accordance with an exemplary embodiment of the present disclosure. Computing device 50 may be a user computing device or a server computing device.
[0088] Computing device 50 includes multiple applications (e.g., applications 1 through N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API across all applications).
[0089] The central intelligence layer includes multiple machine learning models. For example, as shown in FIG. 1C , a respective machine learning model (e.g., model) can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model (e.g., a single model) for all applications. In some implementations, the central intelligence layer is included within or implemented by the operating system of computing device 50.
[0090] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for computing device 50. As shown in FIG. 1C , the central device data layer can communicate with numerous other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0091] Exemplary System Configuration FIG. 2A illustrates an example of an exemplary query indicator, according to an exemplary embodiment of the present disclosure. In particular, FIG. 2A illustrates a query input box 204 within a search interface 202. The query input box 204 can be configured to receive and / or display an input text string utilized as a search query. For example, a user may have provided one or more inputs to generate the search query "clutch with floral pattern." The search query can be processed to determine whether one or more particular words 208 are associated with a visual intent. The one or more particular words 208 can then be provided for display with an indicator (e.g., the one or more particular words 208 can be provided in a different color and / or highlighted). One or more other words 206 within the search query can be provided for display in a regular format and / or in a different indicator.
[0092] An indicator associated with the visual intent can be selected to initiate the generated and / or provided image selection interface. The indicator may be provided in real time during input and / or as the search query is processed and search results are provided for display.
[0093] In some implementations, a search query may be entered via a keyboard (e.g., a physical keyboard and / or a graphical keyboard), via a mouse, and / or via voice input (e.g., a user may select the voice command icon 210 to begin recording a voice utterance for processing and transcription). Additionally and / or alternatively, visual intent determination and / or ranking of search results may be based in part on the user profile 212.
[0094] FIG. 2B illustrates an example of an exemplary image selection interface 220 according to an exemplary embodiment of the present disclosure. In particular, FIG. 2B illustrates an example of an image selection interface 220 for selecting an image from a user-specific image gallery. For example, an indicator may be provided in a search query input box 222 that may be selected to transition from a search results page 224 to an initial image selection page 226. The image selection page 226 may include multiple panels, including a recent images panel, an all images panel, and / or a relevance panel. The recent images panel may include recently saved images. The all images panel may include an interface for accessing all images in the user-specific image gallery. The all images panel may include images sorted based on the image's save date, the image's name, and / or the image's relevance to one or more specific words associated with a visual intent. The relevance panel may include one or more images from the user-specific image gallery that are determined to be most relevant to one or more specific words and / or visual intent. Relevance may be determined based on one or more features detected in the image, image metadata, the image's source, the image's name, and / or the image capture location.
[0095] Once an image is selected, the selected image may be processed to determine a region of interest. An indicator may be provided for display with each candidate region of interest in the region selection interface 228. The region of interest may be determined based on the image being processed by one or more machine learning models to detect one or more features within the image. The user may then select a particular candidate region, which will provide the cropping interface 230. The cropping interface 230 may provide a suggested crop region based on the selected candidate region and / or based on one or more other user inputs.
[0096] Once the cropped area is confirmed, the image 232 (or a thumbnail of the image) can replace one or more particular words and can be provided for display in a query input box. Search results can then be refined based on the image, and an updated search results page 234 can be provided for display.
[0097] 2C illustrates an example of an exemplary image selection interface 240, according to an exemplary embodiment of the present disclosure. In particular, FIG. 2C illustrates an example of an image selection interface 240 for capturing an image. For example, a search query can be provided, the search query can be processed to determine visual intent, and an indicator 242 can be provided. Selecting the indicator 242 can transition the search interface from a search results interface 244 to an image capture interface 248. The image capture option can be selected from multiple options 246 provided by the image selection interface 240.
[0098] An image can then be captured using one or more image sensors of the user's computing device. Image selection interface 240 may then provide cropping options 250 to the user. Cropping options 250 may include an automatically suggested cropping area. Alternatively and / or additionally, cropping options 250 may allow the user to manually crop the captured image to provide a more specific input area.
[0099] The cropped region can then be added to the search query (e.g., to replace and / or complement visually descriptive terms) to generate a multimodal query 252. Multiple search results can then be provided in an updated search results interface 254 based on the multimodal query 252.
[0100] 2D illustrates an example of an exemplary image selection interface 260, according to an exemplary embodiment of the present disclosure. In particular, FIG. 2D illustrates an example image selection interface 260 for selecting images using a search engine. For example, a search query can be provided, the search query can be processed to determine visual intent, and an indicator 262 can be provided. Selecting the indicator 262 can transition the search interface from a search results interface 264 to an image search interface 268. Image search options can be selected from multiple options 266 provided by the image selection interface 260.
[0101] The image search interface 268 can process one or more particular words of the search query associated with the visual intent to determine a plurality of candidate images. The user can then select a particular image, which can transition the image selection interface 260 to a region selection stage 270. The user can select a region, and the image selection interface 260 can provide cropping options 272, which can enable automatic cropping and / or manual cropping.
[0102] Once cropping is complete, an updated search results page 276 may be provided. The search results on the updated search results page 276 may be based on a multimodal query 274 that includes one or more words of the original search query and at least a portion of the selected image.
[0103] 3 illustrates a block diagram of an exemplary search interface 300, according to an exemplary embodiment of the present disclosure. The systems and methods disclosed herein may enable the expansion of a search query 304 to generate a multimodal search query that can be processed by a search engine 302. The search query 304 may be entered into a query input box of the search engine 302 and may include visual descriptors associated with visual intent.
[0104] The search query 304 can be processed to determine multiple search results, which can be utilized to generate a search results page 306. The search results page 306 can include a query input box including the search query with an indicator 308 indicating one or more determined visual descriptors. The search query with indicator 308 can indicate that a visual intent has been determined and that an interface can be opened to refine the search by generating a multimodal search query. The search results page 306 can include a first search result 310, a second search result 312, a third search result 314, and / or an nth search result 316. Based on the refinement of the search by generating a multimodal search query, the search results page 306 can be updated to include the same search results with different rankings, different search results, and / or a mix of new and previously displayed search results.
[0105] FIG. 4 illustrates an example of an exemplary image selection interface 400, according to an exemplary embodiment of the present disclosure. In some implementations, a user-specific image gallery option 410, an image capture option 420, and / or an image search option 430 may be provided in response to a selection of a text replacement option. The user-specific image gallery option 410, the image capture option 420, and the image search option 430 may each have their own respective icon that can be associated with the particular option. The icons may be selectable to move from one option to another. For example, the user-specific image gallery option 410 may be associated with an overlapping tile icon 412, the image capture option 420 may be associated with a camera icon 422, and the image search option 430 may be associated with a globe icon 432 indicating a global search for images.
[0106] Each option may provide different and / or overlapping sources for images. The user-specific image gallery option 410 may provide images from one or more image galleries specifically associated with the user. The image galleries may be stored locally on the user device and / or on a server computing system. The user-specific image gallery option 410 may include different panels for interaction, which may include a recent screenshots panel 414, a recent camera captures panel, and / or an all images panel 416.
[0107] Image capture options 420 may utilize one or more image sensors of the user device and may include image capture user interface elements 424 for determining when and / or what to capture in the environment.
[0108] The image search option 430 can utilize a search engine to retrieve image data from multiple sources on the Internet. The image search option 430 can utilize one or more words of an input search query to query the search engine. In some implementations, a new query can be entered via a dedicated search query box 434. Alternatively and / or additionally, one or more words can be adjusted. Multiple image search results can be displayed and / or interacted with by the user.
[0109] 5 illustrates a block diagram of an exemplary text-to-image substitution system 500, according to an exemplary embodiment of the present disclosure. The text-to-image substitution system 500 can process text data 502 to generate augmented data 516. The text data 502 can describe a plurality of characters associated with one or more words. The one or more words can be associated with a search query, a text string in a blog, a text string in a message, and / or a response to a question or prompt.
[0110] The text data 502 can be processed to determine whether one or more particular words associated with the text data 502 are associated with a visual intent (e.g., the one or more words are visually descriptive words (e.g., describing one or more visual features)). This determination can be made based on historical data 504, heuristics, and / or one or more machine learning models (e.g., visual intent determination model 508). For example, the historical data 504 can describe past interactions by a user when one or more particular words were used. In some implementations, a user and / or multiple users can narrow search results to images when using one or more particular words. Alternatively and / or additionally, the one or more particular words may be commonly used in describing images (e.g., in image captions). The one or more particular words can be determined to be associated with a visual intent based on a common association with images and / or image features. In some implementations, the natural language meaning of a word or phrase can be utilized to determine that one or more particular words are associated with a visual intent.
[0111] Additionally and / or alternatively, one or more machine learning models (e.g., visual intent determination model 508) can be utilized to determine that one or more particular words are associated with the visual intent. The visual intent determination model 508 can analyze the text data, process each segment to provide a classification for each segment, and generate output data 510 that describes whether the text data includes one or more particular words associated with the visual intent. Alternatively and / or additionally, the visual intent determination model 508 can include a natural language processing model that can process the text data as a whole and / or in various syntactically determined segments to generate the output data 510.
[0112] Based on the determination of one or more particular words associated with the visual intent, an indicator 506 may be provided for display. The indicator 506 may include one or more particular words with different and / or changing colors. The indicator 506 and / or one or more other user interface elements may be selected. A text-to-image substitution interface 512 may then be provided. The user may then select to search a user-specific image gallery, capture a new image, and / or search the web (e.g., a network of computing systems) for a particular image 514 to be utilized in place of and / or with the portion of the text data 502.
[0113] The selected particular image 514 can then be utilized to enhance the text data 502 and generate enhanced data 516, which can include both text data and image data. In some implementations, the selected particular image 514 can be processed prior to enhancing the text data 502. For example, the particular image 514 can be processed by one or more machine learning models (e.g., a cropping model 518) to generate an enhanced image to add to the text data 502. In particular, the particular image 514 can be processed by the cropping model 518 to determine one or more portions of the particular image 514 to segment and generate a cropped image 520. The cropped image 520 can then be utilized to generate the enhanced data 516. The cropping model 518 can include one or more detection models, one or more classification models, and / or one or more segmentation models. The cropping model can determine whether one or more objects are depicted in the particular image 514, can determine one or more regions associated with the one or more objects, and can provide suggested cropping regions to the user. Alternatively and / or additionally, the cropping model 518 may determine which of multiple regions of a particular image 514 are associated with one or more particular words. For example, if the one or more particular words include "pattern," the cropping model 518 may determine to segment a portion of a striped dress rather than segmenting the solid wallpaper on the wall.
[0114] Exemplary Methods 6 shows a flowchart diagram of an exemplary method performed in accordance with an exemplary embodiment of the present disclosure. While FIG. 6 shows steps performed in a particular order for purposes of illustration and discussion, the methods of the present disclosure are not limited to the particularly illustrated order or arrangement. Various steps of method 600 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.
[0115] At 602, the computing system may obtain text data. The text data may describe a plurality of text characters. The plurality of text characters may describe one or more words. The plurality of characters may be obtained via one or more inputs to a user interface. Alternatively and / or additionally, the text data may be generated by processing audio data associated with verbal speech.
[0116] At 604, the computing system can process the text data to determine whether a subset of the plurality of text characters includes a visually descriptive term. The visually descriptive term can be associated with one or more visual features. In some implementations, the visually descriptive term can be determined based on historical search data. The historical search data can describe a plurality of terms utilized to retrieve one or more image search results. In some implementations, the visually descriptive term can be determined based on processing the text data with a semantic understanding model. The visually descriptive term can be determined based on historical click data. The historical selection data can be global selection data, user-specific historical selection data, region-specific historical selection data, and / or context-specific historical selection data. In some implementations, the historical selection data can describe how often the image search tab is selected when a particular term is entered.
[0117] At 606, the computing system may provide an indicator for display. The indicator may describe text replacement options for visually replacing the descriptive term with the image data. The indicator may include a subset of the plurality of text characters displayed in one or more colors different from the remaining characters of the plurality of text characters. In some implementations, the indicator may include a pop-up user interface element. The indicator may include highlighting one or more words, underlining one or more words, circling one or more words, and / or flashing one or more words.
[0118] At 608, the computing system may obtain first input data. The first input data may describe a first selection of a text replacement option. The first input data may describe audio input (e.g., a voice command), touch input (e.g., input to a touchscreen), keyboard input, and / or mouse input. The first input data may include a selection of an indicator.
[0119] At 610, the computing system may provide an image selection interface for display. The image selection interface may include a plurality of images for selection. In some implementations, the plurality of images may be obtained based at least in part on visually descriptive terms. The plurality of images may be obtained by determining that image data in a user-specific image database is associated with one or more visual features. The computing system may determine that the image data associated with the one or more visual features includes a plurality of images. In some implementations, the plurality of images may be obtained based on the one or more visually descriptive terms.
[0120] In some implementations, the plurality of images can be obtained by querying a search engine using a subset of the plurality of text characters and receiving the plurality of images. The query utilized in querying the search engine can include visually descriptive terms. Additionally and / or alternatively, one or more contexts can be obtained and / or determined. The one or more contexts can then be utilized to refine the search. The one or more contexts can include user-specific information (e.g., user location, application history, user search history, user purchase history, user preferences, and / or user profile). In some implementations, the one or more contexts can include time of day, day of the week, time of year, global trends, and / or past selection of images when particular visually descriptive terms are used.
[0121] Additionally and / or alternatively, providing an image selection interface for display can include providing an image search option, a user image database option, and an image capture option. The image search option can include querying the web using a subset of the plurality of text characters. The user image database option can include retrieving images from a user image database. The image capture option can include utilizing one or more image sensors of the user device. The user image database can be associated with one or more user profiles and can also be associated with one or more image gallery applications. In some implementations, the user image database option allows selection of locally stored data. Alternatively and / or additionally, the user image database option can allow a user to select images stored associated with the user in one or more image storage applications, which can include cloud storage, server storage, and / or local storage.
[0122] In some implementations, the computing system may provide the image selection interface without providing an indicator and / or without obtaining the first input data. For example, the computing system may perform 604 and then perform 610 without performing 606 and 608.
[0123] At 612, the computing system may obtain second input data. The second input data (or selection data) may describe a second selection of an image. The second input data may describe audio input (e.g., a voice command), touch input (e.g., input to a touchscreen), keyboard input, and / or mouse input. The first input data may include a selection of a selection icon, a selection of a thumbnail, and / or a drop-and-drag selection.
[0124] At 614, the computing system may provide an image for display in place of the subset of text characters. For example, the computing system may delete the subset of text characters and add an image in place of the subset of text characters before the deletion.
[0125] In some implementations, the plurality of text characters may include a subset of the plurality of text characters and a second subset. The computing system may include processing the image and the second subset to determine a plurality of search results. In some implementations, the plurality of search results may be determined based on the image and the second subset. The plurality of search results may then be provided in a search results page interface.
[0126] 7 shows a flowchart diagram of an exemplary method performed in accordance with an exemplary embodiment of the present disclosure. While FIG. 7 shows steps performed in a particular order for purposes of illustration and discussion, the methods of the present disclosure are not limited to the particularly illustrated order or arrangement. Various steps of method 700 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.
[0127] At 702, a computing system may obtain a search query. The search query may include one or more words. In some implementations, obtaining the search query may include obtaining the search query via a query box of a search interface. The search interface may be provided by a web platform, a mobile application, and / or a desktop application. The search query may include Boolean terms, syntax, and / or natural language constructs.
[0128] At 704, the computing system may determine that one or more words comprise visual intent. The visual intent may be associated with one or more visual features. The visual intent may be based on one or more words associated with a color, pattern, design, object, and / or visual feature. The association may be based on one or more words that are visual descriptors, one or more words associated with a label for a particular visual feature, and / or one or more words associated with past image search queries. Words that describe a color, pattern, shape, and / or other visual descriptors may be determined to comprise visual intent.
[0129] At 706, the computing system can provide a user interface element. In some implementations, the user interface element can be descriptive of text replacement options. The user interface element can be an indicator that the systems and methods have determined that one or more words are associated with the visual intent. The user interface element can include a visual effect. The user interface element can include a pop-up element, a drop-down menu, a change in the display of one or more words, and / or an appearance of an icon.
[0130] At 708, the computing system may obtain first input data. The first input data may describe a first selection of a text replacement option. The first input data may include sensor data. The first input data may describe an interaction with a user interface element (e.g., a tap input, a gesture input, and / or a lack of input due to a threshold time elapsed without input being obtained).
[0131] At 710, the computing system may provide an image selection interface for display. The image selection interface may include multiple images for selection. In some implementations, the image selection interface may be provided for display based on determining one or more words that include visual intent. The image selection interface may include one or more different tabs for viewing and selecting images from different databases and / or different media or types. The image selection interface may include one or more panels for providing different types of media content items and / or media content items from different sources.
[0132] In some implementations, the computing system may provide the image selection interface without providing an indicator and / or without obtaining the first input data. For example, the computing system may perform 704 and then perform 710 without performing 706 and 708.
[0133] At 712, the computing system may obtain selection data. The selection data (e.g., second input data) may describe a second selection of an image. The selection data may include sensor data. The selection data may describe an interaction with the image selection interface (e.g., a tap input, a gesture input, and / or a lack of input due to a threshold time period elapsed without any input being obtained).
[0134] At 714, the computing system may provide an image for display in place of the one or more words. For example, a preview and / or thumbnail of the image may be provided for display in a query box of a search interface.
[0135] At 716, the computing system may determine one or more search results associated with the image. In some implementations, the one or more search results may be provided via a search results page. The search results page may include a query box that displays the image. Additionally and / or alternatively, the search results page may include a search results panel for displaying information associated with the one or more search results. The search query may include one or more additional words. In some implementations, the one or more search results may be determined based at least in part on the one or more additional words. The one or more search results may include one or more image search results. Additionally and / or alternatively, the one or more search results may include one or more product search results that describe products associated with one or more visual features of the image.
[0136] At 718, the computing system may provide one or more search results as output. The one or more search results may be provided for display in a search results page interface. The search results may be provided in different panels based on the type of search result, the source of the search results, and / or the classification of the search results.
[0137] 8 shows a flowchart diagram of an exemplary method performed in accordance with an exemplary embodiment of the present disclosure. While FIG. 8 shows steps performed in a particular order for purposes of illustration and discussion, the methods of the present disclosure are not limited to the particularly illustrated order or arrangement. Various steps of method 800 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.
[0138] At 802, the computing system can obtain a plurality of words. The plurality of words can include one or more specific words and one or more additional words. The one or more specific words can include visually descriptive terms. The one or more additional words can complement the one or more specific words and / or target different descriptive aspects of the search query or phrase.
[0139] At 804, the computing system may determine that one or more particular words of the plurality of words comprise visual intent. The visual intent may be associated with one or more visual features. The determination may be based on processing the plurality of words with one or more machine learning models to generate one or more outputs. The one or more machine learning models may include one or more detection models, one or more segmentation models, one or more classification models, and / or one or more expansion models. In some implementations, the one or more machine learning models may include one or more natural language processing models. The one or more machine learning models may include one or more transformer models. In some implementations, the determination may be based on historical search data.
[0140] At 806, the computing system can provide an indicator identifying one or more particular words to the plurality of words for display. The indicator can be a visual indicator describing one or more possible actions that can be performed based on the identified one or more particular words. The indicator can include an explanation, can include a change in text color, can include highlighting, and / or can include a pop-up element.
[0141] At 808, the computing system can determine a plurality of images associated with one or more particular words. Additionally and / or alternatively, the plurality of images can be associated with a visual intent. This determination can be based on querying a database using the one or more particular words. The database can be a local database stored on the user's device and / or a database accessed via a network connection. The one or more images can be cropped to isolate a particular portion of the image associated with the one or more particular words.
[0142] At 810, the computing system can provide a user interface panel with multiple images. The user interface panel can include multiple interactive user interface elements associated with the multiple images. The user interface panel can be a pop-up panel and / or can replace a portion of an initially displayed interface.
[0143] At 812, the computing system may obtain a selection of a particular image of the plurality of images. In some implementations, the particular image may be a cropped image from an image database. The cropped image may be generated by processing an uncropped image with one or more machine learning models to detect relevant portions of the image and segmenting the relevant portions from the uncropped image.
[0144] At 814, the computing system may provide a specific image for output that does not include the one or more additional words and the one or more specific words. The specific image may be positioned where the one or more specific words were previously displayed. In some implementations, a thumbnail and / or preview may be provided for display in place of the complete specific image.
[0145] In some implementations, the computing system can process the output to generate a translation output. The translation output can be generated based at least in part on the particular image.
[0146] Alternatively and / or additionally, the computing system may include providing the output to a search engine and receiving a plurality of search results, which may be associated with one or more additional words and a particular image.
[0147] Additional Disclosures The technology described herein refers to servers, databases, software applications, and other computer-based systems, as well as actions performed and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of configurations, combinations, and divisions of tasks and functions among components. For example, the processes described herein may be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications may be implemented on a single system or distributed across multiple systems. Distributed components may operate sequentially or in parallel.
[0148] While the present subject matter has been described in detail with reference to various specific exemplary embodiments thereof, each example is provided by way of illustration and not as a limitation of the present disclosure. Those skilled in the art, upon understanding the foregoing, will be able to readily create modifications, variations, and equivalents to such embodiments. Accordingly, the present disclosure is not intended to preclude inclusion of such modifications, variations, and / or additions to the present subject matter as would be readily apparent to one skilled in the art. For example, features illustrated or described as part of one embodiment can be used with another embodiment to yield yet another embodiment. Accordingly, the present disclosure is intended to cover such modifications, variations, and equivalents. [Explanation of symbols]
[0149] 10. Computing Devices 50 computing devices 100 Computing Systems 102 User Computing Devices 112 processors 114 memory 116 Data 118 Command 120 Visual Intention Decision Model 120 Machine Learning Models 122 User Input Components 130 Server Computing System 132 processors 134 memory 136 Data 138 Command 140 Visual Intention Decision Model 140 Machine Learning Models 150 Training Computing System 152 processors 154 memory 156 Data 158 Command 160 Model Trainer 162 training data 162 training datasets 180 Network 202 Search Interface 204 Query Input Box 206 One or more other words 208 One or more specific words 210 Voice Command Icon 212 User Profile 220 Image Selection Interface 222 Search query input box 224 search results page 226 Initial image selection page 228 Area Selection Interface 230 Cropping Interface 232 images 234 updated search results page 240 Image Selection Interface 242 indicator 244 Search Results Interface 246 options 248 Image Capture Interface 250 Cropping Options 252 Multimodal Queries 254 Updated search results interface 260 Image Selection Interface 262 indicator 264 search results interface 266 options 268 Image Search Interface 270 Area Selection Stage 272 Cropping Options 274 Multimodal Queries 276 updated search results page 300 Search Interface 302 Search Engines 304 Search Queries 306 Search Results Page 308 Indicator 310 first search result 312 second search result 314 third search result 316 nth search result 400 Image Selection Interface 410 User-specific Image Gallery Options 412 overlapping tile icons 414 Recent Screenshots Panel 416 full image panels 420 Image Capture Options 422 camera icon 424 Image Capture User Interface Element 430 Image Search Options 432 Globe Icon 434 Dedicated Search Query Box 500 Substitution System 502 Text Data 504 Historical Data 506 Indicator 508 Visual Intention Decision Model 510 Output Data 512 Text to Image Replacement Interface 514 Specific Images 516 Extended Data 518 Cropping Model 520 Cropped Images 600 ways 700 methods 800 ways
Claims
1. 1. A computer-implemented method for multimodal search, comprising: obtaining, by a computing system comprising one or more processors, a search query, the search query comprising one or more words; determining, by the computing system, that the one or more words have a visual intent, the visual intent being associated with one or more visual features; providing, by the computing system, an image selection interface for display, the image selection interface comprising a plurality of images for selection, the image selection interface provided for display based on the determination that the one or more words comprise the visual intent, the plurality of images obtained based at least in part on the one or more words comprising the visual intent; obtaining, by the computing system, selection data, the selection data describing a selection of images; providing, by the computing system, the image for display in place of the one or more words in a query box from which the search query was obtained; determining, by the computing system, one or more search results associated with the image; providing, by the computing system, the one or more search results as output; A computer-implemented method comprising:
2. providing, by the computing system, the image selection interface for display; providing, by the computing system, a user interface element, the user interface element describing text replacement options; obtaining, by the computing system, first input data, the first input data describing a first selection of the text replacement options; providing, by the computing system, the image selection interface for display based on the first input data; The method of claim 1 , comprising:
3. 10. The method of claim 1, wherein the one or more search results are provided via a search result page, the search result page comprising a query box that displays the image, and the search result page comprising a search result panel for displaying information associated with the one or more search results.
4. The method of claim 1 , wherein the search query comprises one or more additional words, and the one or more search results are determined based at least in part on the one or more additional words.
5. The method of claim 1 , wherein obtaining the search query comprises obtaining the search query via a query box of a search interface.
6. The method of claim 1 , wherein the one or more search results comprise one or more image search results.
7. The method of claim 1 , wherein the one or more search results comprise one or more product search results describing products associated with the one or more visual features of the image.
8. 1. A computing system for text-to-image substitution, comprising: one or more processors; one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations; and wherein the operation comprises: obtaining text data, the text data describing a plurality of text characters; processing the text data to determine whether a subset of the plurality of text characters comprises a visually descriptive term, the visually descriptive term being associated with one or more visual features; providing an image selection interface for display, the image selection interface comprising a plurality of images for selection, the plurality of images obtained based at least in part on the visually descriptive terms; obtaining selection data, the selection data describing a selection of images; providing the image for display in place of the subset of the plurality of text characters; and A computing system comprising:
9. providing the image selection interface for display; providing a display indicator, the indicator describing text replacement options for replacing the visually descriptive term with image data; obtaining first input data, the first input data describing a first selection of the text replacement options; providing the image selection interface for display based on the first input data; The system of claim 8, comprising:
10. The system of claim 9 , wherein the indicator comprises the subset of the plurality of text characters displayed in one or more colors different from the remaining characters of the plurality of text characters.
11. the plurality of text characters comprises the subset and a second subset of the plurality of text characters; The operation is processing the image and the second subset to determine a plurality of search results, the plurality of search results being determined based on the image and the second subset; providing the plurality of search results in a search result page interface; The system of claim 8 further comprising:
12. the plurality of images submitting a query to a search engine using the subset of the plurality of text characters; receiving the plurality of images; The system of claim 8, wherein the system is obtained by:
13. 9. The system of claim 8, wherein the plurality of images is obtained by determining that image data in a user-specific image database is associated with one or more visual features, and the image data associated with the one or more visual features comprises the plurality of images.
14. providing the image selection interface for display; 10. The system of claim 8, comprising providing an image search option, a user image database option, and an image capture option, wherein the image search option comprises querying a network of computing systems using the subset of the plurality of text characters, the user image database option comprises retrieving an image from a user image database, and the image capture option comprises utilizing one or more image sensors of a user device.
15. The system of claim 8 , wherein the visually descriptive terms are determined based on historical search data.
16. The system of claim 15 , wherein the historical search data describes a plurality of terms previously utilized to retrieve one or more image search results.
17. The system of claim 8 , wherein the visually descriptive terms are determined based on processing the text data with a semantic understanding model.
18. One or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, the operations including: obtaining a plurality of words, the plurality of words comprising one or more particular words and one or more additional words; determining that the one or more particular words of the plurality of words have a visual intent, the visual intent being associated with one or more visual features; providing an indicator identifying the one or more particular words to the plurality of words for display; determining a plurality of images associated with the one or more particular words, the plurality of images being associated with the visual intent; and providing the plurality of images on a user interface panel, the user interface panel comprising a plurality of interactive user interface elements associated with the plurality of images; obtaining a selection of a particular image from the plurality of images; providing the particular image for output without the one or more additional words and the one or more particular words; 1. One or more non-transitory computer-readable media comprising:
19. The operation is 20. The one or more non-transitory computer-readable media of claim 18, further comprising processing the output to generate a translated output, the translated output being generated based at least in part on the particular image.
20. The operation is providing the output to a search engine; receiving a plurality of search results, the plurality of search results being associated with the one or more additional words and the particular image; and 20. The one or more non-transitory computer-readable media of claim 18, further comprising:
Citation Information
Patent Citations
Text and image-based search
JP2020534597A
Content recommendation based on color match
US10083521B1
Automatic color palette based recommendations for affiliated colors
US20180158128A1
Image search and retrieval using object attributes
US20190121879A1