Intelligent system and method for visual search queries

By constructing a user-centric visual interest map and analyzing users' historical image data, the problem of existing visual query systems failing to accurately reflect user intent is solved, resulting in more intelligent and personalized search results and providing more accurate visual query processing.

CN114329069BActive Publication Date: 2026-04-07GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-22
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing visual query systems are unable to intelligently process visual queries, resulting in an inaccurate reflection of the user's true search intent, and the results often focus on items with similar visual features while ignoring the user's personalized needs.

Method used

By constructing a user-centric visual interest map, using a computing system to analyze users' historical image data, generating personalized search result rankings, and processing visual queries based on contextual signals and user interest data, identifying multiple candidate search results, and providing personalized visual result notifications.

Benefits of technology

It achieves smarter and more personalized search results, accurately understands user intent, reduces user interface clutter, provides content related to multiple canonical items, eliminates ambiguity in object-specific and category queries, and returns a content set of composite entities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114329069B_ABST
    Figure CN114329069B_ABST
Patent Text Reader

Abstract

A user can submit a visual query that includes one or more images. Within or in connection with the visual query, various processing techniques, such as optical character recognition (OCR) techniques, can be used to recognize text (e.g., in the image(s), around the image(s), etc.) and / or various object detection techniques (e.g., machine learning object detection models, etc.) can be used to detect objects (e.g., products, landmarks, animals, humans, etc.). Content related to the detected text or object(s) can be identified and potentially provided to the user as search results or proactive content feeds. Thus, various aspects of the present disclosure enable visual search systems to more intelligently process visual queries to provide improved search results and content feeds, including those that are more personalized and / or take into account implicit features of the visual query and / or user search intent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to systems and methods for processing visual search queries. More specifically, this disclosure relates to a computer vision search system capable of detecting and recognizing objects in an image included in a visual query, and providing more personalized and / or intelligent search results. Background Technology

[0002] Text-based or term-based search involves users entering words or phrases into a search engine and receiving various results. Term-based queries require users to explicitly provide search terms in the form of words, phrases, and / or other terms. Therefore, term-based queries are inherently limited by text-based input patterns, and users cannot search based on the visual features of an image.

[0003] Alternatively, a visual query search system can provide search results to a user in response to a visual query that includes one or more images. Computer vision analysis techniques can be used to detect and identify objects in images. For example, optical character recognition (OCR) technology can be used to identify text in images, and edge detection techniques or other object detection techniques (e.g., machine learning-based methods) can be used to detect objects in images (e.g., products, landmarks, animals, etc.). Content related to the detected objects can be provided to the user (e.g., images captured of the detected objects or images submitted in other ways or images associated with the visual query).

[0004] However, some existing visual query systems have several drawbacks. As an example, current visual search query systems and methods may provide users with results that may only relate to the visual query's explicit visual characteristics, such as color schemes, shapes, or depictions of items / objects identical to the image(s) depicting the visual query. In other words, some existing visual query systems focus specifically on identifying other images containing visual features similar to the query image, which may not reflect the user's true search intent.

[0005] Therefore, there is a need for a system that can process visual queries more intelligently to provide users with improved search results. Summary of the Invention

[0006] Aspects and advantages of embodiments of this disclosure will be set forth in part in the description which follows, or may be learned from the description or by practice of the embodiments.

[0007] One example aspect of this disclosure relates to a computer-implemented method for providing personalized visual search query result notifications within a user interface overlaid on an image. The method includes: obtaining a visual search query associated with a user via a computing system including one or more computing devices, wherein the visual search query includes an image. The method includes: identifying a plurality of candidate search results for the visual search query via the computing system, wherein each candidate search result is associated with a specific sub-part of the image, and wherein a plurality of candidate visual result notifications are associated with the plurality of candidate search results respectively. The method includes: accessing user-specific user interest data and a description of the user's visual interests associated with the user via the computing system. The method includes: generating a ranking of the plurality of candidate search results via the computing system based at least in part on a comparison of the plurality of candidate search results with the user-specific user interest data associated with the user. The method includes: selecting at least one of the plurality of candidate search results as at least one selected search result via the computing system based at least in part on the ranking. The method includes: providing at least one selected visual result notification, respectively associated with at least one selected search result, overlaid on a specific sub-part of the image associated with the selected search result via the computing system.

[0008] Another exemplary aspect of this disclosure relates to a computing system that returns content for a plurality of canonical items in response to a visual search query. The computing system includes: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations include: obtaining a visual search query, wherein the visual search query includes an image depicting an object. The operations include: accessing a plurality of graphs describing a plurality of distinct items, wherein a corresponding set of content is associated with each of the plurality of distinct items. The operations include: selecting a plurality of selected items from the graphs depicting objects in the image based on the visual search query. The operations include: returning a combined set of content as a search result in response to the visual search query, wherein the combined set of content includes at least a portion of the corresponding set of content associated with each of the plurality of selected items.

[0009] Another exemplary aspect of this disclosure relates to a computational system for disambiguating between object-specific and categorical visual queries. The computational system includes: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the computational system to perform operations. The operations include: obtaining a visual search query, wherein the visual search query includes an image depicting one or more objects. The operations include: identifying one or more constituent characteristics of the image included in the visual search query. The operations include: determining, at least in part, based on one or more constituent characteristics of the image included in the visual search query, whether the visual search query includes an object-specific query specifically belonging to one or more objects identified in the image included in the visual search query, or determining whether the visual search query includes a categorical query belonging to a general category of one or more objects identified in the image included in the visual search query. The operations include: when it is determined that the visual search query includes an object-specific query, returning one or more object-specific search results specifically belonging to one or more objects identified in the image included in the visual search query. The operations include: when it is determined that the visual search query includes a categorical query, returning one or more categorical search results belonging to a general category of one or more objects identified in the image included in the visual search query.

[0010] Another exemplary aspect of this disclosure relates to a computer-implemented method for returning content of multiple composite entities to a visual search query. The method includes: obtaining a visual search query, wherein the visual search query includes an image depicting a first entity. The method includes: identifying one or more additional entities associated with the visual search query, at least in part based on one or more contextual signals. The method includes: determining a combined query of content associated with a combination of the first entity and one or more additional entities. The method includes: returning a set of content in response to the visual search query, wherein the set of content includes at least one content item in response to the combined query and associated with a combination of the first entity and one or more additional entities.

[0011] Other aspects of this disclosure relate to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.

[0012] These and other features, aspects, and advantages of the various embodiments of this disclosure will be better understood by referring to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the disclosure and, together with the description, serve to explain the relevant principles. Attached Figure Description

[0013] The embodiments for those skilled in the art are discussed in detail in the specification with reference to the accompanying drawings, in which:

[0014] Figure 1 A block diagram of an example computing system according to an example embodiment of the present disclosure is described.

[0015] Figure 2 A block diagram of an example visual search system including a query processing system according to an example embodiment of the present disclosure is described.

[0016] Figure 3 A block diagram of an example visual search system including a query processing system and a ranking system according to an example embodiment of the present disclosure is described.

[0017] Figure 4 A block diagram of an example visual search system and contextual components according to an example embodiment of this disclosure is described.

[0018] Figure 5 Screenshots of client systems with contrasting augmented reality visual effects according to some embodiments are shown to illustrate the differences.

[0019] Figure 6 A screenshot of a client system, according to some embodiments, showing comparative example search results based on an object of interest, is illustrated to show the differences.

[0020] Figure 7 A client system with screenshots of example search results is shown according to some embodiments.

[0021] Figure 8 A screenshot of a client system with example search results based on multiple objects of interest, according to some embodiments, is shown to illustrate the differences.

[0022] Figure 9 A flowchart is described as an exemplary method for performing more personalized and / or intelligent visual search using a user-centric visual interest model, according to an exemplary embodiment of the present disclosure.

[0023] Figure 10 A flowchart is described as an exemplary method for performing a more personalized and / or intelligent visual search through a visual search query using multiple canonical items, according to an exemplary embodiment of the present disclosure.

[0024] Figure 11 A flowchart is described as an example method for performing more personalized and / or intelligent visual search by deambiguity between specific objects and categorized visual queries, according to an example embodiment of the present disclosure.

[0025] Figure 12 A flowchart is described as an example method for performing more personalized and / or intelligent visual search by combining multiple objects of interest into a search query, according to an example embodiment of the present disclosure.

[0026] The repeated reference numerals in multiple figures are intended to identify the same features in various implementations. Detailed Implementation

[0027] Overview

[0028] Generally, this disclosure relates to a computer-implemented visual search system that can detect and identify objects in or related to a visual query, and then provide more personalized and / or intelligent search results in response to the visual query (e.g., in enhanced visual query coverage). For example, a user may submit a visual query including one or more images. In or related to the visual query, various processing techniques (such as optical character recognition (OCR) techniques) may be used to recognize text (e.g., in images, surrounding images, etc.) and / or various object detection techniques (e.g., machine learning object detection models, etc.) may be used to detect objects (e.g., products, landmarks, animals, humans, etc.). Content related to the detected text or objects (or more) can be identified and provided to the user as search results. Therefore, aspects of this disclosure enable visual search systems to process visual queries more intelligently to provide improved search results, including more personalized and / or contextual signals to interpret implicit features of the visual query and / or the user's search intent.

[0029] The example aspects of this disclosure provide smarter search results in response to visual queries. A visual query may include one or more images. For example, the images included in a visual query may be images captured simultaneously or previously existing images. In one example, a visual query may include a single image. In another example, a visual query may include ten image frames from approximately three seconds of video capture. In yet another example, a visual query may include an image library of images, such as all images included in a user's photo library. For example, such a library may include images of zoo animals recently captured by the user, images of cats captured by the user not long ago (e.g., two months ago), and images of tigers saved by the user from existing sources (e.g., from a website or screen capture) to the library. These images may represent a set of high-affinity images for the user and embody (e.g., through graphics) an abstract idea that the user may have a "visual interest" in similar animals. Any given user may have many such clusters of nodes, each cluster representing an interest that cannot be well captured by words.

[0030] According to one example aspect, a visual search system can construct and utilize a user-centric visual interest graph to provide more personalized search results. In one example use, the visual search system can use the user interest graph to filter visual discovery notifications, announcements, or other opportunities. Therefore, in an exemplary embodiment where search results are presented as visual result notifications (e.g., which may be referred to in some cases as "gleams") within an enhanced overlay of the query image, personalization of search results based on user interests may be particularly advantageous.

[0031] More specifically, in some implementations, the visual search system may include or provide an enhanced overlay user interface for providing visual result notifications for search results as overlays of images included in the visual query. For example, the visual result notification may be provided at a location corresponding to the image portion associated with the search result (e.g., the visual result notification may be displayed "top" of the object associated with the corresponding search result). Thus, in response to a visual search query, multiple candidate search results can be identified, and multiple candidate visual result notifications may be associated with multiple candidate search results separately. However, in cases where the underlying visual search system is powerful and extensive, a large number of candidate visual result notifications may be available, making the presentation of all candidate visual result notifications result in a cluttered and congested user interface or otherwise undesirably obscure the underlying image. Therefore, according to one aspect of this disclosure, the computer visual search system may construct and utilize a user-centric visual interest graph to rank, select, and / or filter candidate visual result notifications based on observed user visual interests, thereby providing a more intuitive and simplified user experience.

[0032] In some implementations, user-specific interest data (e.g., represented using graphs) can be aggregated, at least in part, by analyzing images that a user has engaged with over time. In other words, a computational system can attempt to understand a user's visual interests by analyzing images that the user has engaged with over time. When a user engages with an image, it can be inferred that certain aspects of the image are of interest to the user. Therefore, items (e.g., objects, entities, concepts, products, etc.) that are included in or associated with such images can be added to or otherwise labeled in the user-specific interest data (e.g., graphs).

[0033] As an example, user-engaged images can include photos captured by the user, screenshots captured by the user, or images included in web-based or app-based content viewed by the user. In another potentially overlapping example, user-engaged images can include actively engaged images where the user actively participates by requesting an action to be performed on the image. For example, the requested action could include performing a visual search relative to the image, or the user explicitly marking the image as including aspects of their visual interest. As another example, user-engaged images can include passively observed images presented to the user without the user's specific participation. Visual interests can also be inferred from text content entered by the user (e.g., text- or term-based queries).

[0034] In some implementations, the user interests reflected in the user-centric visual interest map may be precise, categorical, or abstract items (e.g., entities). For example, a precise interest may correspond to a specific item (e.g., a particular artwork), a categorical interest may be a category of an item (e.g., Art Nouveau painting), and an abstract interest may correspond to interests that are difficult to capture through categorization or text (e.g., "artworks that are visually similar to Gustav Klimt's 'The Kiss'").

[0035] Interest can be indicated in a user-centric visual interest map by overlaying or defining (e.g., and then updating periodically) variable weighted levels of interest on exact, categorical, or abstract items of interest. By evaluating the variable weighted levels of interest for items and / or related items (e.g., items connected to items in the graph or items within n hops of items in the graph), a user's interest in a variety of items can be inferred or determined.

[0036] In some implementations, the variable weighted interest bias assigned to the identified visual interests decays over time, such that user-specific interest data is at least partially based on the time frame of the expressed interest. For example, a user might express strong interest in a particular topic for a period of time, and then express no interest at all afterwards (e.g., a user might express strong interest in a particular frequency band for a year). The decay of user interest data over time reflects changes in user interests, and if sufficient decay occurs, the visual search system may be unable to continue showing the user query results related to topics that the user is no longer interested in.

[0037] Therefore, in some implementations, a visual interest graph can be a collection of descriptions of user interests / personalizations (e.g., based on many images they have seen historically). This graph can be constructed by analyzing the properties of many historical images and using this information to find other images of interest (e.g., and often associated content such as a collection of news articles).

[0038] Visual search systems can use user-specific interest data, at least in part, to generate rankings of multiple candidate search results based on comparisons of multiple candidate search results and user-specific interest data associated with the user. For example, item weights can be applied to modify or reweight the initial search scores associated with candidate search results.

[0039] A search system can select at least one of multiple candidate search results as at least one selected search result, at least partially based on ranking, and then provide at least one selected visual result notification, each associated with the at least one selected search result, overlaid on a sub-section of the image associated with the selected search result. In this way, user interests can be used to provide personalized search results and reduce clutter in the user interface.

[0040] In another example, a user-centric visual interest graph can be used to curate user-specific feeds based on visual information and interests. Specifically, in a feed, a personalized set of content can be displayed to the user without being based on a single, specific image. Instead, analysis of a previous set of images can build a graph as described above, for example, with images (and / or image metadata, such as the entities depicted) as nodes, and then the connections between these nodes as edges to determine the intensity of interest. A new visual media item (e.g., an image or video) can also have an “advantage” associated with the user’s interests in a given graph. This new media can then be proactively recommended to the user without prior querying (e.g., as part of a feed provided to the user in certain contexts, such as opening a new tab in a browser application).

[0041] In addition to the above descriptions, users may be given control over whether and when the systems, programs, or features described herein may collect user information (e.g., information about the user's social networks, social behaviors, activities, occupation, user preferences, or the user's current location), and whether content or communications may be sent to the user from the server. Furthermore, some data may be processed in one or more ways before storage or use to remove personally identifiable information. For example, a user's identity may be processed so that personally identifiable information about the user cannot be determined, or a geographic location may be generalized to the location where location information is obtained (such as city, postal code, or state), thus making it impossible to determine the user's specific location. Therefore, users have control over what information about themselves is collected, how that information is used, and what information is provided to them.

[0042] On the other hand, computer-implemented visual search systems can identify and return combined sets of content (e.g., user-generated content) for multiple canonical items in response to a visual search query. In particular, understanding the granularity and object of user intent is a challenging task because visual search queries support more expressive and fluid search input patterns. For example, suppose a user has just watched a particular movie. The user can submit a large amount of visual content as a visual query to reflect their interest in receiving information about the movie. This content could be end credits, the movie's physical media (e.g., a disc), packaging covers, movie receipts, or trailers reminding the user of the movie the next day. Therefore, mapping an image of world value to a specific item is a challenging problem. Conversely, understanding the intended granularity of a user's query is challenging. For example, a visual query that includes an image depicting the cover of a particular movie might be intended to search for content about that particular movie, content about the actors in the movie, content about the director of the movie, content about the artist who made the cover, content about the same genre as the movie (e.g., horror movies), or even more specific content, such as content specifically about a particular version of the movie (e.g., the 2020 “Director’s Cut” version versus all versions of the movie; DVD version versus Blu-ray version; etc.).

[0043] This disclosure addresses these challenges by enabling the return of a combined set of content for multiple canonical items in response to a visual search query. Specifically, in response to a visual search query that includes an image describing an object, the visual search system can access a graph describing multiple distinct items, where a corresponding set of content (e.g., user-generated content such as product reviews) is associated with each of the multiple distinct items. The visual search system can select multiple chosen items from the graph representing the objects depicted in the image, based on the visual search query, and then return a combined set of content as the search result, where the combined set of content includes at least a portion of the corresponding set of content associated with each chosen item. By returning content related to multiple canonical items, the visual search system avoids providing results that are overly specific to a particular entity that might be identified in the visual query. Continuing with the example above, while some existing systems might only return content related to the 2020 “Director’s Cut” movie, the proposed system might return content related to the 2020 “Director’s Cut” version, or content related to other relevant entities, such as content related to the actors in the movie, content related to the movie director, content related to the artist who generated the packaging cover, etc.

[0044] Various techniques can be used to select items from a graph. In one example, “aesthetic-assisted” visual searches (e.g., searches looking for information about specific items rather than abstract aesthetic characteristics) can be advantageously handled, where the graph can be a hierarchical representation of multiple distinct items. Selecting multiple items from a graph can include identifying a principal item in the graph corresponding to an object depicted in an image (e.g., a specific film shown in an image) based on a visual search query. Next, the visual search system can identify one or more supplementary items in the graph’s hierarchical representation that are associated with the principal item, and select the principal item and one or more supplementary items as multiple selections. Supplementary items can be at the same level (e.g., other films directed by the same director), possibly at a “higher” level (e.g., other films of the same genre), and / or possibly at a “lower” level (e.g., the 2020 “director’s cut” version of a film and the original 1990 TV series version).

[0045] In another example, where “aesthetic-primary” visual searches can be advantageously handled (e.g., searches seeking content related to abstract visual or aesthetic features rather than specific canonical items), a graph describing multiple distinct items can include multiple nodes corresponding to multiple indexed images. Multiple nodes can be arranged within the graph based at least in part on the visual similarity between indexed images, such that the distance between a pair of nodes within the graph is negatively correlated with the visual similarity between the corresponding pair of indexed images (i.e., nodes of more similar images are “closer” to each other in the graph). In one example, multiple nodes of the graph can be arranged into multiple clusters, and selecting multiple selections from the graph based on a visual search query can include performing an edge thresholding algorithm to identify the dominant cluster among the multiple clusters and selecting the nodes included in the dominant cluster as multiple selections. In another example, the visual search system can perform an edge thresholding algorithm to directly identify multiple visually similar nodes that are visually similar to the objects depicted in the images (e.g., the opposite of identifying clusters). The visual search system can then select visually similar nodes as multiple selections. Examples of “edges” or “sizes” that a product image search might match include identifying derived attributes such as category (e.g., “dress”), attributes (e.g., “sleeveless”), or other semantic dimensions and / or visual attributes such as “dark with light”, “light is the thin line that makes up 40% of the entire color space”, etc., including machine-generated visual attributes such as machine-extracted visual features or machine-generated visual embeddings.

[0046] Therefore, some example techniques are provided to enable visual search systems to process visual queries more intelligently and return content related to multiple canonical items, rather than “overfitting” the user’s query and returning only content about a single specific item that may not be the intended focus of the user’s query.

[0047] On the other hand, computer-implemented visual search systems can intelligently dissolve ambiguities in visual queries between a specific object depicted in a query image and the classification result. In particular, dissolving ambiguities between object-specific and classification queries is another example of the challenges of understanding the granularity of user intent and object relevance. For example, suppose a user submits an image of a cereal box and a query requesting "which has the most fiber" (e.g., a text or voice query). It is difficult to determine whether the user's intent is to identify the cereal with the highest fiber content from those specifically depicted in the visual query, or whether the user's intent is to identify the cereal that generally has the highest fiber content. Significant differences between the submitted image and the query can lead to the same result being returned to the user, further emphasizing the difficulty of the task.

[0048] This disclosure addresses these challenges by determining whether a visual search query includes an object-specific query or a category query and then returning content that is inherently object-specific or category-specific. Specifically, computer vision search systems can use additional contextual signals or information to provide smarter search results that interpret relationships between multiple distinct objects present in a visual query. For example, a visual search system can identify one or more compositional features of an image contained within a visual query. The visual search system can use these compositional features to predict whether a visual query is a category query relevant to an extended corpus of search results or an object-specific query specifically belonging to one or more objects identified within the visual query. Continuing with the example above, a visual search system can use combined features of an image describing grains to determine whether a visual query image belongs to all grains, all grains of a particular brand or type, or only to those grains contained in the image.

[0049] When a visual search query is determined to include an object-specific query, the visual search system may return one or more object-specific search results that specifically point to one or more objects identified in the visual query (e.g., the grain with the highest fiber content among the grains captured in the image). Optionally, when a visual search query is determined to include a category query, the visual search system may return one or more category search results that point to the general category of one or more objects identified in the visual query (e.g., the grain with the highest fiber content among all grains, or, as another example, the grain with the highest fiber content among all grains of the same type or brand).

[0050] The compositional characteristics used by visual search systems can include a variety of attributes of an image. In one example, the compositional characteristics of an image can include distances to one or more objects identified in the image (e.g., from the camera that took the photo). For example, an image that includes objects closer to the camera is more likely to be specific to the object being depicted, while an image that includes objects farther from the camera is inherently more likely to be categorized. For instance, a user looking for information about a specific grain might stand near a specific grain box and capture a visual query along the entire grain aisle.

[0051] In another example, the compositional characteristics of an image may include the number of one or more objects identified in the image. In particular, an image with a larger number of identified objects may be more likely to indicate that a visual query points to a categorical query, while an image with a smaller number of identified objects may be more likely to indicate that a visual query points to an object-specific query (e.g., an image with 3 cereal boxes is more likely to indicate that a visual query points to one or more objects identified in the visual query, while an image with 25 cereal boxes may be more likely to indicate that a visual query points to a general category of one or more objects identified in the visual query).

[0052] In another example, the compositional characteristics of an image can include the relative similarity between one or more objects identified in the image. Specifically, an image containing multiple objects and exhibiting high similarity to other objects in the image is more likely to indicate a categorical query, whereas an image containing multiple objects and exhibiting low similarity to other objects is more likely to indicate an object-specific query. For example, an image containing a cereal box and a bowl is more likely to indicate that the visual query specifically targets one or more objects identified in the visual query, whereas an image containing multiple cereal boxes is more likely to indicate that the visual query targets a general category of one or more objects identified in the visual query.

[0053] As another example, the compositional characteristics of an image can include the angular orientation of one or more objects in the image. In particular, an image containing objects with arbitrary angular orientations may be more likely to indicate that a visual query points to a category query, while an image containing objects with specific angular orientations may be more likely to indicate that a visual query points to a specific object query. For example, an image of a cereal box at a 32-degree angle to the image edge may be more likely to indicate that a visual query points to a category query, while an image containing a cereal box at a 90-degree angle to the image edge (e.g., clearly showing one side of the box facing the camera) may be more likely to indicate that a visual query specifically points to one or more objects identified in the visual query.

[0054] As another example, compositional characteristics of an image can include the centering of one or more objects in the image (i.e., the degree to which an object is centered in the image). Specifically, images containing non-centered objects or objects not within a centering threshold are more likely to indicate the likelihood that a visual query points to a categorical query, whereas images containing centered objects or objects within a centering threshold are more likely to indicate the likelihood that a visual query points to an object-specific query. Furthermore, the degree of centering of objects in a visual query can be obtained using a measurement ratio from the image edge to the identified object (e.g., an image of a cereal box with a 1:6:9:3 ratio is more likely to indicate that a visual query points to a categorical query, whereas an image containing a cereal box with a 1:1:1:1 ratio is more likely to indicate that a visual query points to one or more objects identified in the visual query).

[0055] In some embodiments, contextual signals or information used to provide smarter search results (interpreting the relationships between multiple different objects present in a visual query) may include the user's location at the time of the visual search query. Specifically, certain locations where the user is conducting a visual search query (e.g., a grocery store) may be more likely to indicate that the visual query points to a category query, while other locations (e.g., a private residence) may be more likely to indicate that the visual query points to an object-specific query. If multiple options are available to the user, the location may be more likely to indicate a category visual query, while a location with limited options may be more likely to indicate an object-specific visual query.

[0056] In some embodiments, the visual search query may be further determined based on filters associated with it, indicating whether the visual search query includes an object-specific query or a category query. Specifically, the filter may incorporate user history as information as to whether the user is more likely to perform a category query or an object-specific query on one or more objects in the image that includes the visual query. Optionally or additionally, the filter may include textual or verbal queries that the user entered in connection with the visual query. For example, a verbal query “Which grain is the healthiest?” is more likely to be straightforward, while a verbal question “Which of these three is the healthiest?” is more likely to be object-specific.

[0057] In some embodiments, the visual search system may return one or more categorized search results pointing to a general category of one or more objects identified in the visual query. Specifically, returning one or more categorized search results pointing to a general category of one or more objects identified in the visual query may include first generating a set of discrete category objects (e.g., category of an item, brand within an item's category, etc.) upon which at least one category of one or more objects in the image is based. Furthermore, the visual search system may then select multiple selected discrete object categories from the set. More specifically, the visual search system may use at least one contextual signal or information to determine which of the discrete category categories to select. Finally, the visual search system may return a combined content set as search results, wherein the combined content set includes results associated with each of the multiple selected discrete object categories. Specifically, the visual search system may return multiple results to the user, wherein the results may be displayed in a maximum likelihood hierarchy.

[0058] Therefore, some example techniques are provided to enable visual search systems to process visual queries more intelligently and return content related to categorized content or object-specific content based on contextual information provided in the visual query provided by the user.

[0059] From another perspective, computer-implemented visual search systems can return the content of multiple combined entities to a visual search query. In particular, understanding when a user is looking for information about a specific combination of multiple entities in a visual search query is another example of how understanding the granularity and object of user intent is a challenging task. For example, suppose a user submits or otherwise selects an image of Emma Watson and Daniel Radcliffe at the Oscars to include in a visual search query. It is difficult to determine whether the user's intent is to query for Emma Watson, Daniel Radcliffe, Oscars, Emma Watson at the Oscars, Harry Potter, or various other combinations of entities. Significant differences between the submitted image and the query could lead to the same result being returned to the user, further highlighting the difficulty of the task. Some existing systems simply cannot interpret any combination of multiple entities and instead will simply return the visually most similar image (e.g., at the pixel level, such as similar background colors) to such a visual query.

[0060] In contrast, this disclosure addresses these challenges by enabling visual search systems to return content to users based on identifying the composition of multiple entities and querying for such entity composition. Specifically, the computer visual search system can identify one or more entities associated with a visual search query based on one or more contextual signals or information. After identifying the multiple entities associated with the visual search query, the visual search system can determine a combined query for content related to a combination of a first entity and one or more additional entities (e.g., “2011 Harry Potter Oscars”). Entities can include people, objects, and / or abstract entities such as events. After determining the combined query for content related to the combination of the first entity and one or more additional entities, the visual search system can obtain and return a set of content including at least one content item that responds to the combined query and is associated with the combination of the first entity and one or more additional entities. Continuing with the example given above, in response to an image of Emma Watson and Daniel Radcliffe at the Oscars, the visual search system can construct a combined query and return search results for the nominations and awards received by the Harry Potter cast at the 2011 Academy Awards.

[0061] Contextual signals or information used to determine whether a visual query involves a combination of multiple entities can include various attributes of the image, information about where the user obtained the image from, information about other uses or instances of the image, and / or various other contextual information. In one example, the image used in the visual search query appears in a web document (e.g., a webpage). More specifically, the web document may reference entities in one or more sections. In particular, these references can be text (e.g., “2011 Oscar designer”) or images (e.g., photos from the 2011 Oscars red carpet), and these entities can be identified as additional entities associated with the visual search (e.g., “Emma Watson, 2011 Oscar costume designer”). Therefore, if a user selects an image of Emma Watson contained in a webpage as their visual query submission, referencing other entities (e.g., textual and / or visual references) can be used to identify potential other entities that can be used to form a combination of multiple entities.

[0062] As another example, the contextual signals or information of an image may include supplementary web documents that include additional instances of the image associated with the visual search query (e.g., multiple articles discussing how Harry Potter swept the Oscars). Specifically, one or more supplementary entities referenced by one or more supplementary web documents can be identified as supplementary entities associated with the visual search (e.g., "Harry Potter 2011 Oscars"). Therefore, if a user selects a first instance of an Emma Watson image contained in a first webpage as the visual query submission, other instances of that image can be identified in other different webpages (e.g., by performing a typical reverse image search), and references to other entities contained in such other different webpages (e.g., textual and / or visual references) can then be used to identify potential other entities that can be used to form a combination of multiple entities.

[0063] As another example, contextual signals or information about an image can include textual metadata (e.g., "Emma Watson and Daniel Radcliffe after winning the Oscars"). Specifically, textual metadata can be accessed and identified as additional entities relevant to the visual search (e.g., "Harry Potter 2011 Oscars"). In particular, textual metadata can include the title of the image used in the user-submitted visual query.

[0064] As another example, contextual signals or information about an image may include location or time metadata (e.g., the Kodak Theatre in Los Angeles). Specifically, location or time metadata can be accessed and identified as an additional entity associated with a visual search, where the location of the image source used in a visual search query may indicate a relevant thematic reference that may not be pointed out elsewhere in the image itself (e.g., an image of Emma Watson and Daniel Radcliffe might be a generic image on a red carpet with no symbolic meaning behind it indicating that they were at the Oscars, and it triggers search queries such as "Emma Watson and Daniel Radcliffe at the Kodak Theatre in Los Angeles, 2011").

[0065] As another example, contextual signals or information from an image can include an initial search. More specifically, an initial search can be performed using multiple identified entities (e.g., “Emma Watson and Daniel Radcliffe at the Kodak Theatre in Los Angeles”), and after obtaining the first set of initial search results, other entities referenced by the initial search results can be identified. In particular, entities identified in a number exceeding a threshold of initial results can be determined to be sufficiently relevant to be included in subsequent queries (e.g., “Emma Watson and Daniel Radcliffe at the Academy Awards”).

[0066] Therefore, some example techniques are provided to enable visual search systems to process visual queries more intelligently and return content related to multiple synthetic entities based on contextual signals or information in the visual query provided by the user.

[0067] The identification of relevant images or other content can be performed in response to an explicit search query, or it can be performed proactively in response to a general query for user content (e.g., as part of a feed, such as a "discovery feed" that proactively identifies and provides content to the user without an explicit query). The term "search results" is intended to include content identified in response to a specific visual query and / or proactively identified as proactive results included in a feed or other content moderation mechanism. For example, a feed may include content based on a user's visual interest map without requiring a specific initial intent statement from the user.

[0068] Exemplary embodiments of this disclosure will now be discussed in further detail with reference to the accompanying drawings.

[0069] Example devices and systems

[0070] Figure 1 A block diagram of an exemplary computing system 100 according to an exemplary embodiment of the present disclosure is depicted, which performs personalized and / or intelligent search in response to at least a portion of a visual query. System 100 includes a user computing device 102 and a visual search system 104 communicatively coupled via a network 180.

[0071] User computing device 102 can be any type of computing device, such as a personal computing device (e.g., a laptop or desktop computer), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0072] User computing device 102 includes one or more processors 112 and memory 114. The one or more processors 112 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. Memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 114 can store data 116 and instructions 118 executed by processor 112 to cause user computing device 102 to perform operations.

[0073] In some implementations, the camera application 126 of the user computing device 102 presents content related to objects identified in the viewfinder of the camera 124 of the user computing device 102.

[0074] Camera application 126 may be a native application developed for a specific platform. Camera application 126 can control camera 124 of user computing device 102. For example, camera application 126 may be a dedicated application for controlling the camera, i.e., a camera-first application for controlling camera 124 used in conjunction with other functions of the application or other types of applications that can access and control camera 124. Camera application 126 may present the viewfinder of camera 124 in the user interface 158 of camera application 126.

[0075] Typically, camera application 126 allows a user to view search content (e.g., information or user experience) related to objects depicted in the viewfinder of camera 124, and / or view content related to objects depicted in images stored on user computing device 102 or stored in another location accessible from user computing device 102. The viewfinder is part of the display of user computing device 102 and presents a live image of the field of view of the camera lens. When the user moves camera 124 (e.g., by moving user computing device 102), the viewfinder is updated to present the current field of view of the lens.

[0076] Camera application 126 includes object detector 128, user interface generator 130, and on-device tracker 132. Object detector 128 can detect objects in the viewfinder using edge detection and / or other object detection techniques. In some implementations, object detector 128 includes a coarse classifier for determining whether the image includes objects of one or more specific classes (e.g., categories). For example, the coarse classifier can detect that the image includes objects of a specific class, regardless of whether actual objects are identified.

[0077] A coarse classifier can detect the presence of a class of objects based on whether an image includes (e.g., a description) one or more features that indicate the class of objects. A coarse classifier can include a lightweight model for performing low-computational analysis to detect the presence of objects within its object class. For example, for each object class, the coarse classifier can detect a limited set of visual features described in the image to determine whether the image includes objects belonging to that object class. In a particular example, the coarse classifier can detect whether the image describes objects categorized in one or more categories, including but not limited to: text, barcodes, landmarks, people, food, media objects, plants, etc. For barcodes, the coarse classifier can determine whether the image includes parallel lines of varying widths. Similarly, for machine-readable codes (e.g., QR codes, etc.), the coarse classifier can determine whether the image includes patterns that indicate the presence of machine-readable codes.

[0078] A coarse classifier can output data indicating whether a class of objects was detected in an image. It can also output confidence values ​​representing the confidence level of detecting a class of objects in the image and / or confidence values ​​representing the confidence level of the actual objects depicted in the image (e.g., a cereal box).

[0079] Object detector 128 can receive image data representing the field of view of camera 124 (e.g., content displayed in the viewfinder) and detect the presence of one or more objects in the image data. If at least one object is detected in the image data, camera application 126 can provide (e.g., send) the image data to visual search system 104 via network 180. As described below, visual search system 104 can identify objects in the image data and provide object-related content to user computing device 102.

[0080] The visual search system 104 includes one or more front-end servers 136 and one or more back-end servers 140. The front-end server 136 can receive image data from a user computing device (e.g., user computing device 102). The front-end server 136 can provide the image data to the back-end server 140. The back-end server 140 can identify content related to the identified objects in the image data and provide that content to the front-end server 136. Furthermore, the front-end server 136 can provide that content to the mobile device receiving the image data.

[0081] Backend server 140 includes one or more processors 142 and memory 146. The one or more processors 142 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operable in series. Memory 146 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 146 can store data 148 and instructions 150 executed by processor 142 to cause the visual search system 104 to perform operations. Backend server 140 may also include object recognizer 152, query processing system 154, and content ranking system 156. Object recognizer 152 can process image data received from mobile devices (e.g., user computing device 102, etc.) and, if present, identify objects in the image data. For example, object recognizer 152 can use computer vision and / or other object recognition techniques (e.g., edge matching, pattern recognition, grayscale matching, gradient matching, etc.) to identify objects in the image data.

[0082] In some implementations, the visual search system 104 includes one or more server computing devices, or is otherwise implemented by one or more server computing devices. In instances where the visual search system 104 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0083] In some implementations, object recognizer 152 includes multiple object recognizer modules, for example, one object per class of objects in their respective classes. For example, object recognizer 152 may include a text recognizer module that recognizes text in image data (e.g., recognizes characters, words, etc.), a barcode recognizer module that recognizes (e.g., decodes) barcodes in image data (including machine-readable codes such as QR codes), a landmark recognizer module that recognizes landmarks in image data, and / or other object recognizer modules that recognize specific object classes.

[0084] In some implementations, the query processing system 154 includes multiple processing systems. One example system may allow it to identify multiple candidate search results. For example, the system may identify multiple candidate search results upon first receiving a visual query image. Alternatively, the system may identify multiple search results after further processing has been completed. Specifically, the system may identify multiple search results based on a more targeted query already generated by the system. More specifically, the system may generate multiple candidate search results upon first receiving a visual query image, and then regenerate multiple candidate search results after further processing based on a more targeted query generated by the system.

[0085] As another example, query processing system 154 may include a system for generating user-specific interest data (e.g., represented using graphs). More specifically, user-specific interest data can be used in part to determine which of a number of candidate results are most likely of interest to the user. Specifically, user-specific interest data can be used in part to determine which results exceed a threshold so that they are worth showing to the user. For example, a visual search system may not be able to output all candidate results in an overlay on the user interface. User-specific interest data can help determine which candidate results will be output and which will not.

[0086] As another example, query processing system 154 may include a system associated with a composite content set. More specifically, the composite content set may reference multiple canonical items responding to a visual search query. In some implementations, the system may include a graph describing multiple items, where a corresponding content set (e.g., user-generated content, such as product reviews) is associated with each of multiple distinct items. A composite content set can be applied when a user provides a visual query image having objects containing multiple canonical items responding to a visual search query. For example, if a user provides a visual query image of the packaging cover of a Blu-ray of a particular movie, the composite content set system may output results not only related to the Blu-ray of that particular movie, but also results related to movies in general, casts, or any number of other content related to that movie.

[0087] As another example, query processing system 154 may include systems related to the compositional characteristics of a visual query image. More specifically, the compositional characteristics used by the visual search system may include various attributes of the image (the number of objects in the image, the distance between the objects and the camera, the angular orientation of the image, etc.). The compositional characteristic system can serve as part of intelligently dissolving visual query ambiguities between a specific object described in the query image and the classification result. For example, a user might submit images of three cereal boxes with the text query "Which one has the most fiber?" The compositional characteristic system can help dissolve ambiguities based on the compositional characteristics identified by the system, whether the query is for the three types of cereal, all cereal, or a specific brand of cereal.

[0088] As another example, query processing system 154 may include a system related to multiple entities associated with a visual query. More specifically, multiple entities refer to multiple topics contained within the image and the context surrounding it in any way (GPS coordinates of the location where the image was taken, text titles accompanying the image, web pages where the image is found, etc.). The multi-entity system can further compose auxiliary queries in the form of entity combinations to include all candidate search results that the user might want. For example, in response to multiple identified entities, a photo of Emma Watson and Daniel Radcliffe in front of the Kodak Theatre in Los Angeles could have many underlying user intents beyond searching for “Emma Watson” or “Daniel Radcliffe,” such as “Harry Potter at the Oscars” or “Emma Watson, Oscar designer.”

[0089] In some implementations, the content ranking system 156 can be used to rank candidate search results at multiple different points in the visual search system's processing. One example application is generating a ranking of search results after multiple search results are initially identified. On the other hand, the initial search results may only be preliminary, and the ranking system 156 can generate a ranking of search results after the query processing system has created a more targeted query. More specifically, the ranking system 156 can generate ranked search results for multiple candidate search results when the system first identifies the set of candidate search results, and then search again after a more targeted query is performed (e.g., preliminary rankings can be used to determine the most likely combinations of multiple entities). The rankings created by the ranking system 156 can be used to determine the final output of candidate search results to the user by determining the order in which the search results will be output and / or whether candidate search results will be output at all.

[0090] The multiple processing systems included in query processing system 154 can be arbitrarily combined and used in any order to process user-submitted visual queries in the most intelligent way, so as to provide users with the most intelligent results. Furthermore, ranking system 156 can also be used in any combination with query processing system 154.

[0091] After selecting content, it can be provided to the user computing device 102 from which image data is received, stored in the content cache 130 of the visual search system 104, and / or stored on top of the memory stack of the front-end server 136. This allows content to be quickly presented to the user in response to a user request. If content is provided to the user computing device 102, the camera application 126 can store the content in the content cache 134 or other fast-access memory. For example, the camera application 126 can store the content of an object and reference the object so that the camera application 126 can identify the appropriate content of the object in response to determining the content to be presented.

[0092] Camera application 126 can present the content of an object in response to a user's interaction with a visual indicator of the object. For example, camera application 126 can detect user interaction with a visual indicator of an object and request the content of the object from visual search system 104. In response, front-end server 136 can fetch content from content cache 130 or the top of the memory stack and provide the content to user computing device 102 that received the request. If the content is provided to user computing device 102 before user interaction is detected, camera application 126 can obtain content from content cache 134.

[0093] In some implementations, the visual search system 104 includes an object detector 128, for example, instead of a camera application 126. In these examples, the camera application 126 may continuously send image data to the visual search system 104 while the camera application 126 is active or when the user puts the camera application 126 into a content request mode, such as continuously sending image data in an image stream. The content request mode allows the camera application 126 to continuously send image data to the visual search system 104 to request content of objects identified in the image data. The visual search system 104 may detect objects in the image, process the image (e.g., select visual indicators for detected objects), and send the results (e.g., visual indicators) to the camera application 126 for presentation in a user interface (e.g., viewfinder). The visual search system 104 may also continue processing the image data to identify objects, select content for each identified object, and cache or send the content to the camera application 126.

[0094] In some implementations, camera application 126 includes an on-device object recognizer that identifies objects in image data. In this example, camera application 126 can identify objects and request content recognition of those objects from visual search system 104, or recognize content from on-device content data storage. Compared to object recognizer 152 of visual search system 104, the on-device object recognizer can be a lightweight object recognizer capable of recognizing a more limited set of objects or using computationally less expensive object recognition techniques. This enables mobile devices with processing power lower than typical servers to perform object recognition processing. In some implementations, camera application 126 can use the on-device recognizer to perform initial object recognition and provide image data to visual search system 104 (or another object recognition system) for confirmation. On-device content data storage can also store a more limited set of content than content data storage unit 138 or links to resources containing that content, to preserve data storage resources of user computing device 102.

[0095] User computing device 102 may also include one or more user input components 122 for receiving user input. For example, user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). Touch-sensitive components can be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other means by which the user can provide input.

[0096] Network 180 can be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., the Internet), or a combination thereof, and can include any number of wired or wireless links. Typically, communication on Network 180 can be conducted via any type of wired and / or wireless connection, using various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).

[0097] Figure 1 An example computing system that can be used to implement this disclosure is shown. Other different distributions of the components may also be used. For example, some or all of the different aspects of the visual search system may alternatively be located and / or implemented at the user computing device 102.

[0098] Example model layout

[0099] Figure 2 A block diagram of an exemplary visual search system 200 according to exemplary embodiments of the present disclosure is depicted. In some implementations, the visual search system 200 is configured to receive an input dataset including a visual query 204 and, as a result of receiving the input data 204, provide output data 206 that offers more personalized and / or intelligent results to the user. For example, in some implementations, the visual search system 200 may include a query processing system 202 operable to facilitate the output of more personalized and / or intelligent visual query results.

[0100] In some implementations, query processing system 202 includes or utilizes a user-centric visual interest graph to provide more personalized search results. In one example use, visual search system 200 may use the user interest graph to rank or filter search results, including visual discovery notifications, announcements, or suggestions of other opportunities. In exemplary embodiments, user interest-based search result personalization may be particularly advantageous, where search results are presented as visual result notifications (e.g., referred to in some cases as “gleams”) on the query image in the form of enhanced overlays.

[0101] In some implementations, user-specific interest data (e.g., using graph representations) can be aggregated over time, at least in part, by analyzing images the user has previously engaged with. In other words, the visual search system 200 can attempt to understand the user's visual interests by analyzing images the user has engaged with over time. When a user engages with an image, it can be inferred that some aspects of the image are of interest to the user. Therefore, items (e.g., objects, entities, concepts, products, etc.) that are included in or related to such images can be added to or otherwise labeled in the user-specific interest data (e.g., graphs).

[0102] For example, user-engaged images may include photos captured by the user, screenshots captured by the user, or images included in web-based or app-based content viewed by the user. In another potentially overlapping example, user-engaged images may include actively engaged images where the user actively participates by requesting an action to be performed on the image. For example, the requested action may include performing a visual search relative to the image or having the user explicitly mark the image as including the user's visual interest. As another example, user-engaged images may include passively observed images that have been presented to the user but in which the user has not specifically participated. Visual interest may also be inferred from text content entered by the user (e.g., text- or term-based queries).

[0103] In addition to the above descriptions, users may be given control over whether and when a system, program, or feature described herein may collect user information (e.g., information about the user's social networks, social behavior, activities, occupation, user preferences, or the user's current location), and whether content or communications may be sent to the user from the server. Furthermore, some data may be processed in one or more ways before storage or use to remove personally identifiable information. For example, a user's identity may be processed to the point that their personally identifiable information cannot be determined, or their geographic location may be generalized to the location from which location information is obtained (such as city, zip code, or state), thus making it impossible to determine the user's specific location. Therefore, users have control over what information about themselves is collected, how that information is used, and what information is provided to them.

[0104] The visual search system 200 can use the query processing system 202 to select search results for the user. The user interest system 202 can be used at various stages of search processing, including query modification, result recognition, and / or other stages of visual search processing.

[0105] As an example, Figure 3A block diagram of an exemplary visual search system 400 according to an exemplary embodiment of the present disclosure is depicted. The visual search system 400 operates in a two-stage process. In a first stage, a query processing system 202 may receive input data 204 (e.g., a visual query including one or more images) and generate a set of candidate search results 206 in response to the visual query. For example, the candidate search results 206 may be obtained without considering user-specific interests. In a second stage, a ranking system 402 may be used by the visual search system 400 to help rank the one or more candidate search results 206 to return to the user as the final search result (e.g., as output data 404).

[0106] As an example, the visual search system 400 can use the ranking system 402 to generate a ranking of multiple candidate search results 206 based at least in part on comparisons between multiple candidate search results 206 and user-specific user interests associated with the user, obtained in the query processing system 202. For example, weights of certain items captured within the user interest data can be used to modify or reweight the initial search scores associated with the candidate search results 206, which can further result in the search results 206 being reordered before being output to the user 404.

[0107] A visual search system 400 can select at least one of a plurality of candidate search results 206 as at least one selected search result, at least partially based on ranking, and then provide at least one selected visual result notification, each associated with the at least one selected search result, to be displayed to the user (e.g., as output data 404). In one example, each of the selected search results (or multiple results) can be provided to overlay a specific sub-section of an image associated with the selected search result. In this way, user interests can be used to provide personalized search results and reduce clutter in the user interface.

[0108] As another example variant Figure 4 A block diagram of an example visual search system 500 according to an exemplary embodiment of the present disclosure is depicted. The visual search system 500 is similar to... Figure 3 The visual search system 400, except that the visual search system 500 also includes a context component 502, which receives and processes context information 504 to describe the implicit features of the visual query and / or the user's search intent.

[0109] Contextual information 504 may include any other available signals or information that help understand the implicit characteristics of the query. For example, location, time, input pattern, and / or various other information can be used as context.

[0110] As another example, contextual information 504 may include various attributes of the image, information about where the user obtained the image, information about other uses or instances of the image, and / or various other contextual information. In one example, the image used in a visual search query is presented in a web document (e.g., a webpage). References to other entities included in the web document (e.g., textual and / or visual references) can be used to identify potential other entities that can be used to form a combination of multiple entities.

[0111] In another example, contextual information 504 may include information obtained from other web documents that include additional instances of the image associated with the visual search query. As another example, contextual information 504 may include textual metadata (e.g., EXIF ​​data) associated with the image. Specifically, the textual metadata can be accessed and identified as other entities associated with the visual search. Specifically, the textual metadata may include the title of the image used in the user-submitted visual query.

[0112] As another example, contextual information 504 may include information obtained through an initial search based on a visual query. More specifically, the initial search can be performed using information from the visual query, and after obtaining a first set of preliminary search results, other entities referenced by the preliminary search results can be identified. In particular, entities identified in some preliminary results that exceed a threshold can be determined to be sufficiently relevant to be included in subsequent queries.

[0113] refer to Figure 2 , Figure 3 and Figure 4 The visual search system 200, 400, and / or 500 can implement an edge detection algorithm to process objects depicted in an image provided as visual query input data 204. Specifically, the acquired image can be filtered using an edge detection algorithm (e.g., a gradient filter) to obtain a resulting image representing a binary matrix that can be measured in both horizontal and vertical directions, determining the position of the object contained in the image within the matrix. Furthermore, the resulting image can be further filtered using Laplacian and / or Gaussian filters to improve edge detection. Boolean operators, such as "AND" and "OR" Boolean operators, can then be used to compare the object with multiple training images and / or any kind of historical images and / or contextual information 504. Utilizing Boolean comparisons provides very fast and efficient comparisons, which is desirable; however, in some cases, non-Boolean operators may be required.

[0114] Furthermore, similarity algorithms can be derived from... Figure 2 , Figure 3 and / or Figure 4The visual search systems 200, 400, and / or 500 are accessible, where the algorithms can access the aforementioned edge detection algorithms and store the output data. Furthermore, additionally and / or optionally, the similarity algorithm can estimate a pairwise similarity function between each image and / or query input data 204 and multiple other images and / or queries and / or contextual information 504, which can be any type of training data and / or historical data. The pairwise similarity function describes whether two data points are similar.

[0115] Alternatively or optionally, Figure 2 , Figure 3 and / or Figure 4 The visual search systems 200, 400, and / or 500 can implement clustering algorithms to process images provided as visual query input data 204. The search system can execute the clustering algorithm and assign images and / or queries to clusters based on an estimated pairwise similarity function. Before executing the clustering algorithm, the number of clusters can be unknown and can vary from one execution of the clustering algorithm to the next, based on the image / visual query input data 204, the estimated pairwise similarity function for each pair of images / queries, and the random or pseudo-random selection of the initial images / queries assigned to each cluster.

[0116] Visual search systems 200, 400, and / or 500 can perform clustering algorithms once or multiple times on the image / query input dataset 204. In some exemplary embodiments, visual search systems 200, 400, and / or 500 can perform clustering algorithms for a predetermined number of iterations. In some exemplary embodiments, visual search systems 200, 400, and / or 500 can perform clustering algorithms and aggregate the results until a distance metric with a non-transitive pairwise similarity function is reached.

[0117] Example Method

[0118] Figure 9 A flowchart depicts an example method 1000 for providing more personalized search results according to exemplary embodiments of this disclosure. Although for illustrative and discussion purposes, Figure 9 The steps are described in a specific order, but the method disclosed herein is not limited to the specific order or arrangement described. Various steps of method 1000 may be omitted, rearranged, combined, and / or modified in various ways without departing from the scope of this disclosure.

[0119] In 1002, the computing system can obtain visual queries. For example, the computing system (e.g., Figure 1 The user computing device 102 and the visual search system 104 can obtain visual query input data from the user (e.g., Figure 2 Visual query input data 204).

[0120] In 1004, the computing system can identify multiple candidate search results and corresponding search result notification coverage. For example, as Figure 3 The output of the visual search result model 402 allows the computing system to receive multiple current candidate search results and corresponding enhanced overlays on the user interface, providing visual result notifications for the search results as overlays of images (or multiple images) included in the visual query.

[0121] More specifically, the computational system can input the previously obtained visual query into the query processing system. For example, the computational system can input visual query input data 204 into query processing system 202. Before inputting the visual query, the computational system can access an edge detection algorithm. More specifically, the acquired image can be filtered using an edge detection algorithm (e.g., a gradient filter) to obtain a resulting image that represents a binary matrix that can be measured in the horizontal and vertical directions, determining the position of objects contained in the image within the matrix.

[0122] In 1006, the computing system can utilize a user-centric visual interest map to select and / or filter multiple previously obtained candidate search results and corresponding search result notification overlays based on observed user visual interests. The query processing system may include a user-centric visual interest map. For example, query processing system 202.

[0123] In 1008, the computing system can generate rankings for multiple candidate search results. For example, the computing system can receive the current ranking of candidate search results and their corresponding search result notification coverage as output of a ranking system (e.g., ranking system 402).

[0124] More specifically, a ranking system can generate rankings at least in part based on comparisons of multiple candidate search results with user-specific interest data associated with the user, contained in the query processing system. For example, item weights can be applied to modify or reweight the initial search scores associated with the candidate search results.

[0125] In some implementations, visual search systems can consider recurring notification overlays in an image, which can be identified as containing multiple identical objects. The visual search system can then output only one of several potential candidate search result notification overlays that provide the same search result.

[0126] In 1010, the computing system can select at least one of multiple candidate search results as at least one selected search result, for example, output data 404. More specifically, the visual search system 400 can select at least one of multiple candidate search results as at least one selected search result based at least partially on ranking, and then provide at least one selected visual result notification associated with each of the at least one selected search result, to overlay on specific sub-sections of the image associated with the selected search result. In this way, user interests can be used to provide personalized search results and reduce clutter in the user interface.

[0127] At 1012, the computing system can provide the user with at least one selected visual result notification. For example, the computing system can provide the user with output data 404, which includes prediction results based on the output of the query processing system 202.

[0128] Figure 5 It shows Figure 9 The advantages of the example methods described in section 602 are shown. Section 602 demonstrates the advantages of not using... Figure 9 An example augmented reality user interface in the case of the method described herein. User interface 602 is shown without... Figure 9 In the case of the method described herein, the interface 602 and the notification overlay 604 are overly confused, making the interface 602 unreadable and the notification overlay 604 difficult to use.

[0129] Instead, interface 606 displays the use of Figure 9 An example augmented reality user interface of the method is shown. Interface 606 illustrates the use of... Figure 9 In the case of the method described, only the selected notification overlay 604 is displayed, so that the user can still see through the selected notification overlay 604 and easily access all selected notification overlays 604.

[0130] Figure 10 A flowchart of an example method 1100 according to an exemplary embodiment of the present disclosure is described. Although for purposes of illustration and discussion, Figure 10 The steps are described in a specific order, but the method of this disclosure is not limited to the specific order or arrangement described. Various steps of method 1100 may be omitted, rearranged, combined and / or modified in various ways without departing from the scope of this disclosure.

[0131] In 1102, the computing system can obtain visual queries. For example, the computing system (e.g., Figure 1 The user computing device 102 and the visual search system 104 can obtain visual query input data 204 from the user.

[0132] In 1104, the computing system can access a graph describing multiple different items. Specifically, a corresponding content set (e.g., user-generated content, such as product reviews) is associated with each of the multiple different items.

[0133] More specifically, the computing system can input the previously obtained visual query into the query processing system. For example, the computing system can input visual query input data 204 into query processing system 202.

[0134] At 1106, the computational system can select multiple items from the graph. More specifically, the query processing system 202 can utilize a graph that may be a hierarchical representation of multiple distinct items. Selecting multiple items from the graph may include, based on a visual search query, identifying a principal item in the graph corresponding to an object depicted in an image (e.g., a specific movie shown in an image). Next, the visual search system can identify one or more supplementary items in the graph associated with the principal item within the hierarchical representation of the graph, and select the principal item and one or more supplementary items as multiple items.

[0135] At 1108, the computing system can provide the user with a combined set of content as search results. For example, the computing system can provide the user with output data 404, including predictions based on the output of the visual search results model 202.

[0136] Figure 6 It shows Figure 10 The advantages of the example method described herein. User interface 702 displays the unused... Figure 10 Example search results for the method described in [the document]. User interface 702 shows [the results] without [the specified method]. Figure 10 In the case of the method described herein, the search results include only results related to the exact same object used as the visual query. Conversely, user interface 704 uses... Figure 10 The method described in [the document] displays example search results. User interface 704 shows the use of [the method described]. Figure 10 The method described herein expands the search results to include results associated with multiple canonical entities.

[0137] Figure 11 A flowchart of an example method 1200 according to an exemplary embodiment of the present disclosure is described. Although for purposes of illustration and discussion, Figure 11 The steps are described in a specific order, but the method disclosed herein is not limited to the specific order or arrangement described. Various steps of method 1200 may be omitted, rearranged, combined, and / or modified in various ways without departing from the scope of this disclosure.

[0138] In 1202, the computing system can obtain visual queries. For example, the computing system (e.g., Figure 1The user computing device 102 and the visual search system 104 can obtain visual query input data 204 from the user.

[0139] In 1204, the computing system can identify one or more constituent characteristics of a visual query image. Specifically, various attributes of the image (e.g., distance to one or more identified objects, number of objects, relative similarity of objects, angular orientation, etc.)

[0140] More specifically, the computational system can input the previously obtained visual query into the query processing system. For example, the computational system can input visual query input data 204 into query processing system 202. Before inputting the visual query, the computational system can access an edge detection algorithm. More specifically, the acquired image can be filtered using an edge detection algorithm (e.g., a gradient filter) to obtain a resulting image that represents a binary matrix that can be measured in the horizontal and vertical directions, determining the position of objects contained in the image within the matrix.

[0141] At 1206, the computational system can determine whether a visual search query is object-specific or categorical. More specifically, the query processing system 202 can use the identified combinatorial characteristics to predict whether a visual query is a categorical query related to an expanded corpus of search results, or an object-specific query specifically belonging to one or more objects identified in the visual query.

[0142] At 1208, the computing system can provide the user with one or more object-specific search results. For example, the computing system can provide the user with output data 404, including predictions based on the output of the visual search results model 202.

[0143] In 1210, the computing system can provide the user with one or more categorized search results. For example, the computing system can provide the user with output data 404, including predictions based on the output of the visual search results model 202. Figure 7 It shows Figure 11The advantages of the example method described are illustrated below. Images 802 and 804 are two examples of image variations that may have the same accompanying text query (e.g., "Which has the highest fiber content?"). In response to the two example images 802 and 804, despite the same accompanying text query, the visual search system will return two different results. User interface 806 shows that, in response to image 802, based on compositional characteristics (e.g., the cereal boxes are centered in focus, the cereal boxes are at approximately a 90-degree angle, meaning they are not tilted, and all cereal boxes contained in the image are clearly identifiable), the visual search system can return an image of the cereal with the highest fiber content. In contrast, user interface 808 shows that, in response to image 804, based on compositional characteristics (e.g., the image was taken to make the entire channel visible, the cereal boxes are not fully in focus, and the cereal boxes are at a 30-degree angle, meaning they are more tilted), the visual search system can return an image of the cereal with the highest fiber content among all cereals.

[0144] Figure 12 A flowchart of an example method 1300 according to an exemplary embodiment of the present disclosure is described. Although for purposes of illustration and discussion, Figure 12 The steps are described in a specific order, but the method of this disclosure is not limited to the specific order or arrangement described. Various steps of method 1300 may be omitted, rearranged, combined and / or modified in various ways without departing from the scope of this disclosure.

[0145] In 1302, the computing system can obtain visual queries. For example, the computing system (e.g., Figure 1 The user computing device 102 and the visual search system 104 can obtain visual query input data 204 from the user.

[0146] In 1304, the computing system can identify one or more additional entities associated with the visual query. Specifically, the computer vision search system can identify one or more entities associated with the visual search query based on one or more contextual signals or information.

[0147] More specifically, the computational system can input the previously obtained visual query into the query processing system. For example, the computational system can input visual query input data 204 into query processing system 202. Before inputting the visual query, the computational system can access an edge detection algorithm. More specifically, the acquired image can be filtered using an edge detection algorithm (e.g., a gradient filter) to obtain a resulting image that represents a binary matrix that can be measured in the horizontal and vertical directions, determining the position of objects contained in the image within the matrix.

[0148] At 1306, the computing system can determine a combined query relating to a combination of a first entity and one or more additional entities. More specifically, the query processing system 202 can utilize multiple entities to determine a combined query relating to a combination of a first entity and one or more additional entities. Specifically, entities may include people, objects, and / or abstract entities such as events.

[0149] At 1308, the computing system can provide the user with a set of content related to a combination of the first entity and one or more additional entities. For example, the computing system can provide the user with output data 404, including predictions based on the output of the visual search results model 202.

[0150] Figure 8 It shows Figure 12 Advantages of the example method described in [the document]. The user interface 904 [is available] without using [the example method]. Figure 12 In the case of the method described herein, example search results are displayed based on the example visual query 902. User interface 904 shows: without... Figure 12 In the case of the method described, the search results include results related only to a single object identified in the visual query image, because current technology cannot combine queries of multiple entities for a given visual query. In contrast, user interface 906 uses... Figure 12 The method described in section 902 shows example search results based on the same example visual query 902. This method considers all faces identified in the visual query and composes a query containing some or all of them or events related to them, thereby generating search results for a specific awards ceremony. The result is a primary preliminary search result including all identified faces. Section 906 shows the use of... Figure 12 The method described herein expands the search results and allows multiple objects of interest to be considered based on their compositional characteristics.

[0151] Additional Publication

[0152] This article discusses technologies involving servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from these systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, task and function partitioning among components. For example, the processing discussed in this article can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can run sequentially or in parallel.

[0153] While this subject matter has been described in detail with reference to various specific example embodiments, each example is provided by way of explanation and not limitation. Modifications, variations, and equivalents can be readily made to these embodiments upon understanding the foregoing. Therefore, this disclosure does not exclude modifications, alterations, and / or additions to the subject matter of this disclosure that will be apparent to those skilled in the art. For example, features shown or described as part of one embodiment may be used with another embodiment to produce further embodiments. Therefore, this disclosure is intended to cover such changes, variations, and equivalents.

Claims

1. A computer-implemented method for providing personalized visual search query result notifications within a user interface overlaid on an image, the method comprising: A visual search query associated with a user is obtained through a computing system including one or more computing devices, wherein the visual search query includes an image; The computational system identifies multiple candidate search results for a visual search query, wherein each candidate search result is associated with a specific sub-part of an image, and wherein multiple candidate visual result notifications are associated with multiple candidate search results respectively. The computing system accesses user-specific user interest data and descriptions of user visual interests associated with a user, wherein the user-specific user interest data includes a user-centered visual interest map reflecting multiple user interests, and the user-centered visual interest map includes a hierarchical representation of multiple different items. Based on a visual search query, multiple selections are made from a user-centric visual interest map based on at least a portion of an image, wherein selecting multiple selections from the user-centric visual interest map includes: Based on visual search queries, identify one or more principal items in the user-center visual interest map that correspond to the objects depicted in the image; Identify one or more supplementary items related to one or more main items within the hierarchical representation of the user-centered visual interest map; and Select a main item and one or more additional items as multiple selections; The ranking of multiple candidate search results is generated by a computing system based at least in part on a comparison of multiple candidate search results with user-specific user interest data associated with the user, wherein the ranking is determined based on one or more main items and one or more supplementary items associated with one or more of the multiple user interests, wherein the one or more supplementary items are determined to be associated with one or more main items based on a hierarchical representation. The system selects at least one of multiple candidate search results as at least one selected search result by calculating at least part of the ranking and one or more additional factors; and The system calculates and provides at least one selected visual result notification associated with at least one selected search result, based on the ranking, to overlay on a specific sub-section of the image associated with the selected search result.

2. The computer-implemented method according to claim 1, wherein, At least in part, this is achieved by analyzing images from users' past engagements, aggregating user-specific interest data over time.

3. The computer-implemented method according to claim 2, wherein, Images that the user has previously engaged with include photos captured by the user, screenshots captured by the user, or images included in web-based or app-based content that the user has viewed.

4. The computer-implemented method according to claim 2, wherein, Images from the user's past engagement include passively observed images presented to the user but in which the user did not actively participate.

5. The computer-implemented method according to claim 2, wherein, Images in which the user has previously participated include actively participated images, where the user has proactively engaged by requesting an action to be performed on the image.

6. The computer-implemented method according to claim 2, wherein, Images that the user has previously engaged with include images that the user has explicitly indicated contain the user's visual interests.

7. The computer-implemented method according to claim 1, wherein, The stored user-specific interest data includes a continuously updated entity graph that identifies one or more specific entities, categorized entities, or abstract entities.

8. The computer-implemented method according to claim 1, wherein, User-specific interest data is based at least in part on user-captured images, wherein repeated user captures include: Variable weighted interest bias overlaid on the identified user visual interests.

9. The computer-implemented method according to claim 1, wherein: The variable weighted interest bias assigned to the identified visual interests decays over time, such that user-specific interest data is at least partially based on time frames expressing those interests.

10. The computer-implemented method according to claim 1, wherein, Visual queries include passive queries, which include the presence of an image included in the visual query on a display screen without any specific indication of user interest.

11. The computer-implemented method according to claim 1, wherein, Visual queries include proactive queries, which include specific indications of user interests when receiving search results in response to a visual query.

12. A computing system for providing personalized visual search query result notifications within a user interface overlaid on an image, the computing system comprising: One or more processors; as well as One or more non-transitory computer-readable media store instructions that, when executed by one or more processors, cause a computing system to perform operations, said operations including: Obtain visual search queries associated with the user, wherein the visual search queries include images; Identify multiple candidate search results for a visual search query, wherein each candidate search result is associated with a specific sub-part of an image, and wherein multiple candidate visual result notifications are associated with multiple candidate search results respectively; Access user-specific user interest data associated with a user and describing the user's visual interests, wherein the user-specific user interest data includes a user-centered visual interest map reflecting multiple user interests, wherein the user-centered visual interest map includes a hierarchical representation of multiple different items; Based on a visual search query, multiple selections are made from a user-centric visual interest map based on at least a portion of an image, wherein selecting multiple selections from the user-centric visual interest map includes: Based on visual search queries, identify one or more principal items in the user-center visual interest map that correspond to the objects depicted in the image; Identify one or more supplementary items related to one or more main items within the hierarchical representation of the user-centered visual interest map; and Select a main item and one or more additional items as multiple selections; The ranking of multiple candidate search results is generated at least in part based on a comparison of multiple candidate search results with user-specific user interest data associated with the user, wherein the ranking is determined based on one or more main items and one or more supplementary items associated with one or more of the multiple user interests, wherein the one or more supplementary items are determined to be associated with one or more main items based on a hierarchical representation; At least one of a plurality of candidate search results is selected as at least one selected search result, based at least in part on ranking and one or more additional factors; and Based on ranking, at least one selected visual result notification is provided, each associated with at least one selected search result, to overlay on a specific sub-section of the image associated with the selected search result.

13. The computing system according to claim 12, wherein, The operation also includes: Processing images to identify one or more constituent features of the image; and The query type of an image is determined based on one or more compositional characteristics. Among them, based on ranking, one or more additional items and query type, at least one of multiple candidate search results is selected as at least one selected search result.

14. The computing system according to claim 13, wherein, The query type includes at least one of category query or object-specific query.

15. The computing system according to claim 12, wherein, The ranking-based notification provides at least one selected visual result, including: Return a combined set of content as the search results, wherein the combined set of content includes at least one selected search result.

16. The computing system according to claim 12, wherein, The computing system includes one or more computing devices, wherein the one or more computing devices include mobile computing devices.

17. One or more non-transitory computer-readable media storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, the operations including: Obtain a visual search query associated with a user, wherein the visual search query includes an image; Identify multiple candidate search results for a visual search query, wherein each candidate search result is associated with a specific sub-part of an image, and wherein multiple candidate visual result notifications are associated with multiple candidate search results respectively; Access user-specific user interest data associated with a user and describing the user's visual interests, wherein the user-specific user interest data includes a user-centered visual interest map reflecting multiple user interests, and wherein the user-centered visual interest map includes a hierarchical representation of multiple different items; Based on a visual search query, multiple selected items are selected from a user-centric visual interest map based on at least a portion of an image, wherein selecting multiple selected items from the user-centric visual interest map includes: Based on visual search queries, identify one or more principal items in the user-center visual interest map that correspond to the objects depicted in the image; Identify one or more supplementary items related to one or more main items within the hierarchical representation of the user-centered visual interest map; and Select a main item and one or more additional items as multiple selections; The ranking of multiple candidate search results is generated at least in part based on a comparison of multiple candidate search results with user-specific user interest data associated with the user, wherein the ranking is determined based on one or more main items and one or more supplementary items associated with one or more of the multiple user interests, wherein the one or more supplementary items are determined to be associated with one or more main items based on a hierarchical representation; At least one of a plurality of candidate search results is selected as at least one selected search result, based at least in part on ranking and one or more additional factors; and Based on ranking, at least one selected visual result notification is provided, each associated with at least one selected search result, to overlay on a specific sub-section of the image associated with the selected search result.

18. One or more non-transitory computer-readable media according to claim 17, wherein, A user-centric visual interest map describing multiple different items includes multiple nodes corresponding to multiple indexed images.

19. One or more non-transitory computer-readable media according to claim 18, wherein, Based at least in part on the visual similarity between indexed images, multiple nodes are arranged within the user-centric visual interest map such that the distance between a pair of nodes within the user-centric visual interest map is negatively correlated with the visual similarity between the corresponding pair of indexed images.

20. One or more non-transitory computer-readable media according to claim 17, wherein, The operation also includes: Use an object detector to process the image to determine the objects depicted in the image.

Citation Information

Patent Citations

  • Architecture for responding to a visual query

    US20110125735A1

  • Mapping images to search queries

    US20170300495A1