Object Filtering and Information Display in Augmented Reality Experiences

The computing system addresses user inefficiencies in searching for specific information within scenes by processing image data from mobile devices to recognize objects, obtain object-specific information, and provide augmented reality overlays for filtering, thus enhancing user experience and efficiency.

JP7675127B2Active Publication Date: 2025-05-12GOOGLE LLC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023076289
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-12-20
Filing Date
2023-05-02
Publication Date
2025-05-12
Estimated Expiration
2043-05-02

AI Technical Summary

Technical Problem

Users face inefficiencies in searching for specific information within scenes, such as products in a grocery store, due to repetitive and time-consuming searches, and difficulty in tracking preferences.

Method used

A computing system that acquires image data from a mobile device, processes it to recognize objects, obtains object-specific information, and provides user interface elements overlaid on the image data to display this information, allowing users to filter objects based on selected tags.

Benefits of technology

Enables users to efficiently find and track specific objects or product information within a scene, reducing search time and improving user experience through augmented reality overlays and filtering capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007675127000001
    Figure 0007675127000001
  • Figure 0007675127000002
    Figure 0007675127000002
  • Figure 0007675127000003
    Figure 0007675127000003
Patent Text Reader

Abstract

To provide a user interface that provides information associated with a scene.SOLUTION: Systems and methods for providing scene understanding can include obtaining a plurality of images, stitching images associated with the scene, detecting objects in the scene, and providing information associated with the objects in the scene. The systems and methods can include determining filter tags or query tags that can be selected to filter the plurality of objects, which can then be provided as information to the user to provide further insight on the scene. The information may be provided in an augmented-reality experience via text or other user-interface elements anchored to objects in the images.SELECTED DRAWING: Figure 2A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Related Applications This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 340,078, filed May 10, 2022. U.S. Provisional Patent Application No. 63 / 340,078 is incorporated herein by reference in its entirety.

[0002] The present disclosure relates generally to providing a user interface that provides information associated with a scene. More particularly, the present disclosure relates to recognizing objects in a scene, generating tags associated with the objects, filtering the objects based on selection of particular tags, and providing object information for the filtered objects. [Background technology]

[0003] Understanding a scene and the objects in the scene can be difficult. In particular, understanding a scene can require repeated and endless searching for objects in the scene, and sometimes it can be difficult to determine what to search for. Furthermore, a user may ask the same questions at a particular location during each visit to the location. The user may be forced to inefficiently search for the same queries during each visit.

[0004] For example, a user may go shopping at a local grocery store. During the shopping trip, the user may want to select a new type or brand of coffee to try, which the user may do at each visit. The user may end up judging the name and selecting each bag, searching each coffee type and brand to see which coffee meets the user's preferences. The searches may be tedious and time consuming. Furthermore, the user may have difficulty keeping track of which coffees meet their preferences and which do not. This may result in possible inefficiencies during each shopping visit. Summary of the Invention [Means for solving the problem]

[0005] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the description that follows, or may be learned from the description, or may be learned through practice of the embodiments.

[0006] One exemplary aspect of the present disclosure is directed to a computing system. The system can include one or more processors and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations can include obtaining image data generated by a mobile image capture device. The image data can represent a scene. The operations can include processing the image data to determine a plurality of objects in the scene. In some implementations, the plurality of objects can include one or more consumer products. The operations can include obtaining object-specific information for one or more objects of the plurality of objects. The object-specific information can include one or more details associated with each of the one or more objects. The operations can include providing one or more user interface elements overlaid on the image data. In some implementations, the one or more user interface elements can describe the object-specific information.

[0007] Another exemplary aspect of the present disclosure is directed to a computer-implemented method. The method may include obtaining, by a computing system including one or more processors, video stream data generated by a mobile image capture device. In some implementations, the video stream data may include a plurality of image frames. The method may include determining, by the computing system, that a first image frame and a second image frame are associated with a scene. The method may include generating, by the computing system, scene data including a first image frame and a second image frame of the plurality of image frames. In some implementations, the method may include processing, by the computing system, the scene data to determine a plurality of objects in the scene. The plurality of objects may include one or more consumer products. The method may include obtaining, by the computing system, object-specific information for one or more objects of the plurality of objects. The object-specific information may include one or more details associated with each of the one or more objects. The method may include providing, by the computing system, one or more user interface elements overlaid on the one or more objects. In some implementations, the one or more user interface elements may describe the object-specific information.

[0008] Another example aspect of the present disclosure is directed to one or more non-transitory computer-readable media collectively storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations. The operations may include obtaining image data. The image data may be indicative of a scene. The operations may include processing the image data to determine a plurality of filters. The plurality of filters may be associated with a plurality of objects in the scene. In some implementations, the operations may include providing one or more particular filters of the plurality of filters for display in a user interface. The operations may include obtaining input data. The input data may be associated with a selection of a particular filter of the plurality of filters. The operations may include providing one or more indicators overlaid on the image data. The one or more indicators may describe one or more particular objects associated with the particular filter.

[0009] Other aspects of the present disclosure are directed to various systems, apparatus, non-transitory computer-readable media, user interfaces, and electronic devices.

[0010] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain related principles.

[0011] A detailed discussion of embodiments, directed to persons skilled in the art, is described herein, which refers to the accompanying drawings. [Brief description of the drawings]

[0012] [Figure 1A]FIG. 1 is a block diagram of an exemplary computing system for implementing object recognition and filtering in accordance with an exemplary embodiment of the present disclosure. [Figure 1B] FIG. 2 is a block diagram of an exemplary computing device that implements object recognition and filtering in accordance with an exemplary embodiment of the present disclosure. [Figure 1C] FIG. 2 is a block diagram of an exemplary computing device that implements object recognition and filtering in accordance with an exemplary embodiment of the present disclosure. [Figure 2A] FIG. 1 illustrates an exemplary object filtering and information display system according to an exemplary embodiment of the present disclosure. [Figure 2B] FIG. 1 illustrates an exemplary object filtering and information display system according to an exemplary embodiment of the present disclosure. [Figure 3A] FIG. 1 illustrates an exemplary object filtering and information display system according to an exemplary embodiment of the present disclosure. [Figure 3B] FIG. 1 illustrates an exemplary object filtering and information display system according to an exemplary embodiment of the present disclosure. [Figure 4A] FIG. 1 illustrates an exemplary object filtering and information display system according to an exemplary embodiment of the present disclosure. [Figure 4B] FIG. 1 illustrates an exemplary object filtering and information display system according to an exemplary embodiment of the present disclosure. [Figure 5A] FIG. 1 illustrates an exemplary question and answer dialogue according to an exemplary embodiment of the present disclosure. [Figure 5B] FIG. 1 illustrates an exemplary question and answer dialogue according to an exemplary embodiment of the present disclosure. [Figure 6] FIG. 2 is a flow chart diagram of an exemplary method for implementing object recognition and information display according to an exemplary embodiment of the present disclosure. [Figure 7] FIG. 2 is a flow chart diagram of an exemplary method for implementing object recognition and information display according to an exemplary embodiment of the present disclosure. [Figure 8]1 is a flow chart diagram of an exemplary method for performing object filtering in accordance with an exemplary embodiment of the present disclosure. [Figure 9] 1 illustrates an exemplary zoom interaction according to an exemplary embodiment of the present disclosure. [Figure 10A] FIG. 2 illustrates an exemplary mobile map application usage according to an exemplary embodiment of the present disclosure. [Figure 10B] FIG. 2 illustrates an exemplary mobile map application usage according to an exemplary embodiment of the present disclosure. [Figure 11] FIG. 1 illustrates an exemplary book filtering based on ratings according to an exemplary embodiment of the present disclosure. [Figure 12] 1 illustrates an exemplary object specific information display according to an exemplary embodiment of the present disclosure. [Figure 13] 1 illustrates an exemplary object specific information display according to an exemplary embodiment of the present disclosure. [Figure 14] FIG. 1 illustrates an exemplary book filtering based on ratings according to an exemplary embodiment of the present disclosure. [Figure 15] 1 illustrates an exemplary object-specific search user interface according to an exemplary embodiment of the present disclosure. [Figure 16] 1A-1C illustrate exemplary user interface elements according to an exemplary embodiment of the present disclosure. [Figure 17] 1A-1C illustrate exemplary user interface elements according to an exemplary embodiment of the present disclosure. [Figure 18] 1A-1C illustrate exemplary user interface elements according to an exemplary embodiment of the present disclosure. [Figure 19] 1A-1C illustrate exemplary user interface transitions according to exemplary embodiments of the present disclosure. [Figure 20] FIG. 1 illustrates an exemplary focus interaction according to an exemplary embodiment of the present disclosure. [Figure 21]1A-1C illustrate exemplary user interface elements according to an exemplary embodiment of the present disclosure. [Figure 22] 1A-1C illustrate exemplary user interface elements according to an exemplary embodiment of the present disclosure. [Figure 23] 1A-1C illustrate example toggle elements for turning object tagging on and off, according to an example embodiment of the present disclosure. [Figure 24] FIG. 2 illustrates an example rating filtering element for rating-based filtering, according to an example embodiment of the present disclosure. [Diagram 25] 1 illustrates an example rating filtering slider element for filtering based on ratings, according to an example embodiment of the present disclosure. [Figure 26] FIG. 2 illustrates an exemplary search interface according to an exemplary embodiment of the present disclosure. [Figure 27] FIG. 2 is a block diagram of an exemplary tag generation model, according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] Reference numbers repeated among the figures are intended to identify like features in various implementations.

[0014] overview In general, the present disclosure is directed to systems and methods for providing object-specific information via an augmented reality overlay. In particular, the systems and methods disclosed herein can leverage image processing techniques (e.g., object detection, optical character recognition, reverse image search, image segmentation, video segmentation, etc.) and augmented reality rendering to provide a user interface that overlays object-specific details on objects shown in image data. For example, the systems and methods disclosed herein can be used to obtain image data, process the image data to understand a scene, and provide details about the scene through an augmented reality experience. In some implementations, the systems and methods disclosed herein can provide suggested filters or candidate queries that can be used to provide more information about the recognized object. Additionally and / or alternatively, object-specific information (e.g., a rating or composition for a particular object) can be obtained and overlaid on an image of the object. For example, the systems and methods can include obtaining image data generated by a mobile image capture device. The image data can depict a scene. The image data can be processed (e.g., with one or more machine learning-based models stored locally on the device) to determine multiple objects in the scene. In some implementations, the plurality of objects may include one or more consumer products (e.g., products for sale at a grocery store (e.g., coffee, chocolate, soda, books, toothpaste, etc.)). The systems and methods may include obtaining object-specific information for one or more objects of the plurality of objects. The object-specific information may include one or more details associated with each of the one or more objects. The systems and methods may include providing one or more user interface elements overlaid on the image data. In some implementations, the one or more user interface elements may describe the object-specific information.

[0015] In particular, a user may open a mobile application. The user may capture one or more images using an image sensor on the mobile device. The images may be processed with a machine learning model stored on the mobile device to determine one or more tags (e.g., one or more queries and / or one or more filters). The tags may be provided to the user via a user interface. The user may select a particular tag, which may cause the user interface to provide an augmented reality experience that includes object-specific information overlaid on a particular object in the captured image.

[0016] The systems and methods can acquire image data (e.g., one or more images associated with a scene and / or multiple image frames including a first image frame and a second image frame). In some implementations, the image data can include video stream data (e.g., a live stream video of the scene). The video stream data can include multiple image frames. The image data (e.g., the video stream data) can be generated by a mobile image capture device (e.g., a mobile computing device with an image sensor). In some implementations, the image data can be indicative of a scene.

[0017] In some implementations, the image data may include multiple frames. The multiple frames may be processed to determine that a first image frame and a second image frame are associated with a scene. The first image frame may include a first set of objects, and the second image frame may include a second set of objects. Determining that the first image frame and the second image frame are associated with a scene may include determining that the first set of objects and the second set of objects are associated with a particular object class.

[0018] Alternatively and / or additionally, determining that the first image frame and the second image frame are associated with the scene may include determining that the first image frame and the second image frame were captured at a particular location. The particular location may be determined based on a time between the image frames that is less than a threshold time. Alternatively and / or additionally, the location may be determined based on one or more location sensors (e.g., a global positioning system on a mobile computing device). In some implementations, determining that the image frames are associated with each other may include processing the multiple image frames with one or more machine learning models (e.g., an image classification model, an image segmentation model, an object classification model, an object recognition model, etc.). The one or more machine learning models may be trained to determine a semantic understanding of the image frames based on context and / or features detected in the scene.

[0019] In some implementations, multiple image frames may be associated with one another based on determining that the image frames capture overlapping portions of a scene and / or determining that the image frames capture portions of a scene that are proximate to one another. Systems and methods may use various techniques to determine that the image frames show different portions of the same scene. Various techniques may include image analysis (e.g., pixel-by-pixel analysis), timestamp analysis (e.g., comparing metadata associated with the image frames), and / or motion data analysis (e.g., acquiring and processing motion sensor data (e.g., inertial data from an inertial motion sensor)).

[0020] The acquisition and / or generation of image frames may occur in response to input data received from a user. The input data may include text data, user interface selections, audio data (e.g., audio data describing a voice command), or another form of input. Additionally and / or alternatively, image frame association may be facilitated based in part on the received input (e.g., user interface selections, touch screen interactions, text input, voice commands, and / or gestures).

[0021] In some implementations, the system and method may include generating scene data based on a first image frame and a second image frame of the plurality of image frames. The scene data may include and / or describe the first image frame and the second image frame. In some implementations, generating the scene data may include stitching the image frames together. Alternatively and / or additionally, the image frames may be concatenated. The stitched image frames may then be cropped to remove data that may not be relevant to semantic understanding of the scene. The stitched frames may be provided for display. Alternatively and / or additionally, the stitched frames may be used only for scene understanding in a backend.

[0022] The image data may be processed to determine a plurality of objects in a scene. In some implementations, the plurality of objects may include one or more consumer products. Alternatively and / or additionally, the scene data may be processed to determine a plurality of objects in a scene, the plurality of objects may include a plurality of consumer products (e.g., food, appliances, soap, tools, etc.). The image data and / or scene data may be processed to understand the scene. Processing the image data and / or scene data may include optical character recognition, object detection and recognition, pixel-by-pixel analysis, feature extraction and subsequent processing, image classification, object classification, object class determination, image segmentation, and / or environment or scene classification. In some implementations, processing may occur on-device (e.g., a mobile computing device using machine learning-based models stored on the device with limited computational resources). On-device processing may limit the resource cost of sending large amounts of data over a network to a server computing system for processing.

[0023] In some implementations, the system and method may determine objects in a scene. Additionally and / or alternatively, the system and method may determine object classes or other forms of relationships between objects. The system and method may then ignore objects that are not included in the relationships (e.g., the system and method need only process data associated with objects of a particular object class). In some implementations, multiple object classes may be determined. The system and method may determine more generalized object classes and / or focus on more useful use cases. Alternatively and / or additionally, the system and method may focus on objects associated with an object class in an earlier search. In some implementations, the system and method may include biases based on user preferences or past user interactions.

[0024] In some implementations, the system and method can determine tags associated with multiple object classes and narrow down to a particular object class based on a selection. The system and method can focus on one or more objects in a reticle of an image capture interface or in a focus of a scene. Alternatively and / or additionally, the system and method can focus on determined user preferences and / or determined regional or global preferences. The preferences and tastes can be learned using a machine learning model. The machine learning model can be trained to generate a probability score associated with the processed image data and the processed contextual data. One or more tags can then be selected based on the probability score (e.g., the highest probability score and / or a probability score above a given threshold can be selected).

[0025] A number of tags (e.g., candidate queries, filters, and / or annotations) may be generated based on the determined scene understanding. The tags may include candidate queries, which may include questions asked by other users when having a similar context, questions associated with a particular object class (e.g., food ingredients or book genres), questions associated with a particular detected object, questions associated with a particular location (e.g., grocery store or museum), and / or questions associated with past user interactions (e.g., what the user asked during a previous outing to this location, what are the frequently asked questions by the user, and / or the user's browsing history regarding this location or object class). In some implementations, the tags (e.g., filters, candidate queries, and / or annotations) may include data associated with a user profile, including the user's preferences. The user profile may include allergies, which may be used as context data when the object is a food item. Additionally and / or alternatively, user preferences may include genre preferences (e.g., book genres such as young adult or romance), taste preferences (e.g., sweet vs. salty, and / or citrus vs. earthy), and / or ingredient preferences (e.g., limits on certain percentages of particular ingredients and / or numbers of ingredients).

[0026] The tags, or chips, may be determined and / or selected such that each tag may apply to at least one object in the scene. Additionally and / or alternatively, tags may not be selected that apply to all objects. Tags may be generated and / or determined based on determined salient features between objects in the scene (e.g., tags may include ingredients or flavor notes that differ between objects in the scene).

[0027] The systems and methods can determine one or more tags of a plurality of tags (e.g., one or more candidate queries of a plurality of candidate queries) based on image data and / or scene data. In some implementations, the one or more tags can be determined based at least in part on the obtained context data. The tags can be ranked and / or selected based on scene context, location, data associated with a particular user, and / or tag popularity among multiple users. Popularity can include popularity over all time or for a given time period (e.g., trending tags). The determination of the one or more tags can include user-specific refinements. In some implementations, upon determination, the systems and methods can only indicate annotations or tags for high value items.

[0028] Additionally and / or alternatively, the systems and methods can obtain object-specific information for one or more objects of the plurality of objects. The object-specific information may include one or more details associated with each of the one or more objects. In some implementations, the object-specific information may include one or more consumer product details associated with each of the plurality of objects.

[0029] In some implementations, the systems and methods may include obtaining context data. The context data may be associated with a user. A query may then be determined based on the image data and the context data. The object specific information may be obtained based at least in part on the query. In some implementations, the context data may describe at least one of a user location, a user preference, past user queries, and / or a user shopping history.

[0030] In some implementations, the contextual data can describe a user location. For example, the systems and methods can obtain one or more popular queries associated with the user location. The query can then be determined based at least in part on the one or more popular queries.

[0031] Alternatively and / or additionally, an object class associated with the plurality of objects may be determined. Object-specific information may then be obtained based at least in part on the object class.

[0032] The systems and methods may include providing one or more user interface elements overlaid on image data. The one or more user interface elements may describe object-specific information. In some implementations, the one or more user interface elements may be provided overlaid on one or more objects. The one or more user interface elements may describe object-specific information associated with the object on which the element is overlaid.

[0033] In some implementations, the user interface elements can describe object-specific information associated with one or more objects, and the user interface elements can be associated with multiple consumer products.

[0034] The one or more user interface elements can include and / or describe multiple product attributes associated with a particular object in the scene. The multiple product attributes can include multiple different product types. For example, the systems and methods can obtain input data associated with a selection of a particular user interface element associated with a particular product attribute (e.g., the particular product attribute can include a threshold product rating) and can provide one or more indicators overlaid on the image data. The one or more indicators can describe one or more particular objects associated with the one or more particular product attributes. In some implementations, the particular user interface element can include a slider associated with a range of consumer product ratings.

[0035] In some implementations, providing one or more user interface elements overlaid on the one or more objects may include adjusting a number of pixels associated with an exterior region surrounding the one or more objects, The pixel adjustments may be used to provide a spotlight effect that may indicate objects that meet criteria associated with a selected tag.

[0036] The system and method may provide one or more user interface elements as part of the augmented reality experience. For example, one or more tags may be provided as user interface elements at the bottom of a display overlaid on one or more image frames. Additionally and / or alternatively, the one or more user interface elements may include text or icons overlaid on a particular object. For example, product attributes associated with a particular object may be anchored to the object in the augmented reality experience. The user interface elements may include callouts at the bottom of the user interface and / or text anchored to the object.

[0037] Alternatively and / or additionally, the systems and methods may acquire image data. The image data may represent a scene. The image data may be processed to determine a plurality of filters. The plurality of filters may be associated with a plurality of objects in the scene. One or more particular filters of the plurality of filters may then be provided for display in the user interface. The systems and methods may then acquire input data. In some implementations, the input data may be associated with a selection of a particular filter of the plurality of filters. The systems and methods may then provide one or more indicators overlaid on the image data. The one or more indicators may describe one or more particular objects associated with the particular filter.

[0038] In some implementations, processing the image data to determine a plurality of filters may include processing the image data to recognize a plurality of objects in the scene, determining a plurality of differentiating attributes associated with differentiating factors between the plurality of objects, and determining the plurality of filters based at least in part on the plurality of differentiating attributes.

[0039] Additionally and / or alternatively, processing the image data to recognize multiple objects in the scene may include processing the image data with a machine learning based model.

[0040] The system and method may obtain second input data. The second input data may be associated with a zoom input. In some implementations, the zoom input may be associated with one or more particular objects. The system and method may then obtain second information associated with the one or more particular objects. An augmented image may be generated based at least in part on the image data and the second information. The augmented image may include a zoomed-in portion of the scene associated with an area that includes the one or more particular objects. In some implementations, the one or more indicators and the second information may be overlaid on the one or more particular objects.

[0041] Additionally and / or alternatively, the one or more indicators may include object specific information associated with one or more particular objects. In some implementations, providing the one or more indicators overlaid on the image data may include an augmented reality experience.

[0042] For example, the system and method may determine a plurality of filters associated with a plurality of objects. Each filter may include criteria associated with a subset of the plurality of objects. The plurality of filters may be provided for display in a user interface. The system and method may then obtain a filter selection associated with a particular filter of the plurality of filters. An augmented reality overlay over one or more image frames may then be provided. The augmented reality overlay may include one or more user interface elements being provided over each object that meets the respective criteria of the particular filter.

[0043] In some implementations, the systems and methods may include receiving audio data. The audio data may describe the voice command. The systems and methods may include determining a particular object associated with the voice command and providing an augmented image frame showing the particular object associated with the voice command. Additionally and / or alternatively, the acquired audio data may describe the voice command, which may be processed with one or more images to generate an output. For example, a multimodal query may be acquired that includes one or more captured images and audio data describing the voice command (e.g., one or more images of a scene with a voice command of "What cereal is it from anyway?"). The multimodal query may be processed to generate a response to the voice command that is determined based at least in part on the one or more images. In some implementations, the response may include one or more user interface elements overlaid on the captured images and / or a live stream of images in a viewfinder. The voice input in tandem with the camera input may provide a conversational assistant that is visually aware of the environment, which may enable the user to be informed of the environment as they navigate through it. In some implementations, the processing of image data may be conditioned based on a voice command. For example, an image may be cropped based on a voice command to segment points or points of interest, which may then be processed. Additionally and / or alternatively, voice input and image input may be input and processed in tandem.

[0044] In some implementations, a user may capture an image of an object and may give a voice command to request information about the particular object. The requested information may include asking about the status of the particular object. For example, a user may capture an image of a pear and may give the voice command "Is this ripe?" The systems and methods disclosed herein may process the image and voice command to determine that a ripeness classification should be provided. The systems and methods may then process the image of the pear to output a ripeness classification, which may then be provided to the user. In some implementations, data describing the ripeness determination of the pear and / or data describing nutritional or growing information of the pear may further be provided.

[0045] The voice input may be processed to generate text data describing the voice command, and the voice command may be processed with the image data for search result determination. Text embeddings may be generated based on the transcribed voice command, and image embeddings may be generated based on the captured image, and the text embeddings and image embeddings may be processed to determine one or more search results.

[0046] The systems and methods disclosed herein may involve obtaining one or more inputs from a user. User input may include selection of a particular tag associated with a particular candidate query, text input (e.g., that may be used to generate new queries and / or new filters), voice input, and / or adjustment of a filter slider (e.g., for a price or rating).

[0047] In some implementations, the systems and methods disclosed herein can be used to filter objects in a scene to answer a question and / or determine one or more specific objects that meet one or more criteria. For example, the systems and methods disclosed herein can obtain image data, determine a plurality of objects shown in the image data, and determine one or more objects in the scene that are relevant or associated with the candidate query (e.g., have a given product attribute and / or meet the input criteria). The determination can involve searching the web. The search can include extracting data from a knowledge graph, a local database, a regional database, a global database, a web page, and / or data stored on a processing device. The systems and methods can further obtain object details associated with the object associated with the selected candidate query.

[0048] Additionally and / or alternatively, a user interface may be provided that indicates which objects are associated or relevant to a selected candidate query. The user interface may highlight particular objects associated with a selected tag (e.g., a candidate query and / or filter). In some implementations, the systems and methods may dim pixels that are not associated with a particular object. Additionally and / or alternatively, the systems and methods may provide indicators overlaid on particular objects. The indicators may include object-specific details (e.g., ingredients, flavor notes, ratings, genre, etc.).

[0049] The user interface may include an augmented reality experience. A user interface including an augmented reality experience may be provided as part of a mobile application, a web application, and / or as part of an integrated system for a smart wearable. The systems and methods disclosed herein may be implemented in an augmented reality application including augmented reality translation, object recognition, and / or various other features. Alternatively and / or additionally, the systems and methods disclosed herein may be implemented as a standalone application. Additionally and / or alternatively, the systems and methods disclosed herein may be used by a smart wearable, such as smart glasses, to learn about different scenes and objects while navigating daily routines.

[0050] In some implementations, the systems and methods disclosed herein can be always on and / or can be toggled on and off. The systems and methods can be provided in an application with multiple tabs associated with multiple different functions. The currently open tab during processing can be used as a context for determining one or more tags.

[0051] The systems and methods disclosed herein can use a number of different user interface / user experience features and elements. The elements can include two-dimensional shapes, three-dimensional shapes, text, popups, dynamic elements, input boxes, graphical keyboards, magnifying elements, transition effects, reticles, shading effects, and / or processing indicators. Tags can be below the user interface, above the user interface, and / or to the side. Annotations can be overlaid on objects, placed above or below objects, and / or indicated by symbols, icons, or indicators. In some implementations, the systems and methods can include off-screen indicators that indicate that an object in a scene meets a given criterion or has certain details but is not currently displayed in the user interface. Additionally and / or alternatively, the user interface can include an artificial spotlight feature that is used to indicate objects that meet a given criterion associated with a selected filter or query.

[0052] The systems and methods disclosed herein can be utilized in a variety of different ways. For example, the systems and methods can be used to narrow and select objects in a scene that meet various criteria. In some implementations, the narrowing can be used to select consumer products based on ratings, ingredients, and / or attributes.

[0053] Additionally and / or alternatively, the systems and methods disclosed herein may be used to determine and provide object differentiators for different objects in a scene.

[0054] The systems and methods can be used to provide instructions on how to interact with a scene (eg, car maintenance and / or use of a particular device, such as a blender).

[0055] In some implementations, the systems and methods can be used for shopping (eg, to avoid allergenic ingredients and / or for symptom-based filtering when shopping for medicines).

[0056] Additionally and / or alternatively, the systems and methods can determine and provide information about relevant objects based on scene analysis.

[0057] In some implementations, the systems and methods disclosed herein can generate and / or determine tags such that tags can be automatically generated based on what is in the scene and / or based on the context to provide tags for what a user may be asking about. The systems and methods disclosed herein can process scene data to determine what is the search query or filter that will be most insightful into the scene and / or what separates different objects from each other. For example, an image of a bag of coffee in a shopping aisle can cause the system to automatically generate tags for flavor profile, rating, local, fair trade, etc., and an image of a book can cause the system to automatically generate tags for genre, rating, length, time period, etc. Additionally and / or alternatively, an image of a shopping mall can cause the system to automatically generate tags for restaurants, clothing, chain businesses, local entities, open, etc.

[0058] Additionally and / or alternatively, the tag (e.g., filter and / or candidate query) may include determining a number of candidate tags associated with the image data and / or the contextual data. The number of candidate tags may then be processed to limit the displayed tags to those that (1) are associated (e.g., true) with at least one object in the scene and (2) are not associated with all objects in the scene. Limiting the candidate query based on one or both of the factors can ensure that the selection of tags provides actual information to the user, rather than leaving the user with the same options originally given when capturing the image.

[0059] A selection of a particular object and / or a tag associated with the particular object may be received, and additional information about the particular object may be obtained and displayed. For example, a selection of a particular product may be received, and additional product details may be obtained and displayed. The additional information may be based in part on one or more past user interactions (e.g., purchase history, search history, and / or previously selected filter tags). The additional information may be obtained by using the image data and / or recognition data as a search query to determine one or more search results that may be displayed and / or processed to determine the additional information. The search query may further include text input, voice input, and / or contextual data (e.g., location, other objects in the scene, time, user profile data, and / or image classification).

[0060] In some implementations, the systems and methods disclosed herein can be used to capture (generate or acquire) and process video. The video may be captured and then processed to detect and recognize one or more objects in the video, which may then be annotated during playback. Additionally and / or alternatively, actions performed in the video may be determined and annotated during playback. In some implementations, one or more objects in the video may be segmented and then searched for. Additionally and / or alternatively, annotations may be determined and provided in real time, which may then be provided as augmented reality annotations.

[0061] The systems and methods of the present disclosure provide several technical effects and benefits. As an example, the systems and methods can provide a real-time augmented reality experience that can provide a user with scene understanding. In particular, the systems and methods disclosed herein can acquire image data, process the image data, recognize objects shown in the image data, and provide object-specific information about those objects. Additionally and / or alternatively, the systems and methods can process the image data and provide tags (e.g., filtering tags for filtering objects in the scene and / or query tags for obtaining specific information associated with the objects). The tags can then be selected, and the systems and methods disclosed herein can provide an indicator anchored to the particular object in the image data. The indicator can include an augmented reality rendering that includes object-specific information about the object to which it is anchored.

[0062] Another technical benefit of the systems and methods of the present disclosure is that multimodal search can be leveraged to help a user narrow their selection or learn how to interact with the environment. For example, the systems and methods disclosed herein can be used to extract data from an image and further receive voice commands, text input, and / or user selections, which can then be used to generate queries based on both the features recognized in the image and the input data. Multimodal search can provide a more comprehensive search, which can then be used to understand the scene. For example, a user may capture an image and select one or more tags associated with the user's preferences to determine what objects the user wants. Additionally and / or alternatively, one or more of those tags may be tags entered via a graphical keyboard. Alternatively and / or additionally, a user may capture an image and ask how to complete a particular task. The systems and methods can then process the image and the entered question to provide step-by-step directions with indicators overlaid on portions of the image to give more precise instructions.

[0063] Another example of the technical effects and benefits relates to improved computational efficiency and increased capabilities of computing systems. For example, the systems and methods disclosed herein can leverage on-device machine learning models and capabilities to process locally on the device. Processing locally on the device can limit data sent over a network to a server computing system for processing, which can be more friendly to users with limited network access.

[0064] Referring now to the drawings, exemplary embodiments of the present disclosure will be discussed in further detail.

[0065] Exemplary Devices and Systems 1A illustrates a block diagram of an exemplary computing system 100 for implementing object recognition and filtering in accordance with an exemplary embodiment of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150, communicatively coupled via a network 180.

[0066] The user computing device 102 may be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0067] The user computing device 102 includes one or more processors 112 and memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple processors operatively connected. The memory 114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 may store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0068] In some implementations, the user computing device 102 can store or include one or more machine learning based models 120 (e.g., one or more machine learning based tag generation models). For example, the machine learning based models 120 can be or otherwise include various machine learning based models, such as neural networks (e.g., deep neural networks) or other types of machine learning based models including nonlinear and / or linear models. The neural networks can include feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Exemplary machine learning based models 120 are discussed with reference to FIGS. 2A-5B and 9-26.

[0069] In some implementations, one or more machine learning based models 120 may be received from the server computing system 130 over the network 180, stored in the user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, the user computing device 102 may implement multiple parallel instances of a single machine learning based model 120 (e.g., to perform parallel object recognition and tag generation across multiple instances of object recognition and filtering).

[0070] More specifically, a machine learning model (e.g., a tag generation model) can process the image data to recognize multiple objects in a scene shown in the image data. The machine learning model (e.g., a tag generation model) can determine tags based at least in part on the multiple objects and the context data. Tags can be generated based on a determined general object class, based on previous interactions, based on location, and / or based on comparing details among multiple objects to determine salient features. Tags can include queries or filters. Tags can then be selected to filter objects shown as meeting certain criteria.

[0071] Additionally or alternatively, one or more machine learning based models 140 (e.g., one or more tag generation models) may be included in or otherwise stored and implemented by a server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine learning based models 140 may be implemented by the server computing system 130 as part of a web service (e.g., an object discovery and filter service). Thus, one or more models 120 may be stored and implemented at the user computing device 102 and / or one or more models 140 may be stored and implemented at the server computing system 130.

[0072] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component may be useful for implementing a virtual keyboard. Other exemplary user input components include a microphone, a traditional keyboard, or other means by which a user can provide user input.

[0073] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple processors operatively connected. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 may store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.

[0074] In some implementations, the server computing system 130 includes or is otherwise implemented by one or more server computing devices. In instances in which the server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0075] As described above, the server computing system 130 can store or otherwise include one or more machine learning based models 140 (e.g., one or more machine learning based tag generation models). For example, the models 140 can be or otherwise include a variety of machine learning based models. Exemplary machine learning based models include neural networks or other multi-layer nonlinear models. Exemplary neural networks include feed-forward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Exemplary models 140 are discussed with reference to FIGS. 2A-5B and 9-26.

[0076] The user computing device 102 and / or the server computing system 130 can train the models 120 and / or 140 by interacting with a training computing system 150 that is communicatively coupled via a network 180. The training computing system 150 may be separate from the server computing system 130 or may be part of the server computing system 130.

[0077] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple processors operably connected. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 may store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is otherwise implemented by one or more server computing devices.

[0078] The training computing system 150 may include a model trainer 160 that trains the machine learning based models 120 and / or 140 stored in the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as, for example, backpropagation. For example, a loss function may be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques may be used to iteratively update the parameters for several training iterations.

[0079] In some implementations, performing the error backpropagation may include performing truncated backpropagation over time. The model trainer 160 can implement several generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the trained model.

[0080] In particular, model trainer 160 may train tag generation model 120 and / or 140 based on a set of training data 162. Training data 162 may include, for example, training images, training labels (e.g., ground truth object labels and / or ground truth tags), training context data, and / or training motion data.

[0081] In some implementations, if the user provides consent, the training examples may be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 against user-specific data received from the user computing device 102. In some instances, this process may be referred to as individualizing the model.

[0082] The model trainer 160 includes computer logic used to provide the desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software controlling a general-purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions stored in a tangible computer-readable storage medium, such as a RAM, a hard disk, or an optical or magnetic medium.

[0083] Network 180 may be any type of communications network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. In general, communications over network 180 may be carried over any type of wired and / or wireless connections, using a wide variety of communications protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).

[0084] The machine learning models described herein may be used in a variety of tasks, applications, and / or use cases.

[0085] In some implementations, the input to the machine learning based model of the present disclosure may be image data. The machine learning based model may process the image data to generate an output. As an example, the machine learning based model may process the image data to generate an image recognition output (e.g., recognition of the image data, latent embedding of the image data, encoded representation of the image data, hash of the image data, etc.). As another example, the machine learning based model may process the image data to generate an image segmentation output. As another example, the machine learning based model may process the image data to generate an image classification output. As another example, the machine learning based model may process the image data to generate an image data modification output (e.g., alteration of the image data, etc.). As another example, the machine learning based model may process the image data to generate an encoded image data output (e.g., encoded and / or compressed representation of the image data, etc.). As another example, the machine learning based model may process the image data to generate an upscaled image data output. As another example, the machine learning based model may process the image data to generate a prediction output.

[0086] In some implementations, the input to the machine learning based models of the present disclosure may be text or natural language data. The machine learning based models may process the text or natural language data to generate an output. As an example, the machine learning based models may process the natural language data to generate a language encoding output. As another example, the machine learning based models may process the text or natural language data to generate a latent text embedding output. As another example, the machine learning based models may process the text or natural language data to generate a transformation output. As another example, the machine learning based models may process the text or natural language data to generate a classification output. As another example, the machine learning based models may process the text or natural language data to generate a text segmentation output. As another example, the machine learning based models may process the text or natural language data to generate a semantic intent output. As another example, the machine learning based models may process the text or natural language data to generate a prediction output.

[0087] In some implementations, the input to the machine learning based model of the present disclosure may be voice data. The machine learning based model may process the voice data to generate an output. As an example, the machine learning based model may process the voice data to generate a speech recognition output. As another example, the machine learning based model may process the voice data to generate a speech translation output. As another example, the machine learning based model may process the voice data to generate a latent embedding output. As another example, the machine learning based model may process the voice data to generate an encoded voice output (e.g., an encoded and / or compressed representation of the voice data, etc.). As another example, the machine learning based model may process the voice data to generate a text representation output (e.g., a text representation of the input voice data, etc.). As another example, the machine learning based model may process the voice data to generate a predicted output.

[0088] In some implementations, the input to the machine learning based model of the present disclosure may be latent coding data (e.g., a latent space representation of the input, etc.). The machine learning based model may process the latent coding data to generate an output. As an example, the machine learning based model may process the latent coding data to generate a recognition output. As another example, the machine learning based model may process the latent coding data to generate a reconstruction output. As another example, the machine learning based model may process the latent coding data to generate a search output. As another example, the machine learning based model may process the latent coding data to generate a reclustering output. As another example, the machine learning based model may process the latent coding data to generate a prediction output.

[0089] In some implementations, the input to the machine learning based models of the present disclosure may be statistical data. The machine learning based models may process the statistical data to generate an output. As an example, the machine learning based models may process the statistical data to generate a recognition output. As another example, the machine learning based models may process the statistical data to generate a prediction output. As another example, the machine learning based models may process the statistical data to generate a classification output. As another example, the machine learning based models may process the statistical data to generate a segmentation output. As another example, the machine learning based models may process the statistical data to generate a segmentation output. As another example, the machine learning based models may process the statistical data to generate a visualization output. As another example, the machine learning based models may process the statistical data to generate a diagnostic output.

[0090] In some implementations, the input to the machine learning based models of the present disclosure may be sensor data. The machine learning based models may process the sensor data to generate an output. As an example, the machine learning based models may process the sensor data to generate a recognition output. As another example, the machine learning based models may process the sensor data to generate a prediction output. As another example, the machine learning based models may process the sensor data to generate a classification output. As another example, the machine learning based models may process the sensor data to generate a segmentation output. As another example, the machine learning based models may process the sensor data to generate a segmentation output. As another example, the machine learning based models may process the sensor data to generate a visualization output. As another example, the machine learning based models may process the sensor data to generate a diagnostic output. As another example, the machine learning based models may process the sensor data to generate a detection output.

[0091] In some cases, the machine learning based model can be configured to perform a task that includes encoding input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). In another example, the input includes visual data (e.g., one or more images or videos) and the output includes compressed visual data, and the task is a visual data compression task. In another example, the task may include generating an embedding for the input data (e.g., input audio or visual data).

[0092] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data for one or more images and the task is an image processing task. For example, the image processing task may be image classification, where the output is a set of scores, each score corresponding to a different object class and representing a likelihood that one or more images show an object belonging to the object class. The image processing task may be object detection, where the image processing output identifies one or more regions in one or more images and, for each region, a likelihood that the region shows an object of interest. As another example, the image processing task may be image segmentation, where the image processing output defines, for each pixel in one or more images, a respective likelihood for each category in a predetermined set of categories. For example, the set of categories may be foreground and background. As another example, the set of categories may be object classes. As another example, the image processing task may be depth estimation, where the image processing output defines, for each pixel in one or more images, a respective depth value. As another example, the image processing task may be motion estimation, where the network input includes multiple images and the image processing output defines, for each pixel of one of the input images, the scene motion represented in pixels between the images in the network input.

[0093] In some cases, the input includes audio data representing speech and the task is a speech recognition task. The output may include text output that is mapped to the speech. In some cases, the task includes encrypting or decrypting input data. In some cases, the task includes a microprocessor-implemented task such as branch prediction or memory address translation.

[0094] 1A illustrates one exemplary computing system that can be used to implement the present disclosure. Other computing systems may be used. For example, in some implementations, the user computing device 102 may include a model trainer 160 and a training dataset 162. In such implementations, the model 120 may be both trained and used locally on the user computing device 102. In some such implementations, the user computing device 102 may implement the model trainer 160 to personalize the model 120 based on user-specific data.

[0095] 1B illustrates a block diagram of an exemplary computing device 10 for implementing an exemplary embodiment of the present disclosure. The computing device 10 may be a user computing device or a server computing device.

[0096] The computing device 10 includes several applications (e.g., applications 1-N). Each application includes its own machine learning library and machine learning model. For example, each application may include a machine learning model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.

[0097] 1B, each application may communicate with several other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application may communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0098] 1C illustrates a block diagram of an exemplary computing device 50 for implementing according to an exemplary embodiment of the present disclosure. The computing device 50 may be a user computing device or a server computing device.

[0099] Computing device 50 includes several applications (e.g., applications 1-N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API across all applications).

[0100] The central intelligence layer includes several machine learning models. For example, as shown in FIG. 1C, a respective machine learning model (e.g., model) may be provided to each application and managed by the central intelligence layer. In other implementations, two or more applications may share a single machine learning model. For example, in some implementations, the central intelligence layer may provide a single model (e.g., a single model) to all of the applications. In some implementations, the central intelligence layer is included in or otherwise implemented by the operating system of the computing device 50.

[0101] The central intelligence layer can communicate with a central device data layer, which can be a centralized repository of data for computing devices 50. As shown in FIG. 1C, the central device data layer can communicate with several other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0102] Example model layout 2A and 2B are illustrations of an exemplary object filtering and information display system 200, according to an exemplary embodiment of the present disclosure. In some implementations, the object filtering and information display system 200 may include one or more machine learning based models trained to recognize objects in a captured image 210 describing a scene, and provide, as a result of the object recognition, an augmented image 250 including one or more user interface elements overlaid on the objects that meet the filtering criteria. Thus, in some implementations, the object filtering and information display system 200 may include an intermediate augmented image 230 operable to show objects in the scene that meet a first criterion.

[0103] As shown in Figures 2A and 2B, the systems and methods disclosed herein may be provided as native applications, mobile applications, and / or web applications running on a mobile computing device 212. The mobile computing device 212 may include one or more stored and / or downloaded machine learning-based models for image processing to determine the plurality of tags 222. The mobile computing device 212 may include one or more processors and may be configured to provide a user interface as disclosed herein. For example, the mobile computing device 212 may include a display screen configured to display the user interface, and the display may include displaying one or more images captured by an image sensor (e.g., an image capture device of the mobile computing device). Alternatively and / or additionally, the systems and methods disclosed herein may be implemented in a smart wearable (e.g., smart glasses).

[0104] In particular, a user can open a mobile device application that can be used to capture one or more images 210 of a scene (e.g., a grocery store aisle including multiple coffee options to choose from). The images can be processed to determine that multiple different coffees are in the scene and that the scene is primarily of coffee class objects. Based on the recognition of the multiple coffees and / or based on the determined coffee class, the object filtering and information display system 200 can generate multiple tags 222 associated with flavor profiles for the different coffees and provide the tags 222 (e.g., citrus, earthy, and fruity) for display 220. The user can then select a particular tag (e.g., citrus). The object filtering and information display system 200 can obtain object-specific information for each of the coffees in the scene to determine which coffees have a flavor profile associated with the selected particular tag. The objects (i.e., coffees) having the particular flavor profile (i.e., citrus) can then be shown in the user interface 230. The instructions may include one or more user interface elements overlaid on the particular object and / or may include highlighting the particular object and blurring the surrounding area.

[0105] The object filtering and information display system 200 may then determine one or more new tags (e.g., local and LGBTQ owned) while continuing to provide the selected tags for display 240. The user may then select a second tag (e.g., local). The object filtering and information display system 200 may determine which of the objects meet the first criteria of the first tag and the second criteria of the second tag. The one or more objects that meet both criteria may then be shown in one or more user interface elements and may be highlighted at 250.

[0106] In some implementations, the indicators and highlighting may occur on live stream image data that may differ from the originally processed image data. For example, annotations, tags, and user interface elements may be provided as part of an augmented reality experience that anchors user interface elements and effects to objects in a scene, so that as the camera moves, the user interface elements can accompany the associated object.

[0107] 3A and 3B are diagrams of an exemplary object filtering and information display system 300, according to an exemplary embodiment of the present disclosure. The object filtering and information display system 300 is similar to the object filtering and information display system 200 of FIGS. 2A and 2B, and further includes text input.

[0108] In particular, the object filtering and information display system 300 can capture one or more images of a scene 310 (e.g., a food aisle containing a plurality of different objects (e.g., different chocolates). The one or more images can be processed to recognize the plurality of objects. Object-specific information for each of the plurality of objects (e.g., a rating for a particular chocolate) can then be obtained. Text associated with the object-specific information can then be overlaid on the respective object at 320. Additionally and / or alternatively, a plurality of tags (e.g., fair trade, organic, and local) can be determined based on the recognized object, the object's object class, and / or contextual data (e.g., location, user profile, etc.). At 330, the plurality of tags can be provided for display, and a particular tag (e.g., fair trade) can be selected. The object filtering and information display system 300 can determine the objects associated with the particular tag, and can show the objects that do or do not have an association with the tag (e.g., whether the object was produced and sold in fair trade). A check mark can then be provided next to the text of the selected tag. The user may then select a second tag, such as a text input tag, at 340 to open a text input interface for generating a new tag. The text input interface may include a graphical keyboard, and the user may enter a new filter or candidate query (e.g., 72% dark) at 350. The input text may then be searched along with the recognized objects to determine which of the objects are associated with the particular text input. Objects that meet the criteria of the first tag and are associated with the text input may then be indicated in the user interface with a spotlight feature at 360.

[0109] 4A and 4B are diagrams of an exemplary object filtering and information display system according to an exemplary embodiment of the present disclosure. The object filtering and information display system 400 is similar to the object filtering and information display system 200 of FIGS. 2A and 2B and the object filtering and information display system 300 of FIGS. 3A and 3B.

[0110] For example, one or more images may be obtained and processed. In some implementations, once the one or more images are processed, a processing interface effect may be provided at 410. A plurality of objects in the scene may be recognized and a rating for each of the objects may be obtained. Additionally and / or alternatively, a plurality of tags may be determined based on the image and / or contextual data. The user interface may then provide a rating overlaid on each object at 420 with the tag provided for selection at the bottom of the interface. A tag may be selected and the objects may be filtered to determine a particular object that meets a particular criterion. The particular object may then be indicated at 430 by removing a rating from the object that does not meet the criterion. A second tag may be selected and a second filtering may occur. The user interface may be updated at 440 to remove a rating from the object that does not meet the first criterion and the second criterion. A third tag may be selected and a third filtering may occur. The user interface may be updated at 440 to remove a rating from the object that does not meet the first criterion, the second criterion, and the third criterion.

[0111] In some implementations, determining that an object meets certain criteria may involve obtaining object-specific information about a particular object, parsing the information into one or more segments, and processing the segments to determine a particular segment classification (e.g., that the segment is related to a flavor, an ingredient, a source, a location, etc.). The systems and methods disclosed herein can then process the segments and the given criteria to determine whether there is an association. The processing can involve natural language processing and can involve determining whether one or more segments are associated with the given criteria (e.g., whether the segment includes a word match) based on one or more knowledge graphs, or describing the given criteria (e.g., the segment states "citrus" or a synonym of citrus, and the criteria is an item that has a citrus flavor).

[0112] Alternatively and / or additionally, the object-specific information may include pre-built indexed data into one or more information categories (e.g., ratings, calories, flavors, usage, ingredients, emissions, etc.) The object-specific information may then be crawled when checking for keywords or information associated with the selected tag.

[0113] In some implementations, objects may be associated with a particular tag before the tag is provided for display. For example, multiple objects may be identified and multiple respective object-specific information sets may be obtained. The object-specific information sets may be analyzed and processed to generate a profile set for each object. The profile sets may be compared to each other to determine differentiating attributes between the objects. The differentiating attributes may be used to generate tags that narrow the list of objects. Objects with particular differentiating attributes may be pre-associated with tags such that when a tag is presented and selected, the system and method can automatically highlight or show the particular object associated with that particular tag.

[0114] Additionally and / or alternatively, the object-specific information may include one or more predefined tags indexed in a database and / or knowledge graph. In response to obtaining the object-specific information, the system and method may determine which tags are universal to all objects in the scene and prune those tags. The remaining predefined tags may be offered for display and selection. Once a tag is selected, the system and method may then present each of the objects that includes an indexed reference to the particular predefined tag.

[0115] In some implementations, the tag or tags can be selected so as not to disrupt the user experience. The tag or tags can be based on search queries by other users when searching for a given object class or a particular object. In some implementations, the system and method can store and retrieve data related to the first and last searches associated with a particular object and a particular object class. Additionally and / or alternatively, the search query data of a particular user or users can be indexed with the user's location at the time of a given query or filter. The data can then be used to determine tags for the particular user or other users. The tag or tags can be generated to predict what a user may want to know about a scene, environment, and / or object. The system and method can generate tags based on what the user should search to reach a final action (e.g., a purchase selection, a DIY step, etc.).

[0116] 5A and 5B show illustrations of an exemplary question and answer dialogue according to an exemplary embodiment of the present disclosure. In particular, image data may be acquired. The image data may be processed to determine that the image data describes an engine compartment of a vehicle. Different parts of the vehicle may be identified and annotated in the augmented reality interface 500. For example, a dipstick 502, an engine 504, and a battery 508 may be identified. Additionally and / or alternatively, a positive terminal 506+ and a negative terminal 510 of the battery may be annotated. An input may be received describing a question 554. The question 554 may be determined and provided for display in the augmented reality interface 500. A response to the question 554 may then be determined. The response may include an annotation 552 of an object in the scene that is associated with an answer to the question.

[0117] In some implementations, the question and answer dialogue can be used for DIY projects (e.g., car maintenance, home repairs, and / or everyday activities). Alternatively and / or additionally, the question and answer dialogue can be used to answer questions about the environment in which the user is currently located.

[0118] FIG. 9 shows an illustration of an exemplary zoom interaction 900 according to an exemplary embodiment of the present disclosure. In particular, in some implementations, the systems and methods disclosed herein can provide even more information when an object becomes a relatively large part of the image (e.g., by zooming or moving toward the object). In FIG. 9, a first instance 910 shows one book displayed in its entirety, with detailed information about the object overlaid on the one book. A second instance 920 can show two books displayed in their entirety, with detailed information about the object overlaid on each book. A third instance 930 can show four books displayed in their entirety, but with only a rating overlaid on each book. A fourth instance 940 can show nine books displayed in their entirety, with only a rating overlaid on each book. A fifth instance 950 can include multiple books displayed in their entirety. In response to multiple books with a relatively small part of the image being used for each book, the user interface can remove details until a zoom input is received or until a selection input is received. A zoom interaction interface can enable a user to receive more and more information about objects in the environment by zooming in on an image.

[0119] 10A and 10B show an illustration of an exemplary mobile map application usage according to an exemplary embodiment of the present disclosure. In particular, the systems and methods disclosed herein can be implemented in a map application to inform a user of information associated with different locations. For example, a user can open a map application 1010 and select an augmented reality experience user interface element (e.g., a "what's nearby" user interface element) to open an augmented reality experience. Image data can then be continuously acquired from an image sensor. The image data can be processed to determine which stores, restaurants, landmarks, and / or monuments are depicted in the image data. One or more annotation user interface elements can be generated to label the recognized location. The recognized location data can be processed using a machine learning model to determine one or more suggested tags (e.g., differentiator tags) to provide as user interface elements that can narrow the recognized location. The augmented reality experience can include an initial interface 1020 with an image stream, a location indicator, and multiple tags for selection (e.g., restaurants, coffee, shopping, etc.). The tags may be determined based on the processed image data, may be predetermined, may be determined based on the location, a number of user interface elements (e.g., annotations (e.g., text and / or icons) for the depicted buildings and monuments), and / or may be determined based on various other data. A selection of a particular tag (e.g., a restaurant tag) may be received and a first filtered interface 1030 may be provided that includes the image stream, location indicators, and filtered annotations for the buildings and monuments associated with the selected tag. New tags may be provided to further filter the identified buildings and monuments. The new tags may be determined by determining one or more differentiating factors among the remaining recognized locations.

[0120] A second tag (e.g., an American restaurant tag) and a location (e.g., a building or monument) associated with the second tag may be selected. The second filtered interface 1040 may include an image stream, a location indicator, the selected second tag, annotations about the determined location, and a detailed information user interface element (e.g., a bubble that may provide details about the location's name, rating, distance, and / or opening hours). The location user interface element may be selected and a directional interface 1050 may be provided. The directional interface 1050 may be interacted with to reopen the routing and directions portion of the map application with route information for getting to the location.

[0121] FIG. 11 shows an illustration of an exemplary book filtering based on ratings, according to an exemplary embodiment of the present disclosure. As shown, an image capture interface 1110 can be opened and used to capture an image. The image can be processed to recognize objects in the image. Object-specific information about the object can be obtained and used to generate a plurality of respective user interface elements for the plurality of objects. An annotation interface 1120 can be provided, where the objects in the image are annotated with a plurality of respective user interface elements. A particular object can be selected and a detail bar interface 1130 can be provided. The detail bar interface 1130 can include the particular object shown with a blurred peripheral portion of the image. Additionally and / or alternatively, other user interface elements can be moved to the border of the interface, and a detail bar can be provided at the bottom of the interface. The detail bar can include more detailed information about the particular object, can include a selectable element to transition to a search application, and can be configured such that an up swipe can expand the detail bar.

[0122] 12 shows an illustration of an exemplary object-specific information display according to an exemplary embodiment of the present disclosure. In particular, multiple images associated with multiple different respective objects may be acquired. The multiple images may be generated by segmenting different portions of one or more original images to segment different objects into different images. Alternatively and / or additionally, the multiple images may be generated separately using one or more image sensors.

[0123] In some implementations, a plurality of images may be selected from a set of images. A user may select a plurality of images for processing via a selection interface 1210 that displays thumbnails for the set of images. The selected images may be processed to recognize objects in the images, and object-specific information associated with the objects may be obtained for each object. An object-specific detail interface may then be provided that may display a first details panel 1220 associated with the object of the first image. In some implementations, the object-specific detail interface may include a carousel of thumbnails with rating indicators associated with a plurality of objects in the plurality of images. A thumbnail may be selected, which may then cause an associated image to be displayed along with information about the object in the associated image. For example, a second thumbnail may be selected, and a second details panel 1230 may be provided while displaying the carousel and the second image. Alternatively and / or additionally, the images may be navigated by swipe gestures and / or various other inputs. In some implementations, the interface may include an automatic navigation that displays each image and details panel for a given period of time.

[0124] FIG. 13 shows an illustration of an exemplary object-specific information display according to an exemplary embodiment of the present disclosure. FIG. 13 may use a user interface similar to FIG. 12. In some implementations, a user may capture a panoramic image and / or video showing multiple objects. The panoramic image and / or video may be processed to detect the objects. The objects may then be segmented from the input data to generate multiple image frames associated with the multiple objects. For example, a panoramic image may start with a first object 1310 and end with a fourth object 1320. The panoramic image may be segmented into four image frames associated with the four objects. The objects may be recognized and object-specific information may then be obtained for each of the objects. The object-specific information and the image frames may then be used to provide detailed information about the objects via an object-specific detail interface. The object-specific detail interface may include a first detail panel 1330 for the first object, a second detail panel 1340 for the second object, a third detail panel for the third object, and a fourth detail panel for the fourth object.

[0125] FIG. 14 shows an illustration of an exemplary book filtering based on ratings, according to an exemplary embodiment of the present disclosure. In particular, an image can be acquired via an image capture interface 1410. The image can be used as an image query, and multiple objects (e.g., books) can be recognized. Object-specific information (e.g., ratings) for the multiple objects can be acquired. A suggestion interface 1420 can be provided that provides at least a portion of the object-specific information overlaid on each object. A shutter user interface element (e.g., a shutter button) can be selected. The system and method can determine the focus of the image and provide more detailed information about objects in a focal region of the image via an answer interface 1430. In some implementations, the focal object can be indicated by a refined reticle. The focal region can be a central region, a region within the reticle, a region selected by user input, a region of determined user line of sight, and / or a region determined to be a focal point of the scene. In some implementations, a focal object may be annotated, and objects out of focus may remain unannotated (although unannotated objects may be detected, processed, and recognized with object details determined to be displayed if the object were to come into focus).

[0126] FIG. 15 shows an illustration of an exemplary object-specific search user interface according to an exemplary embodiment of the present disclosure. The systems and methods disclosed herein may include various user interface display alternatives for an object-specific detail interface, which may include an object-specific detail panel based on a user selection. For example, a selected object may be shown with a user interface element for each respective recognized object in the image. The first interface 1510 may include a hover user interface element with text information over the selected object, and a text information user interface element overlaid on each other object. The second interface 1520 may include a hover user interface element with text information over the selected object, and a text information user interface element for each other object at the periphery of the user interface. The third interface 1530 may include a hover user interface element with text information over the selected object, and a non-descriptive user interface element overlaid on each other object.

[0127] FIG. 16 shows an illustration of an exemplary user interface element according to an exemplary embodiment of the present disclosure. The system and method can use a variety of different user interface elements. In particular, the user interface elements can include icon-only user interface elements (e.g., 1602, 1608, and 1614), text-only user interface elements (e.g., 1604, and 1610), user interface elements with text and icons (e.g., 1616, 1606, and 1612), and user interface elements with text of different styles and sizes (e.g., 1618). The user interface elements can have different sizes and shapes. Additionally and / or alternatively, the user interface elements may have a dot, stem, or another indicator of a particular associated object.

[0128] FIG. 17 shows an illustration of exemplary user interface elements according to an exemplary embodiment of the present disclosure. In FIG. 17, a first user interface 1710 includes a plurality of recognized objects that have been annotated with ratings. A user can select a tag request icon to obtain a filter interface 1720 that the user can interact with to filter the annotations to only those of objects that meet a given criterion (e.g., objects whose ratings are above a certain threshold). A second set of tags can then be determined and provided for selection (e.g., tags associated with a genre of a particular object in a scene). One or more tags can be selected to provide a third interface 1730 that can describe annotations that are to be overlaid only on objects that meet the two criteria.

[0129] 18 shows an illustration of an exemplary user interface element according to an exemplary embodiment of the present disclosure. In some implementations, the systems and methods may include a selectable user interface element (e.g., a button) for hiding the annotation user interface element. The hide button can be provided at the bottom 1810 of the user interface, in a corner 1820 of the user interface, or at the top 1830 of the user interface.

[0130] FIG. 19 shows an illustration of an exemplary user interface transition according to an exemplary embodiment of the present disclosure. The user interface transition may include a thinking stage 1910 that may indicate that an image is being processed. Next, the user interface transition may include an annotated stage 1920 that overlays an annotation user interface element over the recognized object. A filter may then be selected and the annotation user interface element may be limited to objects associated with the selected filter to provide a filtered stage 1930. An annotation user interface element may be selected and a retrieved stage 1940 may be provided for display. In the retrieved stage 1940, an area with the selected object may be highlighted with one or more visual effects. In some implementations, a details panel (e.g., a knowledge panel) may be provided for display and may describe information associated with the selected object.

[0131] 20 shows an illustration of an exemplary focus interaction according to an exemplary embodiment of the present disclosure. In some implementations, annotation user interface elements may change appearance based on whether an object associated with the annotation is in the focus of the camera interface. For example, a first stage 2010 may include all annotation user interface elements that are semi-transparent. In a second stage 2020, the camera interface may have a single object 2002 in the reticle. The annotation user interface elements associated with the single object 2002 may then be displayed as entirely opaque.

[0132] 21 shows an illustration of an exemplary user interface element according to an exemplary embodiment of the present disclosure. Annotation user interface elements associated with a recognized object may include one or more icons 2110, text and icons in a bubble 2120, and / or a bubble in multiple text sizes 2130 with more detailed information (e.g., a rating for the object and where the rating came from). Different levels of information provided can be determined based on user preferences, one or more user selections, the number of objects being annotated, the amount of information available, the distance from the object, and / or the screen size.

[0133] In some implementations, the location and / or size of a user interface element overlay may be determined and / or adjusted based on interface display availability. For example, a user interface element may be displayed higher on an object than neighboring user interface elements to avoid overcrowding and / or overlapping elements. Alternatively and / or additionally, the amount of information and / or text size may be adjusted.

[0134] 22 shows an illustration of an exemplary user interface element according to an exemplary embodiment of the present disclosure. The user interface element may include a three-dimensional dynamic element 2210 that may rotate based on where the reticle is. Alternatively and / or additionally, the size, content, and / or size of the user interface element may change based on where the reticle is. For example, at 2220, a dot may be displayed over an object in the scene, and the dot may expand to include a text bubble when the reticle hovers over the dot. At 2230 and 2240, annotation user interface elements may be provided above the object in the augmented reality experience, rather than being overlaid on the object.

[0135] 23 shows an illustration of an example toggle element 2302 for turning object tagging on and off, according to an example embodiment of the present disclosure. In particular, a first interface 2310 may include a number of annotation user interface elements that show information about objects in a scene. The system and method may then receive a selection of the toggle element 2302, and a second interface 2320 may be provided with the annotation user interface elements. Additionally and / or alternatively, the toggle element 2302 may be used to alternate between the first interface 2310 and the second interface 2320.

[0136] 24 shows an illustration of an example rating filtering element for rating-based filtering, according to an example embodiment of the present disclosure. At 2410, a number of annotation user interface elements 2412 can be provided in response to an object in a scene being recognized. The system and method can then receive a selection of a filter tag 2414 (e.g., a top-rated tag only) and transition to 2420. At 2420, annotation user interface elements 2422 can include only user interface elements associated with objects that meet the filtering criteria. The filter tag 2414 can be provided in different colors and / or with different icons based on whether the filter tag 2414 is selected, unselected, or deselected.

[0137] 25 shows an illustration of an example rating filtering slider element for filtering based on ratings, according to an example embodiment of the present disclosure. In some implementations, the filtering may be based on an interaction with a filtering slider 2522. For example, at 2510, a number of annotation user interface elements may be provided for display along with a filter tag 2512. The filter tag 2512 may be selected to open a filtering slider 2522. At 2520, the filtering slider 2522 is interacted with to filter the annotation user interface elements to display only final user interface elements 2524 associated with objects having a rating of about 90%.

[0138] 26 shows an illustration of an exemplary search interface according to an exemplary embodiment of the present disclosure. In particular, the systems and methods disclosed herein can switch between a first interface 2610 and a second interface 2620 based on a search element selection. The first interface 2610 can include one or more out-of-reticle user interface elements 2612 provided for display as semi-transparent and one or more focus user interface elements 2614 provided as entirely opaque to indicate that the associated object is within the reticle. A search element can then be selected to transition to a second interface 2620 providing a details panel associated with the object focus (e.g., the object associated with the focus user interface element 2614).

[0139] 27 illustrates a block diagram of an exemplary tag generation model 2700 according to an exemplary embodiment of the present disclosure. The exemplary tag generation model 2700 may include multiple machine learning based models and may include one or more deterministic functions. The tag generation model 2700 may be trained to receive image data 2702 (e.g., multiple image frames associated with a scene) and output one or more tags 2724 (e.g., filter tags and / or candidate query tags).

[0140] The image data 2702 may be processed by a stitching model 2704 to determine whether two or more image frames describe the same scene. If the image frames are determined to be associated with the same scene, the stitching model may generate scene data 2706 that describes the image frames stitched together. The scene data 2706 and / or the image data 2702 may be processed by a discriminative model to recognize and / or classify objects in the scene and / or images. The discriminative model may include a detection model 2708, a segmentation model 2710, and a recognition model 2712. The image data 2702 and / or the scene data 2706 may be processed by the detection model 2708 to generate bounding boxes around one or more objects detected in the scene. The bounding boxes and the image data 2702 (and / or the scene data 2706) may be processed by the segmentation model 2710 to segment portions of the image associated with the bounding boxes. The segmented portions of the image may be processed by a recognition model 2712 to identify each of the detected objects to generate object data 2714. The object data 2714 may then be used to search 2716 one or more databases for object specific information 2718 for each identified object.

[0141] The object specific information 2718 and / or the contextual data 2720 may then be processed by a tag determination model 2722 to generate one or more tags. The one or more tags may then be used to receive input from a user to provide more tailored data to the user.

[0142] Exemplary Methods 6 shows a flow chart diagram of an exemplary method performed in accordance with an exemplary embodiment of the present disclosure. Although FIG. 6 shows steps performed in a specific order for purposes of explanation and discussion, the method of the present disclosure is not limited to the specifically shown order or sequence. Various steps of method 600 may be variously omitted, rearranged, combined, and / or adapted without departing from the scope of the present disclosure.

[0143] At 602, a computing system may obtain image data generated by a mobile image capture device. The image data may represent a scene.

[0144] At 604, the computing system can process the image data to determine a plurality of objects in the scene. The plurality of objects can include one or more consumer products.

[0145] At 606, the computing system may obtain object specific information for one or more objects of the plurality of objects. The object specific information may include one or more details associated with each of the one or more objects.

[0146] In some implementations, the computing system can obtain context data associated with the user and determine a query based on the image data and the context data. The object specific information can be obtained based at least in part on the query. The context data can describe a user location, a user preference, past user queries, and / or a user shopping history. For example, the context data can describe a user location. The computing system can obtain one or more popular queries associated with the user location. The query can be determined based at least in part on the one or more popular queries.

[0147] Alternatively and / or additionally, the computing system may determine an object class associated with the plurality of objects, and the object specific information may be obtained based at least in part on the object class.

[0148] At 608, the computing system can provide one or more user interface elements overlaid on the image data. The one or more user interface elements can describe object specific information. In some implementations, the one or more user interface elements can include multiple product attributes associated with a particular object in the scene.

[0149] In some implementations, the computing system can obtain input data associated with a selection of a particular user interface element associated with a particular product attribute and provide one or more indicators overlaid on the image data. The one or more indicators can describe one or more particular objects associated with the one or more particular product attributes. In some implementations, the particular product attribute can include a threshold product rating and the particular user interface element can include a slider associated with a range of consumer product ratings. Additionally and / or alternatively, the multiple product attributes can include multiple different product types.

[0150] Alternatively and / or additionally, the computing system may determine a plurality of filters associated with the plurality of objects. Each filter may include criteria associated with a subset of the plurality of objects. The computing system may provide the plurality of filters for display in a user interface. In some implementations, the computing system may obtain a filter selection associated with a particular filter of the plurality of filters and provide an augmented reality overlay over one or more image frames. The augmented reality overlay may include one or more user interface elements being provided over each object that meets the respective criteria of the particular filter.

[0151] In some implementations, a computing system may receive audio data, the audio data may describe a voice command, the computing system may determine a particular object associated with the voice command, and provide an augmented image frame showing the particular object associated with the voice command.

[0152] 7 shows a flow chart diagram of an exemplary method performed in accordance with an exemplary embodiment of the present disclosure. Although FIG. 7 shows steps performed in a specific order for purposes of explanation and discussion, the method of the present disclosure is not limited to the specifically shown order or sequence. Various steps of method 700 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0153] At 702, a computing system can obtain video stream data generated by a mobile image capture device. The video stream data can include a plurality of image frames.

[0154] At 704, the computing system may determine that the first image frame and the second image frame are associated with a scene. The first image frame may include a first set of objects, and the second image frame may include a second set of objects. In some implementations, determining that the first image frame and the second image frame are associated with a scene may include determining that the first set of objects and the second set of objects are associated with a particular object class. Alternatively and / or additionally, determining that the first image frame and the second image frame are associated with a scene may include determining that the first image frame and the second image frame were captured at a particular location.

[0155] At 706, the computing system may generate scene data including a first image frame and a second image frame of the plurality of image frames.

[0156] At 708, the computing system can process the scene data to determine a plurality of objects in the scene. The plurality of objects can include one or more consumer products. In some implementations, the plurality of objects can include a plurality of consumer products.

[0157] At 710, the computing system can obtain object specific information for one or more objects of the plurality of objects. The object specific information can include one or more details associated with each of the one or more objects. In some implementations, the object specific information can include one or more consumer product details associated with each of the plurality of objects.

[0158] At 712, the computing system may provide one or more user interface elements overlaid on the one or more objects. The one or more user interface elements may describe object-specific information. In some implementations, multiple user interface elements overlaid on multiple objects may be provided. The multiple user interface elements may describe object-specific information. The multiple user interface elements may be associated with multiple consumer products. In some implementations, providing the one or more user interface elements overlaid on the one or more objects may include adjusting a number of pixels associated with an exterior region surrounding the one or more objects.

[0159] 8 shows a flow chart diagram of an exemplary method performed in accordance with an exemplary embodiment of the present disclosure. Although FIG. 8 shows steps performed in a specific order for purposes of explanation and discussion, the method of the present disclosure is not limited to the specifically shown order or sequence. Various steps of method 800 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0160] At 802, a computing system may obtain image data. The image data may represent a scene.

[0161] At 804, the computing system may process the image data to determine a plurality of filters. The plurality of filters may be associated with a plurality of objects in the scene. In some implementations, processing the image data to determine a plurality of filters may include processing the image data to recognize a plurality of objects in the scene, determining a plurality of differentiating attributes associated with differentiators between the plurality of objects, and determining a plurality of filters based at least in part on the plurality of differentiating attributes. The image data may be processed with one or more machine learning models (e.g., a detection model, a segmentation model, a classification model, and / or a recognition model).

[0162] In some implementations, the computing system can determine a number of filters based at least in part on the obtained context data. The context data can include the user's current location, a particular user profile, global trends, time of day, season, and / or recent interactions by the user with one or more applications (e.g., recent searches in a search application). For example, the user's recent queries can be used as filters if the queries apply to at least one object in the scene. Additionally and / or alternatively, other users who have previously been at a location may have used a particular tag at a higher rate than another tag. The particular tag may be provided to the user based on previous interactions of other users at a given location.

[0163] At 806, the computing system can provide one or more particular filters of the plurality of filters for display in the user interface. The one or more particular filters can be provided via a user interface tip provided as a selectable user interface element.

[0164] At 808, the computing system may obtain input data. The input data may be associated with a selection of a particular filter among a plurality of filters.

[0165] At 810, the computing system can provide one or more indicators overlaid on the image data. The one or more indicators can describe one or more particular objects associated with a particular filter. In some implementations, the one or more indicators can include object specific information associated with the one or more particular objects. Providing the one or more indicators overlaid on the image data can include an augmented reality experience.

[0166] In some implementations, the computing system can obtain second input data. The second input data can be associated with a zoom input. The zoom input can be associated with one or more particular objects. The computing system can obtain second information associated with the one or more particular objects. An augmented image can be generated based at least in part on the image data and the second information. The augmented image can include a zoomed-in portion of the scene associated with an area that includes the one or more particular objects. In some implementations, the one or more indicators and the second information can be overlaid on the one or more particular objects.

[0167] In some implementations, determining that an object meets certain criteria may involve obtaining object-specific information about a particular object, parsing the information into one or more segments, and processing the segments to determine a particular segment classification (e.g., that the segment is related to a flavor, an ingredient, a source, a location, etc.). The computing system may then process the segments and the given criteria to determine whether there is an association. The processing may involve natural language processing and may involve determining whether one or more segments are associated with the given criteria (e.g., whether the segment includes a word match) based on one or more knowledge graphs, or describing the given criteria (e.g., the segment states "citrus" or a synonym of citrus, and the criteria is an item that has a citrus flavor).

[0168] Alternatively and / or additionally, the object-specific information may include pre-built indexed data into one or more information categories (e.g., ratings, calories, flavors, usage, ingredients, emissions, etc.) The object-specific information may then be crawled when checking for keywords or information associated with the selected tag.

[0169] In some implementations, objects may be associated with a particular tag before the tag is provided for display. For example, multiple objects may be identified and multiple respective object-specific information sets may be obtained. The object-specific information sets may be analyzed and processed to generate a profile set for each object. The profile sets may be compared to each other to determine differentiating attributes between the objects. The differentiating attributes may be used to generate tags that narrow the list of objects. Objects with particular differentiating attributes may be pre-associated with tags such that when a tag is presented and selected, the computing system may automatically highlight or show the particular object associated with that particular tag.

[0170] Additionally and / or alternatively, the object-specific information may include one or more predefined tags indexed in a database and / or knowledge graph. In response to obtaining the object-specific information, the computing system may determine which tags are universal to all objects in the scene and prune those tags. The remaining predefined tags may be offered for display and selection. Once a tag is selected, the computing system may then present each of the objects that include an indexed reference to the particular predefined tag.

[0171] In some implementations, the tag or tags can be selected so as not to disrupt the user experience. The tag or tags can be based on search queries by other users when searching for a given object class or a particular object. In some implementations, the computing system can store and retrieve data related to the first and last searches associated with a particular object and a particular object class. Additionally and / or alternatively, the search query data of a particular user or users can be indexed with the user's location at the time of a given query or filter. The data can then be used to determine tags for the particular user or other users. The tag or tags can be generated to predict what a user may want to know about a scene, environment, and / or object. The computing system can generate tags based on what the user should search to reach a final action (e.g., a purchase selection, a DIY step, etc.).

[0172] Additional Disclosures The techniques discussed herein refer to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functionality among components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications may be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0173] Although the present subject matter has been described in detail with respect to various specific exemplary embodiments thereof, each example is provided as an explanation, not a limitation of the present disclosure. Upon understanding the above, those skilled in the art can easily create modifications, variations, and equivalents of such embodiments. Thus, the present disclosure does not exclude the inclusion of such modifications, variations, and / or additions to the present subject matter as would be readily apparent to those skilled in the art. For example, features shown or described as part of one embodiment may be used with another embodiment to yield yet a further embodiment. Thus, it is intended that the present disclosure cover such modifications, variations, and equivalents. [Explanation of symbols]

[0174] 10. Computing Devices 50 Computing Devices 100 Computing system, system 102 User Computing Devices 112 processors 114 Memory, user computing device memory 130 Server Computing System 132 processors 134 Memory 150 Training Computing System 152 processors 154 Memory 160 Model Trainer 180 Network 200 Object Filtering and Information Display System 212 Mobile Computing Devices 300 Object Filtering and Information Display System 400 Object Filtering and Information Display System

Claims

1. one or more processors; and one or more computer-readable storage media collectively storing instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations including: acquiring image data generated by a mobile image capture device, the image data being indicative of a scene; processing the image data to determine a plurality of objects in the scene; obtaining object specific information for one or more objects of the plurality of objects, the object specific information including one or more details associated with each of the one or more objects; processing the image data to determine a plurality of filters, the plurality of filters being associated with the one or more details; and providing one or more user interface elements overlaid on the image data, the one or more user interface elements describing one or more particular filters of the plurality of filters; obtaining input data, the input data being associated with a selection of a particular filter of the plurality of filters; providing one or more indicators overlaid on the image data, the one or more indicators describing one or more particular objects associated with the particular filter; Computing system.

2. the one or more user interface elements include a plurality of product attributes associated with a particular object in the scene; The operation includes: Obtaining input data associated with a selection of a particular user interface element associated with a particular product attribute; 13. The system of claim 1, further comprising: providing one or more indicators overlaid on the image data, the one or more indicators describing one or more particular objects associated with the one or more particular product attributes.

3. the particular product attribute includes a threshold product rating; The system of claim 2 , wherein the particular user interface element comprises a slider associated with a range of consumer product ratings.

4. The system of claim 2 , wherein the plurality of product attributes comprises a plurality of different product types.

5. The operation includes: Obtaining contextual data associated with a user; determining a query based on the image data and the contextual data; The system of claim 1 , wherein the object specific information is obtained based at least in part on the query.

6. The context data describes a user location; The operation includes: obtaining one or more popular queries associated with the user location; The system of claim 5 , wherein the query is determined based at least in part on the one or more popular queries.

7. The system of claim 5 , wherein the context data describes at least one of a user location, a user preference, a past user query, or a user shopping history.

8. The operation includes: determining a plurality of filters associated with the plurality of objects, each filter including criteria associated with a subset of the plurality of objects; providing the plurality of filters for display in a user interface; obtaining a filter selection associated with a particular filter of the plurality of filters; 10. The system of claim 1, further comprising: providing an augmented reality overlay over one or more image frames, the augmented reality overlay including the one or more user interface elements being provided over each object that meets the criteria of each of the particular filters.

9. The operation includes: determining an object class associated with the plurality of objects; The system of claim 1 , wherein the object specific information is obtained based at least in part on the object class.

10. The operation includes: receiving audio data, the audio data describing a voice command; determining a particular object associated with the voice command; and providing an augmented image frame showing the particular object associated with the voice command.

11. 1. A computer-implemented method comprising: acquiring, by a computing system having one or more processors, video stream data generated by a mobile image capture device, the video stream data including a plurality of image frames; determining, by the computing system, that a first image frame and a second image frame are associated with a scene; generating, by the computing system, scene data including the first image frame and the second image frame of the plurality of image frames; processing, by the computing system, the scene data to determine a plurality of objects in the scene; obtaining, by the computing system, object specific information for one or more objects of the plurality of objects, the object specific information including one or more details associated with each of the one or more objects; processing the image data to determine a plurality of filters, the plurality of filters being associated with the one or more details; providing, by the computing system, one or more user interface elements overlaid on the one or more objects, the one or more user interface elements describing one or more particular filters of the plurality of filters; obtaining input data, the input data being associated with a selection of a particular filter from the plurality of filters; providing one or more indicators overlaid on the image data, the one or more indicators describing one or more particular objects associated with the particular filter. method.

12. the plurality of objects includes a plurality of consumer products; the object specific information includes one or more consumer product details associated with each of the plurality of objects; Providing, by the computing system, the one or more user interface elements overlaid on the one or more objects, comprises:

12. The method of claim 11, comprising providing, by the computing system, a plurality of user interface elements overlaid on the plurality of objects, the plurality of user interface elements describing the object specific information, the plurality of user interface elements being associated with the plurality of consumer products.

13. the first image frame includes a first set of objects and the second image frame includes a second set of objects; The step of determining, by the computing system, that the first image frame and the second image frame are associated with the scene comprises: The method of claim 11 , comprising determining, by the computing system, that the first set of objects and the second set of objects are associated with a particular object class.

14. The step of determining, by the computing system, that the first image frame and the second image frame are associated with the scene comprises: The method of claim 11 , comprising determining, by the computing system, that the first image frame and the second image frame were captured at a particular location.

15. 12. The method of claim 11, wherein providing, by the computing system, the one or more user interface elements to be overlaid on the one or more objects comprises adjusting a number of pixels associated with an exterior region surrounding the one or more objects.

16. One or more computer-readable storage media collectively storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, including: acquiring image data, the image data being indicative of a scene; processing the image data to determine a plurality of objects in the scene; obtaining object specific information for one or more objects of the plurality of objects, the object specific information including one or more details associated with each of the one or more objects; processing the image data to determine a plurality of filters, the plurality of filters being associated with a plurality of objects in the scene; providing one or more particular filters of the plurality of filters for display in a user interface; obtaining input data, the input data being associated with a selection of a particular filter of the plurality of filters; providing one or more indicators overlaid on the image data, the one or more indicators describing one or more particular objects associated with the particular filter; Including, One or more computer-readable storage media.

17. Processing the image data to determine the plurality of filters includes: processing the image data to recognize a plurality of objects in the scene; determining a plurality of differentiating attributes associated with differentiators among the plurality of objects; and determining the plurality of filters based at least in part on the plurality of differentiating attributes.

18. 20. The one or more computer-readable storage media of claim 17, wherein processing the image data to recognize the plurality of objects in the scene comprises processing the image data with a machine learning based model.

19. The operation includes: obtaining second input data, the second input data being associated with a zoom input, the zoom input being associated with the one or more particular objects; obtaining second information associated with the one or more particular objects; and 17. The one or more computer-readable storage media of claim 16, further comprising: generating an augmented image, the augmented image including a zoomed-in portion of the scene associated with an area including the one or more particular objects, and the one or more indicators and the second information are overlaid on the one or more particular objects.

20. 17. The one or more computer-readable storage media of claim 16, wherein the one or more indicators include object specific information associated with the one or more particular objects, and providing the one or more indicators overlaid on the image data includes an augmented reality experience.

Citation Information

Patent Citations

  • Terminal device, distribution device, control method of terminal device, control method of distribution device, control program, and recording medium

    JP2010118019A

  • Location-based searching

    JP2016184446A

  • Commodity search program and method in commodity shipment management system

    JP2018036851A

  • Real-time object detection and tracking

    JP2021523463A

  • Augmented Reality

    US20180341811A1