Rendering augmented reality content based on post-processing of application content

Through the display interface of computerized glasses, the automation assistant generates augmented reality content based on the application content, solving the problem that notifications in the prior art cannot distinguish user situations, and achieving the effect of improving user work efficiency.

CN119948386APending Publication Date: 2025-05-06GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070922.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-10
Filing Date
2023-10-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When existing automation assistants provide notifications to users, they cannot effectively distinguish the user's situation, resulting in notifications that may waste users' valuable bandwidth and affect users' work efficiency.

Method used

Through the computerized display interface of glasses, the automation assistant can generate augmented reality content based on post-processing of the application content, determine that a specific object or classified object is located in the field of view, and render helpful content at corresponding locations to advance the ongoing tasks of the user.

Benefits of technology

It provides targeted and real-time information and suggestions without interfering with the user's work, improving the user's work efficiency and convenience of task completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948386A_ABST
    Figure CN119948386A_ABST
Patent Text Reader

Abstract

Implementations relate to automated assistants that provide augmented reality content resulting from post-processing of application content via a display interface of computerized glasses. The application content may be identified based on prior interactions between a user and one or more applications, and the application content may be processed to determine objects and / or object classifications that may be associated with the application content. When the user is wearing the computerized glasses and the object is detected within the field of view of the computerized glasses, the automated assistant may cause particular content to be rendered at the display interface of the computerized glasses. In some implementations, the content may be generated to supplement and / or differ from existing content that the user may have accessed, to facilitate preventing reuse of applications and / or preserving computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] A human may engage in a human-computer conversation with an interactive software application, referred to herein as an "automated assistant" (also referred to as a "digital agent," "chatbot," "interactive personal assistant," "intelligent personal assistant," "assistant application," "conversational agent," etc.). For example, a human (who may be referred to as a "user" when they interact with an automated assistant) may provide commands and / or requests to the automated assistant using spoken natural language input (i.e., utterances) and / or by providing textual (e.g., typed) natural language input, which in some cases may be converted to text and then processed.

[0002] Some automated assistants may provide notifications to users when certain features of the user's context become apparent to the automated assistant. However, these notifications may only convey information that has been created by the user and / or otherwise readily available to the user when initializing a particular application. For example, a user who relies on navigation to reach a destination may receive a notification about the operating hours of the destination (e.g., "The library closes at 5PM."). In addition, such information conveyed by the automated assistant may be globally communicated to other users who themselves have similar behaviors (e.g., the operating hours may also be presented to the user when other users navigate to the destination). Although such information may still be helpful, providing these notifications indiscriminately may waste the user's precious bandwidth (i.e., limited attention). For example, a user who is actively working to complete their work tasks may no longer operate efficiently when receiving a notification about an upcoming event. However, notifications that reflect a wider range of information that the user has granted access to the automated assistant can equip the user with appropriate information to make their efforts more efficient. Summary of the invention

[0003] The implementations described herein relate to an automated assistant that can provide augmented reality content generated by post-processing of application content via a display interface of computerized glasses. The application content can be accessed by the user during an interaction between the user and the application, and the application content can be processed to generate content that can be helpful in certain contexts, such as when a particular object is located within the field of view of the computerized glasses. In other words, the automated assistant can determine that a particular object and / or a particular classification of objects is located within the field of view of the computerized glasses. Based on this determination, the automated assistant can identify and / or generate content to be rendered within the field of view of the computerized glasses (e.g., depending on the location of the object) to facilitate processing of tasks that the user may have been working to complete using the application.

[0004] As an example, the user may interact with a word processing application to complete a report by 5pm the next day. The report may be associated with a calendar entry named "Report Presentation" accessible via a calendar application. The automated assistant may access application content generated by the word processing application and / or the calendar application with prior permission from the user to generate helpful augmented reality content. For example, helpful content may include, but is not limited to, an estimated amount of time remaining to complete the report, an amount of time to make certain progress on the report, suggested edits to the report, scheduling conflicts that may occur if the user continues to work on the report, and / or any other content that may be helpful to the user. In some implementations, one or more entries may be generated by the automated assistant to associate a corresponding instance of the helpful content (e.g., an estimated amount of time remaining to complete the report) with a classification of objects (e.g., a clock) that the user may be viewing within the field of view of the computerized glasses. Thereafter, when the automated assistant determines that the object or an object of the classification of shared objects is located within the field of view of the computerized glasses being worn by the user, the automation may render the helpful content at the interface of the computerized glasses.

[0005] For example, the user may be wearing their computerized glasses while interacting with a word processing application during the morning of the day a report is due (e.g., the report may be due at 5 p.m. that day). The automated assistant may determine that the user has directed their attention away from the word processing application to view a clock hanging on a nearby wall. In other words, the user may be viewing the word processing application via a computing device (e.g., a laptop), and the automated assistant may determine based on one or more sensors of the computerized glasses that the user has directed their sight away from the computing device and toward the clock. In some implementations, the automated assistant may employ one or more heuristic processes and / or one or more trained machine learning models to classify objects that may be within the field of view of the computerized glasses. The automated assistant may then determine whether the object is associated with an entry associated with the user who is wearing the computerized glasses. For example, the automated assistant may determine that the clock being viewed by the user is associated with an entry that relates a report being drafted by the user to a time measurement device (e.g., a clock).

[0006] In some implementations, when the automated assistant determines that the user is viewing an object corresponding to a previously generated entry or other relevant data, the automated assistant can cause certain content to be rendered at the display interface of the computerized glasses. For example, image data can be generated to convey information to the user so that the user is informed of the estimated amount of time to complete a report before a deadline (e.g., 5 p.m.). The automated assistant can then cause the computerized glasses to render an image that conveys information generated based on prior and / or current interactions between the user and the word processing application. The image can include natural language content and / or other visual content that can, for example, illustrate a stopwatch or status bar that actively conveys the amount of time left to complete the report before the deadline. In some implementations, the image can convey information based on multiple data sources, such as a specific calendar event that may be between the current time and the time of the deadline and therefore affect the user's ability to work uninterruptedly before the deadline.

[0007] In some implementations, the content may be rendered at a position of the display interface that can be dynamically adjusted in real time according to how the user manipulates their head and / or otherwise offsets the field of view of the computerized glasses. Alternatively or additionally, the rendered content may be modified according to any feature of the scene in which the content is being rendered. For example, content rendered above a clock and / or rendered adjacent to a clock may be dynamically rendered so that the content appears to adapt to the passage of time indicated by the clock. In some cases, the content may be a pie-shaped image with a label covering a portion of the clock, the label having a text indicating "Estimated time left to finish the report (estimated time left to finish the report)". In this way, the assistant can assist the user in visualizing the amount of free time that the user may have before the deadline if they continue to work at their current pace. Alternatively or additionally, the content may be a pie-shaped image with a label covering a portion of the clock, the label having a text indicating "Estimated amount of free time if you continue working for 45 minutes (estimated amount of free time if you continue working for 45 minutes)". In some implementations, the content may be selectable and / or modifiable to allow a user to visualize other scenarios in which they may work for different amounts of time.

[0008] In some implementations, content can be rendered according to settings that can control any content being rendered by the automated assistant at the computerized glasses. For example, a user can specify via settings that the automated assistant must limit rendering certain types of augmented reality content (e.g., work-related content) to certain times of the day and / or week. In some implementations, the automated assistant can render content based on certain features of the object such as the distance between the user and the object, the size of the object, whether the object includes clear text, the decorative features of the object, the functional aspects of the object, and / or any other features that can be used as the basis for rendering content in a specific manner. For example, in some cases, the automated assistant can render the content as adjacent to the corresponding object, rather than as being located on top of the corresponding object. In some implementations, the automated assistant can provide settings for rendering content only after certain conditions are met (e.g., the user has stared at a particular classified object for a threshold duration). Alternatively or additionally, when the user interacts with the automated assistant and / or the computerized glasses, one or more heuristic processes and / or one or more trained machine learning models can be used to automatically determine settings for features of the augmented reality content.

[0009] The above description is provided as an overview of some implementations of the present disclosure. Further descriptions of those implementations and other implementations are described in more detail below.

[0010] Other implementations may include a non-transitory computer-readable storage medium storing instructions that are executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform methods such as one or more of the methods described above and / or elsewhere herein. Still other implementations may include a system of one or more computers including one or more processors operable to execute stored instructions to perform methods such as one or more of the methods described above and / or elsewhere herein.

[0011] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are considered part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1A , Figure 1B and Figure 1C An automated assistant is illustrated that renders a view of content for a user based on whether certain objects in the field of view of the computerized glasses are related to application content previously accessed by the user.

[0013] Figure 2 A system facilitating an automated assistant that can provide augmented reality content resulting from post-processing of application content via a display interface of computerized eyewear is illustrated.

[0014] Figure 3 Illustrated is a method for causing certain content to be rendered at an interface of computerized glasses when an object having a particular object classification is determined to be within the field of view of the computerized glasses and the object classification is determined to be associated with prior interaction between a user and an application.

[0015] Figure 4 is a block diagram of an example computer system. DETAILED DESCRIPTION

[0016] Figure 1A , Figure 1B and Figure 1C 1 illustrates views 100, 120, and 140 of content rendered by an automated assistant to a user 102 based on whether certain objects in the field of view of the computerized glasses are related to application content previously accessed by the user 102. Figure 1A As illustrated in view 100 of FIG. 1 , a user 102 may interact with an application, such as a calendar application 106, to access certain application content 108. For example, application content 108 may include reminders 110 and events 114 that the user 102 may wish to be reminded of throughout his or her day. Figure 1A As provided in , user 102 may access calendar application 106 in the morning (e.g., 6:30 a.m.) or at other times prior to their events for the day. In some implementations, application content may be processed using one or more heuristic processes and / or one or more trained machine learning models to generate entries that relate portions of the application content to certain objects in the physical world.

[0017] For example, application content rendered at display interface 112 of computing device 104 may be processed to identify objects that may be associated with the content. In some cases, because the application content includes the time of an event, a data entry may be created that associates a time-related object (e.g., a watch, a wall clock, a timer, etc.) with words and / or phrases provided in the application content. For example, an entry that associates the phrase "mowing the lawn" with a category of objects such as time-recording objects may be generated based at least on the time indicator (e.g., "9:15 a.m.", "11:00 a.m.", etc.) also being in the application content. Alternatively or additionally, another entry that associates the word "breakfast" with a category of objects having a European theme may be generated based at least on the event "conducting a conference call with European parties" also being in the application content.

[0018] In some implementations, other post-processing may be performed to generate assistant data from the entry. The assistant data may include, for example, natural language content and / or images that may assist the user in certain situations. For example, the amount of time used to complete the task of mowing the lawn may be estimated and compared to other information conveyed by the calendar application. Based on this comparison, assistant data may be generated to convey information to the user regarding when the user 102 may incorporate "mowing the lawn" into their schedule. Alternatively or additionally, information that may be relevant to the "conference call with Europe" event may be predicted, and the sources of that information may be identified. Based on these sources, assistant data may be generated to convey the sources of the information to the user 102 as a reminder prior to the "conference call" event.

[0019] For example, and if Figure 1B 120, the user 102 may be wearing computerized glasses 124, which may include one or more sensors and / or one or more interfaces for rendering content to the user 102. After viewing the calendar application 106, the user 102 may be located in their living room 126. While the user 102 is in their living room 126, the user 102 may view various objects such as another computing device 132 and / or a wall clock 130, which may provide an analog clock face 128. When the user 102 is viewing the clock face 128, the clock face 128 may be located in the field of view 122 of the computerized glasses 124. Sensor data from one or more sensors of the computerized glasses 124 may be processed to determine objects and / or classifications of objects within the field of view 122 of the computerized glasses. For example, the sensor data may be processed using one or more object recognition models and / or other trained machine learning models to facilitate determining identifiers of objects within the field of view 122. The object's identifier may then be compared to any entries relating the object identifier to the application content and / or other data that has been generated based on the application content.

[0020] For example, and if Figure 1B, the user 102 may be viewing an analog clock face 128, which may be characterized by the automated assistant as having a classification of "clock." The automated assistant may then determine whether there are any entries that associate an object having the classification of clock with any content. For example, the automated assistant may determine that the clock is associated with an entry that relates a time tracking device to the amount of time that the user 102 may have to mow the lawn. Based on this determination, the automated assistant may use the assistant data associated with the entry to render helpful information at the display interface of the computerized glasses 124. The rendered information may be augmented reality content that may be rendered at a location within the field of view 122 of the computerized glasses 124 based on where a particular object is located within the field of view 122 of the computerized glasses 124.

[0021] For example, and if Figure 1C , the automated assistant may cause the display interface 152 of the computerized glasses to render content at or near an object determined to be associated with the entry and / or application content. For example, the content may include one or more images and / or one or more portions of natural language content. As illustrated in view 140, the one or more images may include a shape 148 that indicates the amount of time that the user 102 may have to complete the task of mowing the lawn before an upcoming event indicated in the application content 108. Alternatively or additionally, one or more portions of the natural language content rendered at the display interface 152 may include phrases 150, such as "Amount of time left to mow thegrass". The content rendered at the display interface 152 may be rendered based on sensor data generated by one or more sensors of the computerized glasses 124, such as a camera 144, a microphone 146, and / or any other sensor that can communicate with a computing device.

[0022] For example, the location of the analog clock face 128 may be determined based on the sensor data, and the shape 148 may be rendered at that location such that the shape 148 at least partially overlaps the analog clock face 128. Alternatively or additionally, the phrase 150 may be rendered adjacent to the analog clock face 128 with an optional "callout" line that may indicate that the phrase 150 is associated with the analog clock face 128. In some implementations, the characteristics of the shape 148 and / or the phrase 150 may be based on characteristics of the context of the user 102. For example, the opacity, angle, font, size, and / or any other characteristics of the content rendered at the display interface 152 may be selected based on the number of other objects that may be present in the field of view 122 of the computerized glasses 124 and / or the physical properties of the other objects.

[0023] In some implementations, the rendered content may be selectable via one or more inputs and / or input gestures. Selection of the rendered content via user input may cause the automated assistant to interact with a separate application that may be the basis for the rendered content. For example, the user 102 may provide a gesture (e.g., a slide away of the phrase 150) to indicate that the user 102 is not interested in the information conveyed by the phrase 150. In response to this gesture, the automated assistant may cause the "mow the lawn" reminder in the reminder 110 to be removed from the reminder 110 of the calendar application 108. Alternatively or additionally, the user 102 may provide a gesture (e.g., a finger press and drag) to cause the shape 148 to move to a different portion of the analog clock face 128, and cause the shape 148 to move from an original position (e.g., covering 8:50 a.m. to 9:15 a.m.) to a different position (e.g., covering 12:00 p.m. to 12:45 p.m.). In response to this gesture, the automated assistant may interact with calendar application 106 to cause a “mow the lawn” reminder to render at an interface of computing device 104 and / or another computing device at 12:00 PM, or any other time associated with a different location.

[0024] Figure 2A system 200 is illustrated that facilitates an automated assistant 204 for providing augmented reality content generated by post-processing of application content via a display interface of computerized glasses. The system 200 may include a computing device 202 for providing access to the automated assistant 204, which may be computerized glasses and / or docked with computerized glasses. The automated assistant 204 may operate as part of an assistant application provided at one or more computing devices (such as the computing device 202 and / or a server device). The user may interact with the automated assistant 204 via an assistant interface 220, which may be a microphone, a camera, a touch screen display, a user interface, and / or any other device capable of providing an interface between the user and the application. For example, the user may initialize the automated assistant 204 by providing speech, text, and / or graphical input to the assistant interface 220 to cause the automated assistant 204 to initiate one or more actions (e.g., providing data, controlling peripheral devices, accessing an agent, generating input and / or output, etc.). Alternatively, the automated assistant 204 may be initialized based on processing context data 236 using one or more trained machine learning models. Context data 236 may characterize one or more characteristics of an environment in which automated assistant 204 is accessible, and / or one or more characteristics of a user predicted to be intent on interacting with automated assistant 204. Computing device 202 may include a display device, which may be a display panel, including a touch interface for receiving touch input and / or gestures to allow a user to control applications 234 of computing device 202 via the touch interface. In some implementations, computing device 202 may lack a display device, thereby providing audible user interface outputs rather than graphical user interface outputs. In addition, computing device 202 may provide a user interface for receiving oral natural language input from a user, such as a microphone. In some implementations, computing device 202 may include a touch interface and may not contain a camera, but may optionally include one or more other sensors.

[0025] Computing device 202 and / or other third-party client devices may communicate with the server device via a network such as the Internet. Additionally, computing device 202 and any other computing devices may communicate with each other via a local area network (LAN) such as a Wi-Fi network. Computing device 202 may offload computing tasks to a server device in order to conserve computing resources at computing device 202. For example, a server device may host automated assistant 204, and / or computing device 202 may send input received at one or more assistant interfaces 220 to a server device. However, in some implementations, automated assistant 204 may be hosted at computing device 202, and various processes that may be associated with automated assistant operations may be performed at computing device 202.

[0026] In various implementations, all or less than all aspects of automated assistant 204 may be implemented on computing device 202. In some of those implementations, aspects of automated assistant 204 are implemented via computing device 202 and may interface with a server device that may implement other aspects of automated assistant 204. The server device may optionally serve multiple users and their associated assistant applications via multiple threads. In implementations in which all or less than all aspects of automated assistant 204 are implemented via computing device 202, automated assistant 204 may be an application separate from (e.g., installed “on top of”) an operating system of computing device 202, or may alternatively be implemented directly by (e.g., treated as an application of, but integrated with) the operating system of computing device 202.

[0027] In some implementations, automated assistant 204 may include an input processing engine 206 that may employ a plurality of different modules to process input and / or output of computing device 202 and / or server devices. For example, input processing engine 206 may include speech processing engine 208 that may process audio data received at assistant interface 220 to identify text embodied in the audio data. The audio data may be sent from, for example, computing device 202 to a server device in order to conserve computing resources at computing device 202. Additionally or alternatively, the audio data may be processed exclusively at computing device 202.

[0028] The process for converting audio data to text may include a speech recognition algorithm that can use a neural network and / or a statistical model to identify grouped audio data corresponding to a word or phrase. The text converted from the audio data can be parsed by the data parsing engine 210 and used as text data for the automated assistant 204, which can be used to generate and / or identify command phrases, intentions, actions, slot values, and / or any other content specified by the user. In some implementations, the output data provided by the data parsing engine 210 can be provided to the parameter engine 212 to determine whether the user has provided input corresponding to a specific intention, action, and / or routine that can be performed by the automated assistant 204 and / or an application or agent that can be accessed via the automated assistant 204. For example, the assistant data 238 can be stored at a server device and / or a computing device 202, and can include data defining one or more actions that can be performed by the automated assistant 204, and parameters necessary for performing the action. The parameter engine 212 can generate one or more parameters for the intention, action, and / or slot value, and provide one or more parameters to the output generation engine 214. Output generation engine 214 may use the one or more parameters to communicate with assistant interface 220 for providing output to a user, and / or to communicate with one or more applications 234 for providing output to one or more applications 234 .

[0029] In some implementations, automated assistant 204 may be an application that may be installed “on top of” an operating system of computing device 202, and / or may itself form part (or the entirety) of an operating system of computing device 202. The automated assistant application includes and / or has access to on-device speech recognition, on-device natural language understanding, and on-device fulfillment. For example, on-device speech recognition may be performed using an on-device speech recognition module that processes audio data (detected by a microphone) using an end-to-end speech recognition machine learning model stored locally at computing device 202. On-device speech recognition generates recognized text for spoken utterances (if any) present in the audio data. In addition, for example, on-device natural language understanding (NLU) may be performed using an on-device NLU module that processes the recognized text generated using on-device speech recognition and optionally context data to generate NLU data.

[0030] The NLU data may include an intent corresponding to the spoken utterance and optionally parameters for the intent (e.g., slot values). On-device fulfillment may be performed using an on-device fulfillment module that utilizes the NLU data (from the on-device NLU) and optionally other local data to determine actions to be taken to resolve the intent (and optionally parameters for the intent) of the spoken utterance. This may include determining local and / or remote responses (e.g., replies) to the spoken utterance, interactions with locally installed applications performed based on the spoken utterance, commands sent to an Internet of Things (IoT) device (directly or via a corresponding remote system) based on the spoken utterance, and / or other resolution actions performed based on the spoken utterance. The on-device fulfillment may then initiate local and / or remote execution / execution of the determined actions to resolve the spoken utterance.

[0031] In various implementations, remote speech processing, remote NLU, and / or remote fulfillment may be at least selectively utilized. For example, the identified text may be at least selectively sent to a remote automated assistant component for remote NLU and / or remote fulfillment. For example, the identified text may be optionally sent for remote execution in parallel with on-device execution or in response to failure of on-device NLU and / or on-device fulfillment. However, on-device speech processing, on-device NLU, on-device fulfillment, and / or on-device execution may be prioritized at least due to the reduced latency they provide when resolving spoken utterances (because resolving spoken utterances does not require client-server round trips). Further, on-device functionality may be the only functionality available when there is no network connection or when network connection is limited.

[0032] In some implementations, computing device 202 may include one or more applications 234, which may be provided by a third-party entity that is different from the entity that provided computing device 202 and / or automated assistant 204. Automated assistant 204 and / or an application state engine of computing device 202 may access application data 230 to determine one or more actions that can be performed by one or more applications 234, as well as a state of each of one or more applications 234 and / or a state of a corresponding device associated with computing device 202. Automated assistant 204 and / or a device state engine of computing device 202 may access device data 232 to determine one or more actions that can be performed by computing device 202 and / or one or more devices associated with computing device 202. Additionally, application data 230 and / or any other data (e.g., device data 232) can be accessed by automated assistant 204 to generate context data 236, which can characterize the context in which a particular application 234 and / or device is executing, and / or the context in which a particular user is accessing computing device 202, accessing application 234, and / or any other device or module.

[0033] While one or more applications 234 are executing at computing device 202, device data 232 may characterize the current operating state of each application 234 being executed at computing device 202. In addition, application data 230 may characterize one or more features of the executing applications 234, such as the content of one or more graphical user interfaces being rendered under the direction of one or more applications 234. Alternatively or additionally, application data 230 may characterize action paradigms based on the current operating state of the respective applications, which action paradigms may be updated by the respective applications and / or by automated assistant 204. Alternatively or additionally, one or more action paradigms of one or more applications 234 may remain static, but may be accessed by the application state engine in order to determine appropriate actions to be initiated via automated assistant 204.

[0034] The computing device 202 may further include an assistant invocation engine 222, which may use one or more trained machine learning models to process application data 230, device data 232, context data 236, and / or any other data accessible to the computing device 202. The assistant invocation engine 222 may process this data to determine whether to wait for the user to explicitly say an invocation phrase to invoke the automated assistant 204, or to treat the data as indicating the user's intention to invoke the automated assistant, rather than requiring the user to explicitly say an invocation phrase. For example, one or more trained machine learning models may be trained using instances of training data based on scenarios in which the user is in an environment where multiple devices and / or applications are exhibiting various operating states. Instances of training data may be generated to capture training data that characterizes contexts in which the user invokes the automated assistant and other contexts in which the user does not invoke the automated assistant. When one or more trained machine learning models are trained based on these instances of training data, the assistant invocation engine 222 may cause the automated assistant 204 to detect or limit detection of verbal invocation phrases from the user based on features of the context and / or environment. Additionally or alternatively, assistant invocation engine 222 can cause automated assistant 204 to detect or limit detection for one or more assistant commands from the user based on characteristics of the context and / or environment. In some implementations, assistant invocation engine 222 can be disabled or limited based on computing device 202 detecting an assistant suppression output from another computing device. In this manner, when computing device 202 is detecting an assistant suppression output, automated assistant 204 will not be invoked based on context data 236, which would otherwise cause automated assistant 204 to be invoked if no assistant suppression output was detected.

[0035] In some implementations, the system 200 may include an object detection engine 216 that can process sensor data from one or more sensors of the computerized glasses and / or any other device to determine whether a particular object and / or a particular classification of objects is within the field of view of the computerized glasses. The object detection engine 216 can use one or more heuristic processes and / or one or more trained machine learning models to determine whether certain objects are present in the field of view. For example, various models for detecting various types of objects can be employed to identify certain objects that may be within the field of view of the computerized glasses and / or the user. When an object is identified, an entry can be generated and / or identified for associating the object with any content that the user may have accessed that is associated with the object and / or the object classification.

[0036] For example, the system 200 may include an interactive correspondence engine 218 that can generate entries for associating certain instances of data such as application content with objects and / or object classifications. Thereafter, the entries can be used to determine whether to render assistant content at the display interface of the computerized glasses to promote assisting the user in doing certain tasks and / or other attempts. For example, the automated assistant can process content being accessed by the user (e.g., when the user is writing an email) with prior permission from the user to determine the classification of objects that may be related to the content. Alternatively or additionally, other sources of content can be identified to determine the assistant content that should be rendered for the user in response to the user viewing a specific object. For example, calendar data from a calendar application can be processed in conjunction with email data with prior permission from the user to promote the generation of content that can be rendered to the user to assist the user in preparing for events stored by the calendar in association with emails. Thereafter, when the user views an object associated with email data, the automated assistant 204 can cause the content to be rendered at the display interface of the computerized glasses.

[0037] In some implementations, the system 200 may include a glasses content engine 226 that can generate content for rendering at a display interface of a computerized glasses and / or any other interface of a computing device. For example, one or more trained machine learning models (e.g., deep learning neural network models) may be utilized to generate content that may be helpful in certain contexts from one or more instances of data associated with a user. In some cases, the generated content may include images, videos, audio, words, and / or phrases that may actively assist the user in achieving data that may be helpful in certain contexts. In some implementations, the historical interactions between the user and the automated assistant 204 may be processed by the glasses content engine 226 with prior permission from the user to determine the requests that the user may expect to make to the automated assistant 204 in certain contexts. Based on this processing, the glasses content engine 226 may generate assistant data representing responses to certain requests, which may be based on application content for use by the user and / or one or more objects that may be located in the field of view of the computerized glasses. In some implementations, the glasses content engine 226 may update the assistant data in real time based on information accessible to the automated assistant. For example, when content is being rendered at a display interface of a computerized pair of glasses (e.g., rendered as a virtual invitation overlaying a portion of a physical calendar) and an email associated with the content is received (e.g., an email delaying a virtual event), the content can be updated to reflect the changes indicated in the body of the email.

[0038] In some implementations, the system 200 may include a content rendering engine 224 that can process assistant data to generate content that can be rendered at one or more interfaces of computerized glasses. In some implementations, the content rendering engine 224 can generate content based on the physical properties of the object, which is at least part of the basis of the content being rendered. For example, the geographic location of the object can be used as the basis for rendering content. The geographic location can be identified in the application content previously accessed by the user and / or otherwise associated with the application content, and when the user is within a threshold distance of the geographic location, the content can be rendered at the computerized glasses. Alternatively or additionally, when the user interacts with the application content (e.g., writing a report on its word processing application) and the user views the object within a threshold duration between viewing the object, the content can be rendered.

[0039] In some implementations, the content rendering engine 224 can render content at the location of the display interface of the computerized glasses based on the location of the object within the field of view of the computerized glasses. Alternatively or additionally, the content rendering engine 224 can select the characteristics of the content based on the physical properties of the object. For example, the size of the content to be rendered can be selected based on the distance of the object from the user and / or the size of the object within the field of view of the computerized glasses. Alternatively or additionally, the content can be rendered to supplement and / or supplement the information that may have been conveyed by the object. For example, when the content to be rendered is related to a calendar event and the object is a calendar, the content rendering engine 224 can bypass the "date" part of the rendered content because the calendar may already include a date. Instead, the content can be rendered at a location of the physical calendar corresponding to the date of the calendar event, and the amount of content rendered can depend on the amount of white space or other areas with a threshold degree of color uniformity. In some implementations, when a particular object embodies a threshold degree of color uniformity (e.g., most areas of the object are a single color), the content rendering engine 224 can choose to render the content as overlapping with most of the object (e.g., at least half). However, if a particular object does not exhibit a threshold degree of color uniformity (e.g., less than a majority area of ​​the object is not a single color), the content rendering engine 224 may choose to render content at a location within the field of view that is adjacent to the location of the object and / or otherwise does not overlap with the object.

[0040] Figure 3 A method 300 is illustrated for causing certain content to be rendered at an interface of computerized glasses when an object with a particular object classification is determined to be located within the field of view of computerized glasses and the object classification is determined to be associated with a prior interaction between a user and an application. The method 300 may be performed by one or more computing devices, applications, and / or any other device or module that may be associated with an automated assistant. The method 300 may include an operation 302 of determining whether an object is detected within the field of view of computerized glasses being worn by a user. The computerized glasses may provide access to an automated assistant that may respond to user input and / or gestures and / or otherwise operate as an interface for the automated assistant. The computerized glasses may include one or more sensors for generating sensor data that may be processed to determine features of the user's context, such as whether the user is providing input, whether a particular object is located within the user's field of view, and / or the user's position relative to an object and / or the user's context. The computerized glasses may also include one or more interfaces for rendering content to the user, such as at a transparent display interface of the computerized glasses, thereby allowing the user to view their surroundings and any content of the display interface.

[0041] In some implementations, an object located in the field of view of the computerized glasses or outside the field of view of the computerized glasses can be detected based on processing sensor data from one or more sensors of the computerized glasses. The sensor data from the one or more sensors can be processed using one or more heuristic processes and / or one or more trained machine learning models to facilitate detection and classification of objects characterized by the sensor data. For example, a user can wear the computerized glasses at a lunch meeting, and a plate of food can be detected as an object within the user's field of view. When the object is detected, the method 300 can proceed from operation 302 to operation 304.

[0042] Operation 304 may include determining whether an object is associated with application content, which may be associated with a user wearing computerized glasses. In some implementations, determining whether an object and / or object classification associated with application content is associated with application content may be performed using one or more heuristic processes and / or one or more trained machine learning models. For example, object data representing an object may be processed using one or more trained machine learning models to generate an object embedding that may be mapped to a latent space, and the latent space may include entry embeddings. The entry embeddings may be generated from entries that have been generated by an automated assistant and / or another application to associate certain application content with certain objects and / or object classifications. For example, an entry embedding may correspond to an entry that associates an upcoming dinner reservation with a food object. Therefore, when the object embedding corresponding to the dish of food is mapped to the latent space, the object embedding may be mapped to within a threshold distance of the entry embedding, thereby demonstrating an association between the object embedding and the entry embedding.

[0043] When the object is determined to be associated with the application content, method 300 may proceed from operation 304 to operation 306. Otherwise, while the user continues to wear the computerized glasses, method 300 may return to operation 302. Operation 306 may include generating assistant data representing an image to be rendered at the computerized glasses. In some implementations, the assistant data may be generated before the object is within the field of view of the computerized glasses. In some implementations, the assistant data may be generated using one or more heuristic processes and / or one or more trained machine learning models. For example, when the object is the plate of food being served at a lunch meeting and the application content includes an upcoming dinner reservation, assistant data may be generated to convey helpful information to a particular user.

[0044] In some implementations, additional context may be considered when generating assistant data. For example, application content from a variety of different sources may be used as a basis for generating assistant data. Another source of application content may be, for example, a health application for tracking the nutrients a user eats each day. Assuming the user eats the plate of food at a lunch meeting, assistant data generated from the application content of the health application may convey information such as suggestions for menu items to select in a subsequent dinner reservation to facilitate assisting the user in reaching their daily nutrient goals. In this way, certain information may be inferred from the detected object and used in conjunction with one or more sources of application content to generate assistant data that may convey further helpful information.

[0045] In some implementations, the generated assistant data may include one or more images, natural language content, and / or any other content that can convey information to the user. For example, when generating assistant data for suggesting menu items selected for an upcoming dinner reservation, the assistant data may characterize an image of a menu item captured from a website of a restaurant with a dinner reservation. Alternatively or additionally, the assistant data may provide natural language content identifying a menu item and indicate the purpose of the automated assistant suggesting the menu item (e.g., "If you finish this lunch meal, you should select the Vegetable Curry at the dinner at Ramsi's tonight to reach your nutrient goal for today. (If you finish this lunch, you should select the vegetable curry at the dinner at Ramsi's tonight to reach your nutrient goal for today.)"). In some implementations, assistant data may be generated based on prior interactions between a user and an automated assistant. For example, assistant data may be generated to provide information that a user may have originally explicitly requested from an automated assistant based on context. For example, a user may have previously accessed their health app to determine what they should eat for dinner based on what they currently eat for lunch. During an interaction with a health app, a user may have asked their automated assistant to provide suggestions for what to eat for dinner (e.g., "Show me recipes for low-carb, high-protein dinners."). This prior interaction can serve as the basis for the automated assistant to generate assistant data regarding the nutrients in the menu items to be selected for a dinner reservation and / or the foods for a luncheon dinner plate.

[0046] From operation 306, method 300 may proceed to operation 308 to determine a location within the field of view of the computerized glasses for rendering an image. In some implementations, the image may be rendered adjacent to and / or at least partially overlapping an object. Alternatively or additionally, the image may be rendered with at least some transparent features, allowing the user to view a portion of the environment behind the image. In some implementations, the selection of the location for rendering the image may be based on the classification of the object, the application content associated with the object, and / or the type of information to be conveyed to the user. For example, an image that is meaningful because the placement of the image overlaps an object in the field of view of the computerized glasses may be rendered at a location that causes the image to at least partially overlap the object. Alternatively or additionally, an image that can appropriately convey information without the presence of an object may be rendered at any location within the field of view of the computerized glasses that does not obstruct the view of other objects (e.g., in a space that includes a threshold degree of color uniformity).

[0047] In some implementations, at least a portion of an image may be selected for rendering within the field of view of the computerized glasses based on one or more features of the user's context. For example, the amount of text of assistant data to be rendered at the display interface of the computerized glasses may be selected based on the user's distance from an object upon which the assistant data is based. Alternatively or additionally, the location of the image may be selected based on the user's preferences as indicated by user-controllable settings. For example, a user may explicitly request an indication of available assistant data at the periphery of the field of view of the computerized glasses. When the user adjusts the field of view to include more peripheral area (e.g., turns their head 90 degrees to their right), the image corresponding to the indication may be rendered entirely in the adjusted field of view of the computerized glasses.

[0048] From operation 308, method 300 may proceed to operation 310 of causing the computerized glasses to render an image at a display interface of the computerized glasses. Method 300 may optionally proceed from operation 310 to operation 312 of causing the image to be modified according to a detected change in the context of the user. For example, the image may be dynamically updated according to changes in application content that may have occurred since the image was rendered at the display interface of the computerized glasses. For example, the cancellation of the aforementioned dinner reservation as indicated by the calendar application content may cause the automated assistant to modify the image in real time. Thus, the image may be transformed from being a suggestion for a menu item to a suggestion for a recipe to cook at home, and / or an indication of the cancellation of the dinner reservation. In some implementations, the image may be selectable via input to the computerized glasses and / or the automated assistant. The content rendered in response to the selection of the image (e.g., the selection of the suggested menu item via a gesture captured by a sensor of the computerized glasses) may then be rendered according to objects that may be within the user's field of view and / or the user's preferences.

[0049] Figure 4 4 is a block diagram 400 of an example computer system 410. The computer system 410 generally includes at least one processor 414 that communicates with a number of peripheral devices via a bus subsystem 412. These peripheral devices may include a storage subsystem 424 (including, for example, a memory 425 and a file storage subsystem 426), a user interface output device 420, a user interface input device 422, and a network interface subsystem 416. The input and output devices allow a user to interact with the computer system 410. The network interface subsystem 416 provides an interface to an external network and is coupled to corresponding interface devices in other computer systems.

[0050] The user interface input device 422 may include a keyboard, a pointing device (such as a mouse, a trackball, a touch pad or a graphics tablet, a scanner, a touch screen integrated into a display), an audio input device (such as a voice recognition system, a microphone), and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways to input information into the computer system 410 or onto a communication network.

[0051] User interface output device 420 can include display subsystem, printer, fax machine, or non-visual display (such as audio output device).Display subsystem can include cathode ray tube (CRT), flat panel device such as liquid crystal display (LCD), projection device or some other mechanisms for creating visible image.Display subsystem can also provide non-visual display such as via audio output device.In general, the use of term "output device" is intended to include all possible types of devices and modes for outputting information from computer system 410 to user or another machine or computer system.

[0052] The storage subsystem 424 stores programming and data constructs that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 424 may include logic for performing selected aspects of the method 300 and / or implementing one or more of the system 200, computerized eyewear 124, computing device 104, automated assistant, and / or any other application, device, apparatus, and / or module discussed herein.

[0053] These software modules are typically executed by the processor 414 alone or in conjunction with other processors. The memory 425 used in the storage subsystem 424 may include multiple memories, including a main random access memory (RAM) 430 for storing instructions and data during program execution and a read-only memory (ROM) 432 in which fixed instructions are stored. The file storage subsystem 426 may provide persistent storage for program and data files, and may include a hard drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of certain implementations may be stored in the storage subsystem 424 by the file storage subsystem 426, or in other machines accessible to the processor 414.

[0054] The bus subsystem 412 provides a mechanism for the various components and subsystems of the computer system 410 to communicate with each other as intended. Although the bus subsystem 412 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple busses.

[0055] Computer system 410 can be of different types, including a workstation, server, computing cluster, blade server, server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, Figure 4 The description of the computer system 410 depicted in FIG. 4 is intended only as a specific example for the purpose of illustrating some implementations. Many other configurations of the computer system 410 are possible, with similar configurations. Figure 4 The computer systems may have more or fewer components than those depicted.

[0056] In cases where the systems described herein collect personal information about users (or "participants" as often referred to herein) or may make use of personal information, users may be provided with the opportunity to control whether a program or feature collects user information (e.g., information about a user's social network, social behavior or activities, occupation, a user's preferences, or a user's current geographic location), or to control whether and / or how content is received from a content server that may be more relevant to the user. In addition, certain data may be processed in one or more ways before it is stored or used so that personally identifiable information is removed. For example, the identity of a user may be processed so that personally identifiable information about the user cannot be determined, or, in the case of obtaining geographic location information, the geographic location of the user may be generalized (such as to a city, zip code, or state level) so that the specific geographic location of the user cannot be determined. Thus, users may control how information about the user is collected and / or used.

[0057] Although several implementations have been described and illustrated herein, a variety of other means and / or structures for performing functions and / or obtaining results and / or one or more of the advantages described herein may be utilized, and each of these changes and / or modifications is considered to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications for which this teaching is used. Those skilled in the art will recognize or confirm many equivalents of the implementations described herein using only routine experiments. Therefore, it should be understood that the aforementioned implementations are presented only by way of example, and it should be understood that within the scope of the attached claims and their equivalents, implementations may be practiced in a manner other than specifically described and claimed. The implementations disclosed herein relate to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more of these features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure if these features, systems, articles, materials, kits, and / or methods do not contradict each other.

[0058] In some implementations, a method implemented by one or more processors is described as including operations such as: determining, based on sensor data, that a field of view of computerized glasses being worn by a user includes an object. The sensor data is generated using one or more sensors that are integrated with (e.g., included as a part of) the computerized glasses and / or integrated with a separate computing device that communicates with the computerized glasses. The method may further include: determining, based on contextual data, that the object is associated with an application accessible via the computerized glasses and / or the separate computing device. The contextual data indicates that the user interacted with the application within a threshold duration when the object was within the field of view of the computerized glasses. The method may further include: generating assistant data representing an image for rendering at a display interface of the computerized glasses based on the object being within the field of view of the computerized glasses. The image is rendered to convey information based on the user interacting with the application. The method may further include: causing the computerized glasses to render the image at the display interface of the computerized glasses when the object is within the field of view of the computerized glasses based on the assistant data. The image is rendered at a location within the field of view of the computerized glasses that is selected based on the location of the object within the field of view of the computerized glasses.

[0059] In some implementations, the context data further indicates that the user interacted with the application to advance completion of a task identified by a calendar application accessible to the user, and the object includes at least a portion of a calendar. In some implementations, the context data further indicates a deadline for the task identified by the calendar application, and the position of the image is selected based on a location indicated by the deadline on the calendar. In some implementations, the context data further indicates a deadline for the task identified by the calendar application, and the size of the image is selected based on a location indicated by the deadline on the calendar. In some implementations, the context data further indicates a date of an upcoming event identified by the calendar application, and the position of the image is selected to not overlap with a location indicated by the date of the upcoming event on the calendar. In some implementations, the image is rendered at the location further selected based on one or more physical properties of the object.

[0060] In some implementations, the method further includes: selecting the location for rendering the image based on the size of the object within the field of view of the computerized glasses. For example, the location can be selected to prevent the image from overlapping a majority of the object within the field of view of the computerized glasses. As another example, the location can be selected to cause the image to be adjacent to the object within the field of view of the computerized glasses. In some implementations, generating the assistant data characterizing the image is further based on one or more physical properties of the object. In some implementations, the one or more physical properties of the object include the size of an area of ​​the object that exhibits a threshold degree of color uniformity. In some implementations, the information included in the image includes natural language content, and the amount of natural language content to be included in the image is selected based on the size of the area of ​​the object that exhibits the threshold degree of color uniformity.

[0061] In other implementations, a method implemented by one or more processors is described as including operations such as: generating an entry that associates content determined based on an interaction with an application on a computing device with a classification of a physical object. The method may further include, after generating the entry that associates the content with the classification of the physical object: determining, based on sensor data generated by one or more sensors of the computerized glasses being worn by the user, that a particular physical object is present in the field of view of the computerized glasses and has the classification of the entry. The method may further include, in response to determining that the particular physical object is present in the field of view of the computerized glasses and has the classification of the entry: causing the content associated with the classification in the entry to be rendered at a display interface of the computerized glasses so that the content is displayed simultaneously with the particular physical object that is located in the field of view of the computerized glasses.

[0062] In some implementations, the content identifies a geographic location, and causing the computerized glasses to render the content is performed when the user is within a threshold distance of the geographic location. In some implementations, the content includes a selectable element selectable via an automated assistant, the automated assistant accessible via the computerized glasses, and selection of the selectable element causes the automated assistant to interact with the application. In some implementations, the method may further include: determining, based on the sensor data, specific information being conveyed by the physical object when the specific physical object is within the field of view of the computerized glasses. In some of those implementations, the content is based on: the specific information being conveyed by the specific physical object, and the interaction between the user and the application. In some implementations, the specific information includes an indication of a current time, and the content includes another indication of an amount of time until an event associated with the application.

[0063] In yet other implementations, a method implemented by one or more processors is set forth as including operations such as: determining that a particular physical object is present in the field of view of the computerized glasses being worn by a user based on sensor data generated by one or more sensors of the computerized glasses. The method may further include: generating an entry associating the particular physical object with a classification of physical objects based on the sensor data. The method may further include, after generating the entry associating the particular physical object with the classification of physical objects: determining that the user is accessing application content associated with the classification of physical objects via a separate computing device or the computerized glasses based on contextual data associated with the user. The method may further include: determining that the particular physical object or another physical object having the classification of physical objects is located within the field of view of the computerized glasses after the user accesses the application content associated with the particular physical object. The method may further include: causing an image associated with the application content to be displayed at a display interface of the computerized glasses based on the contextual data and the particular physical object or other physical objects being located within the field of view of the computerized glasses.

[0064] In some implementations, determining that the particular physical object or the other physical object is within the field of view of the computerized glasses includes determining that the particular physical object or the other physical object is within the field of view of the computerized glasses for a threshold duration since the user accessed the application content. In some implementations, determining that the particular physical object or the other physical object is within the field of view of the computerized glasses includes determining that the particular physical object or the other physical object is within the field of view of the computerized glasses for a duration specified by the user for receiving notifications corresponding to the classification of physical objects. In some implementations, the image is selectable via an interface of the computerized glasses, and selection of the image by the user causes the application content to be edited via the computerized glasses or the separate computing device.

Claims

1. A method implemented by one or more processors, the method comprising: determining based on the sensor data that a field of view of the computerized eyewear being worn by the user includes the object, wherein the sensor data is generated using one or more sensors integrated with the computerized eyewear and / or integrated with a separate computing device in communication with the computerized eyewear; determining, based on contextual data, that the object is associated with an application accessible via the computerized eyewear and / or the separate computing device, wherein the contextual data indicates that the user interacted with the application for a threshold duration in which the object was within the field of view of the computerized eyewear; generating assistant data representing an image for rendering at a display interface of the computerized glasses based on the object being within the field of view of the computerized glasses, wherein the image is rendered to convey information based on the user interacting with the application; as well as causing the computerized glasses to render the image at the display interface of the computerized glasses when the object is within the field of view of the computerized glasses based on the assistant data, Wherein the image is rendered at a location within the field of view of the computerized glasses that is selected based on the location of the object within the field of view of the computerized glasses.

2. The method of claim 1, wherein the context data further indicates that the user interacted with the application to advance a task identified by a calendar application accessible to the user, and the object comprises at least a portion of a calendar.

3. The method of claim 2, wherein the contextual data further indicates a due date for the task identified by the calendar application, and the position of the image is selected based on a position indicated by the due date on the calendar.

4. The method of claim 2 or claim 3, wherein the context data further indicates a due date for the task identified by the calendar application, and the size of the image is selected based on a location indicated by the due date on the calendar.

5. A method as claimed in any one of claims 2 to 4, wherein the context data further indicates a date of an upcoming event identified by the calendar application, and the position of the image is selected so as not to overlap with a position indicated on the calendar for the date of the upcoming event.

6. A method as claimed in any preceding claim, wherein the image is rendered at the location further selected based on one or more physical properties of the object.

7. The method of any preceding claim, further comprising: selecting the location for rendering the image based on a size of the object within the field of view of the computerized eyewear, Wherein the position is selected to prevent the image from overlapping a substantial portion of the object within the field of view of the computerized glasses.

8. The method of any preceding claim, further comprising: selecting the location for rendering the image based on a size of the object within the field of view of the computerized eyewear, Wherein the position is selected to cause the image to be adjacent to the object within the field of view of the computerized glasses.

9. A method as claimed in any preceding claim, wherein generating the assistant data characterizing the image is further based on one or more physical properties of the object.

10. The method of claim 9, wherein the one or more physical properties of the object include a size of a region of the object that exhibits a threshold degree of color uniformity.

11. The method of claim 10, wherein the information included in the image includes natural language content, and the amount of natural language content to be included in the image is selected based on the size of the area of ​​the object that exhibits the threshold degree of color uniformity.

12. A method implemented by one or more processors, the method comprising: generating, based on an interaction of a user with an application on a computing device, an entry associating content determined based on the interaction with a classification of physical objects; After generating the entry associating the content with the category of physical objects: determining, based on sensor data generated by one or more sensors of computerized eyewear being worn by the user, that a particular physical object is present in a field of view of the computerized eyewear and has the classification of the entry; as well as In response to determining that the particular physical object is present in the field of view of the computerized eyewear and has the classification of the entry: The content associated with the category in the entry is caused to be rendered at a display interface of the computerized glasses such that the content is displayed simultaneously with the particular physical object being located in the field of view of the computerized glasses.

13. The method of claim 12, wherein the content identifies a geographic location, and causing the computerized eyewear to render the content is performed when the user is within a threshold distance of the geographic location.

14. The method of claim 12 or claim 13, wherein the content includes selectable elements selectable via an automated assistant accessible via the computerized eyewear, and Wherein selection of the selectable element causes the automated assistant to interact with the application.

15. The method of any one of claims 12 to 14, further comprising: determining, based on the sensor data, specific information being conveyed by the specific physical object while the specific physical object is within the field of view of the computerized eyewear, The content is based on: the specific information being conveyed by the specific physical object, and the interaction between the user and the application.

16. The method of claim 15, wherein the specific information includes an indication of a current time and the content includes another indication of an amount of time until an event associated with the application.

17. A method implemented by one or more processors, the method comprising: determining, based on sensor data generated by one or more sensors of the computerized glasses being worn by the user, that a particular physical object is present in a field of view of the computerized glasses; generating, based on the sensor data, an entry associating the particular physical object with a classification of physical objects; After generating said entry associating said particular physical object with said category of physical objects: determining, based on contextual data associated with the user, that the user is accessing application content associated with the classification of physical objects via a separate computing device or the computerized eyewear; determining that the particular physical object or another physical object having the classification of physical objects is located within the field of view of the computerized eyewear after the user accesses the application content associated with the particular physical object; as well as An image associated with the application content is caused to be displayed at a display interface of the computerized glasses based on the contextual data and the particular physical object or other physical objects being located within the field of view of the computerized glasses.

18. The method of claim 17, wherein determining that the particular physical object or the other physical objects are located within the field of view of the computerized glasses comprises: A determination is made that the particular physical object or the other physical object is within the field of view of the computerized glasses within a threshold duration since the user accessed the application content.

19. The method of claim 17 or claim 18, wherein determining that the particular physical object or the other physical objects are located within the field of view of the computerized glasses comprises: A determination is made that the particular physical object or the other physical objects are within the field of view of the computerized glasses for a duration specified by the user for receiving notifications corresponding to the classification of physical objects.

20. The method of any one of claims 17 to 19, wherein the image is selectable via an interface of the computerized glasses, and selection of the image by the user causes the application content to be edited via the computerized glasses or the separate computing device.

21. A system comprising: one or more hardware processors; as well as A memory storing instructions which, when executed by the one or more hardware processors, cause the one or more hardware processors to perform the method of any one of claims 1 to 20.

22. A non-transitory computer-readable storage medium storing instructions which, when executed by one or more hardware processors, cause the one or more hardware processors to perform the operations of the method according to any one of claims 1 to 20.