Systems and methods for generating and providing suggested actions
By using a memory-based context storage AI system, which utilizes machine learning models to process contextual data to generate and store semantic entities, the problem of existing AI agents being unable to actively remember user actions is solved, resulting in more efficient task completion and a better user experience.
Patent Information
- Application Number
- CN202511508879.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2019-08-02
- Publication Date
- 2026-02-10
AI Technical Summary
Existing AI agents and personal assistants lack the ability to proactively remember user actions or items, thus failing to effectively help users complete tasks.
An AI system employing memory-based context storage processes contextual data through machine learning models to generate and store semantic entities, and then provides suggested actions to help users complete tasks.
It improves users' task completion efficiency, reduces users' memory burden, and helps users complete tasks that were started but not finished earlier by providing meaningful suggested actions at the right time.
Smart Images

Figure CN121503728A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on August 2, 2019, with Chinese application number 201980098094.1 and entitled "System and method for generating and providing suggested actions". Technical Field
[0002] This disclosure generally relates to artificial intelligence systems. More specifically, this disclosure relates to systems and methods for generating and providing suggested actions to users of computing devices. Background Technology
[0003] Artificial intelligence and machine learning have been used to assist users of computing devices, for example, by providing AI agents and personal assistants. However, these AI agents and personal assistants lack the ability to proactively help users remember actions or items. Summary of the Invention
[0004] Aspects and advantages of embodiments of this disclosure will be set forth in part in the description which follows, or may be learned from the description or by practice of the embodiments.
[0005] One aspect of this disclosure relates to a computing system comprising: an artificial intelligence system having a memory-based context storage, wherein the artificial intelligence system includes one or more machine learning models, and wherein the artificial intelligence system is configured to perform operations including: during a first time interval: receiving model input including context data by the artificial intelligence system; processing the model input by the artificial intelligence system using the one or more machine learning models to generate model output describing one or more semantic entities referenced by the context data; storing the model output in memory; and during a second time interval following the first time interval: providing suggested actions regarding the one or more semantic entities described by the model output based on the model output stored in memory.
[0006] Another aspect of this disclosure relates to a non-transitory computer-readable medium, comprising: an artificial intelligence system having a memory-based context storage, wherein the artificial intelligence system includes one or more machine learning models, and wherein the artificial intelligence system is configured to perform operations including: during a first time interval: receiving model input including context data by the artificial intelligence system; processing the model input by the artificial intelligence system using the one or more machine learning models to generate model output describing one or more semantic entities referenced by the context data; storing the model output in memory; and during a second time interval after the first time interval: providing suggested actions regarding the one or more semantic entities described by the model output based on the model output stored in memory.
[0007] Another aspect of this disclosure relates to a computer-implemented method for providing suggested actions to a user, the computer-implemented method comprising: during a first time interval: receiving model input including context data by a computing system including an artificial intelligence system having a memory-based context storage, the artificial intelligence system including one or more machine learning models; processing the model input by the artificial intelligence system using the one or more machine learning models to generate model output describing one or more semantic entities referenced by the context data; storing the model output in memory; and during a second time interval following the first time interval: providing suggested actions regarding the one or more semantic entities described by the model output based on the model output stored in memory.
[0008] Another aspect of this disclosure relates to a computing system including at least one processor and an artificial intelligence system including one or more machine learning models. The one or more machine learning models can be configured to receive model input including context data, and in response to the receipt of the model input, output a model output describing one or more semantic entities referenced by the context data. The computing system may include at least one tangible, non-transitory computer-readable medium storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations. Operations may include, during a first time interval, obtaining the context data; inputting model input including the context data into the one or more machine learning models; receiving the model output describing one or more semantic entities referenced by the context data as the output of the one or more machine learning models; storing the model output in at least one tangible, non-transitory computer-readable medium; and, during a second time interval following the first time interval, providing suggested actions regarding the one or more semantic entities described by the model output for display in a user interface.
[0009] Another aspect of this disclosure relates to a computer-implemented method for generating and providing suggested actions. The method may include obtaining context data by one or more computing devices during a first time interval. The method may include having one or more computing devices input model input, including context data, into one or more machine learning models configured to receive the model input including the context data, and, in response to receiving the model input, output a model output describing one or more semantic entities referenced by the context data. The method may include having one or more computing devices receive the model output describing one or more semantic entities referenced by the context data as the output of one or more machine learning models. The method may include having one or more computing devices store the model output in at least one tangible, non-transitory computer-readable medium. The method may include having one or more computing devices provide suggested actions regarding one or more semantic entities described by the model output during a second time interval following the first time interval, for display in a user interface of one or more computing devices.
[0010] Other aspects of this disclosure relate to a variety of systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.
[0011] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and form a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. Attached Figure Description
[0012] Referring to the accompanying drawings, a detailed discussion of embodiments for those skilled in the art is set forth in the specification, wherein:
[0013] Figure 1A A block diagram of an example computing system for generating and providing suggested actions to a user of the computing system, according to an example embodiment of the present disclosure, is depicted.
[0014] Figure 1B A block diagram of an example computing system for generating and providing suggested actions to a user of the computing system, according to an example embodiment of the present disclosure, is depicted.
[0015] Figure 1C A block diagram of an example computing system for generating and providing suggested actions to a user of the computing system, according to an example embodiment of the present disclosure, is depicted.
[0016] Figure 2 An example artificial intelligence system for generating and providing suggested actions is described according to an example embodiment of the present disclosure.
[0017] Figure 3An example computing system for generating and providing suggested actions according to an example embodiment of the present disclosure is described, which includes one or more computer applications.
[0018] Figure 4 Example suggested actions are described according to aspects of this disclosure.
[0019] Figure 5A , Figure 5B and Figure 5C Additional example suggested actions are described according to aspects of this disclosure.
[0020] Figure 6 An example panel depicting multiple suggested actions displayed on the lock screen of a computing device, according to aspects of this disclosure, is described.
[0021] Figure 7 A computing device is depicted displaying an example notification panel that shows suggested actions using notifications.
[0022] Figure 8 A computing device in a first state according to aspects of this disclosure is depicted, in which multiple category labels corresponding to suggested actions for classification are displayed.
[0023] Figure 9 The description depicts a scenario where, according to aspects of this disclosure, one of a plurality of category labels has been selected and a suggested action corresponding to the selected category label is displayed. Figure 8 Computing devices.
[0024] Figure 10 The proposed actions according to aspects of this disclosure are described, wherein a semantic entity has been selected and a search is performed in response to the selection of the semantic entity.
[0025] Figure 11 The description depicts a computing system that displays a settings panel where users can select a default computer application for a type of suggested action.
[0026] Figure 12 A flowchart is depicted illustrating a method for generating and providing suggested actions to a user of a computing system, according to aspects of this disclosure. Detailed Implementation
[0027] Overview
[0028] Generally, this disclosure relates to an artificial intelligence system for identifying information of interest, storing that information, and providing suggested actions to a user of a computing system at a later time based on the stored information. The AI system can be configured to intelligently process information on behalf of the user, including, for example, visual and / or audio information displayed, played, and / or otherwise processed or detected by a computing device. In other words, the AI system can capture information of interest while the computing device is used throughout the day to perform tasks. For example, while a user navigates between multiple computer applications and / or switches between different tasks or activities, the AI system can identify and store semantic entities. Alternatively, the AI system can identify and store semantic entities referenced or included by the user's surrounding environment (e.g., by analyzing captured imaging, audio, or other data about the surrounding environment). Thus, the AI system can capture and process information actively identified or emphasized by the user (e.g., for identifying semantic entities), while in other instances, the AI system can capture and process information simply referenced or included by the user's ambient environment (e.g., for identifying semantic entities) (e.g., information contained in the surrounding environment but not specifically identified or emphasized by the user).
[0029] As semantic entities are identified over time, AI systems can save or otherwise retain data associated with those entities. For example, saved semantic entities can be ranked, sorted, categorized, prioritized, etc., based on user preferences and / or user plans or schedules. As another example, an AI system can generate one or more suggested actions for a user related to one or more identified semantic entities. For example, suggested actions can include actions that can be taken by the AI system and / or computer applications, guided by the AI system, for the user and / or on behalf of the user, relative to the identified semantic entities. As examples, suggested actions can include communication actions (e.g., sending an email to a contact), information retrieval actions (e.g., searching for options to buy or shop for an item, offering an opportunity to listen to a song, accessing geographic information such as the location of a point of interest), booking actions (e.g., requesting a ride on shared transportation or buying a plane ticket), information storage (e.g., taking notes or inserting an item into the user's calendar), and / or many other suggested actions.
[0030] Saved suggested actions can be provided for display at a later time. For example, suggested actions may be accessible to the user via a specific menu, provided in a notification menu, appear automatically at a later, context-sensitive time, and / or be accessed in other ways. Suggested actions may include links or buttons to perform the suggested action (e.g., in a computer application). Users can also optionally provide feedback and / or instructions to the AI system to customize how the AI system captures information and / or suggested actions. Optionally, the AI system can also learn user preferences based on how the user interacts with suggested actions.
[0031] Importantly, users can be given control over whether and when the systems, programs, or features described herein can enable the collection of user information (e.g., external audio, text presented in the user interface, etc.). Furthermore, some data may be processed in one or more ways before it is stored or used, resulting in the removal of personally identifiable information. For example, a user's identity may be processed in such a way that no personally identifiable information about that user can be determined. Therefore, users can control what information about themselves is collected, how that information is used, and what information is provided to them.
[0032] This disclosure relates to an artificial intelligence system that operates over multiple distinct time periods to provide suggested actions that are meaningful within a context. Specifically, the AI system can identify and store semantic entities during a first time interval, and then display suggested actions during a second time interval following the first. For example, the suggested actions can act as a reminder to the user during the second time interval to complete a task the user started earlier. In this way, aggregating relevant information during the first time interval and then providing multiple suggested actions based on that information during the second and later time intervals can be less disruptive to the user. This is also more useful for the user, as they are more likely to forget the task after some time. Therefore, by storing suggested and context-derived actions for later use, the AI system can function as an intelligent memory assistant, helping the user remember actions they might want to take based on activities they participated in earlier that day, week, month, etc.
[0033] As an example, a user can take a "screenshot" of an item while shopping during a first time interval. In response to this user action, the AI system can generate and store the item's name or description. Later, the user can suggest a specific meeting time and date during a phone call. The AI system can generate and store a second semantic entity that may include names, meeting times, etc. During a second time interval (e.g., after get off work, after dinner, etc.), the AI system can display suggested actions for each item from the screenshot and suggested meeting. Suggested actions for items could include purchasing the item shown in the screenshot, and suggested actions for meetings could include creating a calendar event based on information gathered during the phone call.
[0034] In some implementations, AI systems can generate (multiple) suggested actions based on templates (e.g., predefined templates). Using templates reduces the computational resources required to generate such suggested actions. Instead of training and utilizing machine learning models to generate complete suggested actions, keywords for the suggested actions can be generated using (multiple) machine learning models and then assembled based on the template to generate the suggested action. For example, the template may include verbs, semantic entities described by the model output of the machine learning model, and computer applications, such as:
[0035] [Verb] + [Semantic Entity] + [Computer Application]
[0036] The AI system can select appropriate verbs and computer applications from the corresponding placeholders to be inserted into the template. An example of a suggested action generated based on the template described above is "Add an appointment with Dr. Sherrigan to the calendar." It should be understood that multiple templates can be used. Furthermore, the AI system can learn the user's preferences regarding which templates to use and / or whether to use them. Therefore, the system and method described herein can employ one or more templates to generate suggested actions.
[0037] In some implementations, the same one or more templates can be used to generate multiple suggested actions, ensuring that these actions have the same overall look and feel to the user. This allows users to quickly evaluate suggested actions because they are presented in a known and predictable format. Therefore, the templates described in this paper can facilitate wider use of suggested actions.
[0038] According to aspects of this disclosure, the systems and methods described herein can utilize one or more machine learning models. More specifically, the artificial intelligence system may include one or more machine learning models configured to receive model inputs including contextual data (e.g., external audio, information displayed on a screen of a computing device, etc.). The computing system may be configured to acquire the contextual data and input the model inputs including the contextual data into the machine learning model. The computing system may receive model outputs describing one or more semantic entities referenced by the contextual data as the output of the machine learning model. The computing system may store the model outputs in at least one tangible, non-transitory computer-readable medium. The computing system may provide suggested actions regarding one or more semantic entities.
[0039] The contextual data discussed herein can include a variety of information, such as information currently displayed in the user interface, information previously displayed in the user interface, information extracted from the user's previous actions (e.g., text written or read by the user, content viewed by the user, etc.), and / or the like. Contextual data can include user data describing preferences or other information associated with the user and / or contact data describing preferences or other information associated with the user's contacts. Example contextual data can include messages received by the computing system for the user, the user's previous interactions with one or more of the user's contacts (e.g., text messages mentioning the user's preferences for restaurants or food types), location-related previous interactions (e.g., visiting parks, museums, other attractions, etc.), business-related interactions (e.g., posting reviews of restaurants, reading restaurant menus, making reservations at restaurants, etc.), and / or any other suitable information about the user's preferences or the user. Further examples include audio played or processed by the computing system, audio detected by the computing system, information about the user's location (e.g., the location of the computing system's mobile computing device), and / or calendar data. For example, contextual data can include external audio detected by the computing system's microphone and / or telephone audio processed during a telephone call. Calendar data can describe future events or plans (e.g., flights, hotel bookings, dinner plans, etc.). Example semantic entities that can be described by the model output can include words or phrases recognized in text and / or audio. Additional examples can include information about the user's location, such as city name, state name, street name, names of nearby attractions, etc.
[0040] In some implementations, after collecting context data during a first time interval, multiple suggested actions can be displayed together in a second time interval. More specifically, at least one additional suggested action can be displayed alongside the suggested action in the user interface. These additional suggested actions can be distinguished from the suggested action. These additional suggested actions can also be generated and stored by the AI system based on distinguished semantic entities (e.g., during the first time interval). For example, the AI system can obtain additional context data that is distinguished from the context data and input (multiple) additional model inputs, including the additional context data, into (multiple) machine learning models. The AI system can receive data describing (multiple) additional suggested actions as additional outputs of (multiple) machine learning models. Therefore, the AI system can utilize (multiple) machine learning models to store multiple semantic entities during the first time interval and then display multiple suggested actions during the second time interval.
[0041] In some implementations, AI systems can, for example, rank, categorize, and prioritize suggested actions based on user data. User data may include user preferences, calendar data (e.g., a user's plans or schedules), and / or other information about the user or computing device. For instance, an AI system can rank suggested actions and arrange them within the user interface based on that ranking. Therefore, an AI system can prioritize suggested actions and selectively display a set of the most important and / or relevant suggested actions to the user in the user interface.
[0042] In some implementations, AI systems can categorize suggested actions across multiple categories. Category labels can be displayed in the user interface corresponding to the category, allowing users to navigate between categories of suggested actions using these labels (e.g., in separate panels or pages). For example, the computational system can detect user touch actions related to a category label. In response to detecting a user touch action, the computational system can display suggested actions categorized according to the selected suggested action. Thus, the AI system can categorize suggested actions and provide users with an intuitive way to navigate between categories of suggested actions.
[0043] In some implementations, the computing system can display an explanation of the suggested action. The explanation can describe information about the acquisition of contextual data, including, for example, the time the contextual data was acquired, the location of the computing device at the time the contextual data was acquired, and / or the source of the contextual data. As an example, the explanation could indicate that the suggested action was generated based on audio of a telephone conversation with a specific user contact that occurred at a specific time. As another example, the explanation could indicate that the suggested action was generated based on a user's shopping session using a specific shopping app at a specific time. The sources of the data could include the computer application being displayed when the contextual data was acquired; whether the contextual data was acquired from external audio, text, graphical information, etc.; and / or any other information associated with acquiring the contextual data. Such explanations can provide users with a better understanding of the operation of the artificial intelligence system. Therefore, users may feel more comfortable or trusting of the operation of the artificial intelligence system, which can make the artificial intelligence system more useful.
[0044] In some implementations, in addition to the explanations, the computing system can provide the user with a way to view additional information associated with the acquisition of context data. For example, the computing system can detect user touch input requesting additional information about the explanations. In response to detecting the user touch input, the computing system can display additional explanatory information about the acquisition of the context data. The additional explanatory information may include the time when the context data was acquired or the location of the computing device at the time the context data was acquired. The additional explanatory information may include the source of the context data (if it has not already been displayed). The additional information may include information about other times when the context data was acquired in a similar manner (e.g., from the same source, at a similar time, etc.).
[0045] In some implementations, the computing system can provide users with ways to adjust their preferences regarding how the AI system collects contextual data. Additional explanatory information may also include preferences and / or rules regarding when and how the AI system can obtain contextual data. Users can adjust these rules and / or preferences within the user interface.
[0046] The computing system can automatically or in response to user requests display (multiple) suggested actions in multiple locations. For example, a panel displaying (multiple) suggested actions could appear on the computing device's lock screen or home screen. The panel can be accessed at the system level from dropdown menus, navigation bars, etc. The panel can be automatically displayed at one or more regular times of the day. In other implementations, the artificial intelligence system can intelligently select when to display the panel based on user data (e.g., preferences) and / or contextual data. The artificial intelligence system can also display the panel when the suggested action is most relevant to the user, based on its content.
[0047] In some implementations, the computing system can be configured to interface with one or more computing applications to provide suggested actions that can be performed by the applications. The computing system can provide the applications with data descriptions of the model outputs of the machine learning models that describe the suggested actions. The computing system can receive one or more application outputs from the applications via predefined application programming interfaces. The suggested actions can describe at least one of the application outputs.
[0048] In some implementations, a user can select a portion of a suggested action (e.g., a semantic entity) to perform a separate action distinct from the suggested action, related to the selection of the suggested action. As an example, in a user response that detects a user touch action pointing to a semantic entity of the suggested action, the computing system can display a panel that includes a search for the semantic entity (e.g., a web search). Additional examples include editing the semantic entity, changing the computer application of the suggested action, editing the details of the suggested action, etc. For example, a suggested action could include shopping for a product from a specific brand (e.g., a grill) using a specific shopping app. A user can select a semantic entity (e.g., "Webster Classic Grill") and manually edit the entity, such as changing the brand name or product type. A user can change "Webster Classic Grill" to "Webster Classic Grill Cover" before selecting a suggested action to buy an item or shop using a shopping app. As another example, a user can change the shopping app.
[0049] As an example, the systems and methods of this disclosure may be included within, or otherwise employed within, an application, a browser plugin, or other context. Therefore, in some implementations, the model of this disclosure may be included in, or otherwise stored and implemented by, a user computing device such as a laptop, tablet, or smartphone. As yet another example, the model may be included in, or otherwise stored and implemented by, a server computing device communicating with the user computing device according to a client-server relationship. For example, the model may be implemented by the server computing device as part of a web service (e.g., a web email service).
[0050] For example, the systems and methods disclosed herein can operate at the operating system level, rather than at the level of one or more specific applications that require user selection to initiate their operations. For instance, during normal use of the system, context data can be automatically acquired from one or more sources on the system, such as a microphone, camera, browsed web pages, the system's location, and the system's orientation, without requiring the user to open a specific application. Context data can be acquired whenever the system is powered on and / or when an appropriate password, passphrase, and / or biometric authentication is performed. Therefore, the user does not need to remember to activate specific functions to acquire context data. Context data can be automatically stored on the system-level "clipboard." Users can disable certain sources of context data if needed.
[0051] For example, the systems and methods of this disclosure can limit the amount of data stored in memory and / or allocated to such systems and methods. For example, the systems and methods can automatically remove or overwrite acquired context data and / or suggested actions based on one or more rules. For example, context data earlier than a predetermined period (e.g., one day or one week earlier) can be automatically removed or overwritten. Rules (e.g., predetermined periods) can be user-configurable. Predetermined periods can be dynamically updated for specific context data and / or suggested actions based on prior user interactions. For example, prior user interactions can be user choices associated with suggested actions. For example, if a user chooses to create a calendar appointment from a suggested action more frequently than opening a shopping app from a suggested action, the context data linked to the shopping action can be deleted or overwritten earlier than the context data used for the calendar appointment. A maximum limit can be set on the amount of data stored to avoid consuming application data storage resources. When suggested actions are ranked by priority, only the top N suggested actions can be retained for output.
[0052] For example, by utilizing at least one optional button or similar object displayed alongside or otherwise together with the suggested action, the systems and methods of this disclosure can provide one or more suggested actions linked to or associated with one or more applications, allowing a single or reduced number of clicks, touches, or swipes to influence the suggested action. Here, the number of physical user interactions influencing a suggested action (such as adding an appointment to a calendar event, playing a music or video track, opening a shopping website, etc.) may require less user interaction with the user interface and / or use less power and processing resources than opening a specific application and making a selection and / or manually typing data. For example, opening a shopping application or website for purchasing a specific product can be achieved with a single gesture, rather than opening the relevant shopping application or browser window, typing a search term, and then selecting an item from a list of suggestions.
[0053] For example, suggested actions can be displayed in the relevant section of the user interface, and in addition to optional buttons associated with launching the suggested action, one or more further buttons can be presented in the same section for single-touch launch of related functions that may typically require opening the application. For example, for suggested actions involving playing audio or video, and suggested actions involving opening a music or video application to play audio or video, one or more further buttons associated with the application, such as a "like" button or a "share" button, can be presented, and selecting these causes the associated action to be performed. For example, for suggested actions involving creating a calendar appointment, if a conflict is detected between the proposed appointment and an existing appointment, one or more further buttons can be provided to initiate cancellation and / or rescheduling of the existing appointment.
[0054] Where an AI system intelligently selects when to display a panel based on user data (e.g., user preferences) and / or contextual data, the AI system can display the panel when the suggested action is most relevant to the user or least intrusive or disruptive, based on the content of the suggested action. For example, if the suggested action includes playing audio or video, the AI system can choose not to display such suggested actions during working hours, while "silent" suggested actions such as suggesting appointments or shopping can be displayed at those times. As described above, summarizing suggested actions into a later, second time interval avoids excessively disturbing and / or intrusive to the user.
[0055] For example, the systems and methods disclosed herein can operate at the operating system level, rather than at the level of one or more specific applications that require user selection to initiate their operations. For instance, during normal use of the system, context data can be automatically acquired from one or more sources on the system, such as a microphone, camera, browsed web pages, the system's location, and the system's orientation, without requiring the user to open a specific application. Context data can be acquired whenever the system is powered on and / or when an appropriate password, passphrase, and / or biometric authentication is performed. Therefore, the user does not need to remember to activate specific functions to acquire context data. Context data can be automatically stored on the system-level "clipboard." However, the user can disable certain sources of context data if needed.
[0056] For example, the systems and methods of this disclosure can limit the amount of data stored in memory and / or allocated to such systems and methods. For example, the systems and methods can automatically remove or overwrite acquired context data and / or suggested actions based on one or more rules. For example, context data earlier than a predetermined period (e.g., one day or one week earlier) can be automatically removed or overwritten. Rules (e.g., predetermined periods) can be user-configurable. Predetermined periods can be dynamically updated for specific context data and / or suggested actions based on prior user interactions. For example, prior user interactions can include user selections associated with suggested actions. For example, if a user selects to create a calendar appointment from a suggested action more frequently than from opening a shopping app, context data linked to the shopping action may be deleted or overwritten earlier than context data used for calendar appointments. A maximum limit can be set on the amount of data stored to avoid consuming application data storage resources. When suggested actions are ranked by priority, only the top N suggested actions can be retained for output.
[0057] For example, by utilizing at least one optional button or similar object displayed alongside or otherwise together with the suggested action, the systems and methods of this disclosure can provide one or more suggested actions linked to or associated with one or more applications, allowing a single or reduced number of clicks, touches, or swipes to influence the suggested action. Here, the number of physical user interactions influencing a suggested action (such as adding an appointment to a calendar event, playing a music or video track, opening a shopping website, etc.) may require less user interaction with the user interface and / or use less power and processing resources than opening a specific application and making a selection and / or manually typing data. For example, opening a shopping application or website for purchasing a specific product can be achieved with a single gesture, rather than opening a related shopping application or browser window, typing a search term, and then selecting an item from a list of suggestions.
[0058] For example, suggested actions can be displayed in the relevant section of the user interface, and in addition to optional buttons associated with launching the suggested action, one or more further buttons can be presented in the same section for one-touch launch of related functions that would typically require the user to manually open the application. For example, for suggested actions involving playing audio or video, and for suggesting actions involving opening a music or video application to play audio or video, one or more further buttons associated with the application, such as a "like" button or a "share" button, can be presented. The selection of such buttons allows the associated action to be performed. For example, for suggested actions involving creating a calendar appointment, if a conflict is detected between the proposed appointment and an existing appointment, one or more further buttons can be provided to initiate cancellation and / or rescheduling of the existing appointment.
[0059] Where an AI system intelligently selects when to display a panel based on user data (e.g., preferences) and / or contextual data, the AI system can display the panel when the suggested action is most relevant to the user or least intrusive or disruptive, based on the content of the suggested action. For example, if the suggested action includes playing audio or video, the AI system can choose not to display such suggested actions during working hours, while "silent" suggested actions such as suggesting appointments or shopping can be displayed during those times. As described above, summarizing suggested actions into a later, second time interval avoids excessively disturbing and / or intrusive to the user.
[0060] Exemplary embodiments of this disclosure will now be discussed in further detail with reference to the accompanying drawings.
[0061] Example devices and systems
[0062] Figure 1A A block diagram of an example computing system 100 for generating and providing suggested actions according to an example embodiment of the present disclosure is depicted. System 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.
[0063] User computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0064] User computing device 102 includes one or more processors 112 and memory 114. The one or more processors 112 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be one or more processors operatively connected. Memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 114 can store data 116 and instructions 118 executed by processor 112 to cause user computing device 102 to perform operations.
[0065] User computing device 102 may store or include one or more computer applications 119. The computer applications 119 may be configured to perform various operations and provide application outputs as described herein.
[0066] User computing device 102 may store or include artificial intelligence system 120. Artificial intelligence system 120 may perform some or all of the operations described herein. Artificial intelligence system 120 may be separate from and distinct from one or more computer applications 119, but may be able to communicate with one or more computer applications 119.
[0067] User computing device 102 may store or include one or more machine learning models 122. For example, machine learning model 122 may be, or may otherwise include, various machine learning models, such as neural networks (e.g., deep neural networks) or other multi-layered nonlinear models. Neural networks may include recurrent neural networks (e.g., long short-term memory recurrent neural networks), feedforward neural networks, or other forms of neural networks. Reference Figure 2 and 3 Discuss example machine learning model 122.
[0068] In some implementations, one or more machine learning models 122 may be received from server computing system 130 via network 180, stored in user computing device memory 114, and used or otherwise implemented by one or more processors 112. In some implementations, user computing device 102 may implement multiple parallel instances of a single machine learning model 122 (e.g., performing parallel operations for multiple instances across machine learning model 120).
[0069] Additionally or alternatively, the artificial intelligence system 140 may be included in, or otherwise stored and implemented by, the server computing system 130, which communicates with the user computing device 102 according to a client-server relationship. For example, the artificial intelligence system 140 may include a machine learning model 142. For example, the machine learning model 142 may be implemented by the server computing system 140 as part of a web-based service. Thus, one or more models 122 may be stored and implemented at the user computing device 102, and / or one or more models 142 may be stored and implemented at the server computing system 130.
[0070] User computing device 102 may also include one or more user input components 124 for receiving user input. For example, user input component 124 may be a touch-sensitive component (e.g., a touch-sensitive display or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). Touch-sensitive components can be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other devices that allow the user to type communications.
[0071] Server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be one or more processors operatively connected. Memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 134 can store data 136 and instructions 138 executed by processor 132 to cause server computing system 130 to perform operations.
[0072] In some implementations, the server computing system 130 includes one or more server computing devices or is otherwise implemented using one or more server computing devices. When the server computing system 130 includes multiple server computing devices, these devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0073] As described above, the server computing system 130 may store or otherwise include one or more machine learning models 142. For example, model 142 may be, or may otherwise include, various machine learning models, such as neural networks (e.g., deep recurrent neural networks) or other multi-layered nonlinear models. (See reference...) Figure 2 and 3 Discuss example model 142.
[0074] Server computing system 130 can train model 142 via interaction with training computing system 150, which is communicatively coupled to network 180. Training computing system 150 may be separate from server computing system 130, or it may be part of server computing system 130.
[0075] The training computing system 150 includes one or more processors 152 and memory 154. The one or more processors 152 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be one or more processors operatively connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158 executed by the processors 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes one or more server computing devices or is otherwise implemented by one or more server computing devices.
[0076] Training computing system 150 may include a model trainer 160 that trains machine learning model 142 stored in server computing system 130 using various training or learning techniques, such as backpropagation of error. In some implementations, performing backpropagation of error may include performing truncated backpropagation over time. Model trainer 160 may perform various generalization techniques, such as weight decay, dropout, etc., to improve the generalization ability of the trained model.
[0077] In some implementations, if the user has already provided consent, training data 162 can be obtained from the user computing device 102 (e.g., based on communications previously provided by the user of the user computing device 102). Therefore, in such an implementation, the model 122 provided to the user computing device 102 can be trained by the training computing system 150 on user-specific communication data received from the user computing device 102. In some cases, this process may be referred to as a personalized model.
[0078] Model trainer 160 includes computer logic for providing the required functionality. Model trainer 160 can be implemented using hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, model trainer 160 includes a program file stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, model trainer 160 includes one or more sets of computer-executable instructions stored in a tangible computer-readable storage medium, such as RAM, a hard disk, or an optical or magnetic medium.
[0079] Network 180 can be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Typically, communication over network 180 can be carried via any type of wired and / or wireless connection, using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).
[0080] Figure 1A An example computing system that can be used to implement this disclosure is shown. Other computing systems may also be used. For example, in some implementations, user computing device 102 may include model trainer 160 and training dataset 162. In such an implementation, model 122 may be trained and used locally at user computing device 102. In some such implementations, user computing device 102 may implement model trainer 160 to personalize model 122 based on user-specific data.
[0081] Figure 1BA block diagram is depicted of an example computing device 10 that can be used to implement this disclosure. The computing device 10 may be a user computing device or a server computing device.
[0082] The computing device 10 includes multiple applications (e.g., application 1 to application N). Each application contains its own machine learning library and (multiple) machine learning models. For example, each application may include a machine learning model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.
[0083] like Figure 1B As shown, each application can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is application-specific.
[0084] Figure 1C A block diagram depicts an example computing device 50 implemented according to an example embodiment of the present disclosure. The computing device 50 may be a user computing device or a server computing device.
[0085] Computing device 50 includes multiple applications (e.g., application 1 to application N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the multiple models stored therein) using an API (e.g., a common API across all applications).
[0086] The central intelligence layer comprises multiple machine learning models. For example, such as... Figure 1C As shown, a corresponding machine learning model (e.g., a model) can be provided for each application and managed by a central intelligent layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligent layer can provide a single model (e.g., a single model) for all applications. In some implementations, the central intelligent layer is included within the operating system of computing device 50 or implemented by the operating system of computing device 50 in other ways.
[0087] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data resource library of computing device 50. For example... Figure 1CAs shown, the central device data layer can communicate with many other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0088] Example model arrangement
[0089] Figure 2 A block diagram of an example artificial intelligence system 200 according to an example embodiment of the present disclosure is depicted. The artificial intelligence system 200 may include one or more machine learning models 202 trained to receive context data 204 and, as a result of receiving the context data 204, provide model output 206 describing one or more semantic entities referenced by the context data 204.
[0090] The context data 204 discussed herein may include a variety of information, such as information currently displayed in the user interface, information previously displayed in the user interface, information extracted from the user's previous actions (e.g., text written or read by the user, content viewed by the user, etc.), and / or the like. Context data 204 may include user data describing preferences or other information associated with the user and / or contact data describing preferences or other information associated with the user's contacts. Example context data 204 may include messages received by the computing system for the user, the user's previous interactions with one or more of the user's contacts (e.g., text messages mentioning the user's preferences for restaurants or food types), location-related previous interactions (e.g., visiting parks, museums, other attractions, etc.), business-related interactions (e.g., posting reviews of restaurants, reading restaurant menus, making reservations at restaurants, etc.), and / or any other suitable information about the user's preferences or the user. Further examples include audio played or processed by the computing system, audio detected by the computing system, information about the user's location (e.g., the location of the computing system's mobile computing device), and / or calendar data. For example, context data 204 may include external audio detected by the computing system's microphone and / or telephone audio processed during a telephone call. Calendar data can describe future events or plans (e.g., flights, hotel bookings, dinner plans, etc.). Example semantic entities that can be described by the model output 206 can include words or phrases recognized in text and / or audio. Additional examples can include information about the user's location, such as city name, state name, street name, names of nearby attractions, etc.
[0091] In some implementations, model output 206 can more directly describe the suggested action. For example, (multiple) machine learning models 202 can be trained to output text (e.g., including verbs, applications, and semantic entities) describing the suggested action, such as reference data. Figures 4 to 9 As described above, a single machine learning model 202 can be trained to receive context data 204 and output such model output 206.
[0092] In some implementations, multiple machine learning models 202 can be trained (e.g., end-to-end) to produce such a model output 206. For example, a first model of machine learning model 202 can output data describing semantic entities included in context data 204. A second model of machine learning model can receive data describing semantic entities included in context data 204 and output model output 206 describing suggested actions about the semantic entities. Multiple second machine learning models can additionally receive some or all of the context data 204. Those skilled in the art will understand that additional configurations are possible within the scope of this disclosure.
[0093] Figure 3 A block diagram of an example computing system 300 including an artificial intelligence system 301 is depicted. According to an example embodiment of this disclosure, the artificial intelligence system 301 may include one or more machine learning models 302. The machine learning models(s) 302 may be trained to receive context data 304 and, as a result of receiving the context data 304, provide model outputs 306 describing one or more semantic entities referenced by the context data 304.
[0094] The computing system 300 can be configured to interface with one or more computer applications 308 to provide suggested actions that can be performed with the computer applications(s) 308. The computing system 300 can provide the computer applications(s) 308 with data descriptions of the model outputs 306 of the suggested actions, representing descriptions of the machine learning models(s) 302. The computing system 300 can receive one or more application outputs 310 from the computing applications(s) 308 via predefined application programming interfaces. The suggested actions can describe or correspond to at least one of the application outputs 310.
[0095] Figure 4An example suggested action 400 according to aspects of this disclosure is depicted. Suggested action 400 can describe an available action that can be performed by a computer application. In this example, suggested action 400 may include creating a calendar event using a calendar application. More specifically, in this example, suggested action 400 includes the text "Add an appointment with Dr. Sherrigan to the calendar". The computing system can be configured to perform the action in response to a user touch input requesting the action. For example, a user can click button 401, slide a slider, or otherwise interact with suggested action 400 to request the computing system to perform the action.
[0096] In some implementations, the artificial intelligence system can generate (multiple) suggested actions 400 based on a template (e.g., a predefined template). Using templates can reduce the computational resources required to generate such suggested actions 400. Instead of training and utilizing a machine learning model to generate all the text for the suggested actions 400, multiple machine learning models can be used to generate keywords for the suggested actions 400, which are then assembled according to the template to generate the suggested actions 400. For example, the template may include verbs 402, semantic entities 404, and / or corresponding placeholders for computer applications 406. Semantic entities 404 can be described by the model outputs 206 of the multiple machine learning models 202, as referenced above. Figure 2 The templates can be arranged as follows:
[0097] [Verb] + [Semantic Entity] + [Computer Application]
[0098] The artificial intelligence system can select appropriate verbs and / or computer applications to insert into the corresponding placeholders in the template. The verbs and / or computer applications can be selected based on contextual data and / or semantic entities. It should be understood that various template variations can be employed within the scope of this disclosure. Furthermore, the artificial intelligence system can learn user preferences regarding which templates to use and / or whether to use templates at all. Therefore, the systems and methods described herein can employ one or more templates to generate suggested actions.
[0099] In some implementations, the same one or more templates can be used to generate multiple suggested actions, ensuring that these actions have the same overall look and feel to the user. This allows users to quickly evaluate suggested actions because they are presented in a known and predictable format. Therefore, the templates described in this paper can facilitate wider use of suggested actions.
[0100] An explanation 410 can be displayed regarding the suggested action 400. Explanation 410 can describe information about obtaining the context data, including, for example, the time the context data was obtained, the location of the computing device when the context data was obtained, and / or the source of the context data. In this example, explanation 410 states that the context data was "saved at home" at "8:06 AM". Explanation 410 can indicate the source of the context data used to generate the suggested action 400 by displaying an icon. In this example, explanation 410 can include a phone icon 412 to indicate that the context data was collected from audio during a phone call. Explanation 410 can encourage users to be more comfortable or trusting of the operation of the AI system, which can make the AI system more useful to the user.
[0101] In some implementations, in addition to the explanation 410, the computing system may provide the user with a way to view additional information associated with obtaining the context data. For example, the computing system may detect user touch input requesting additional information about the explanation 410. In this example, in response to receiving user touch input for "details" 414, the computing system may display additional explanatory information about obtaining the context data. The additional explanatory information may include the time when the context data was obtained or the location of the computing device when the context data was obtained. The additional explanatory information may include the source of the context data (if it has not already been displayed). The additional information may include information about other times when the context data was obtained in a similar manner (e.g., from the same source, at a similar time, at a similar time, while the computing device was in the same location, etc.).
[0102] Figure 5A Additional example suggested action 500 is depicted according to aspects of this disclosure. This example suggested action 500 describes an available action that can be performed using a shopping app. In this example, suggested action 500 may include purchasing the item "Webster Classic Grill" using a shopping app. More specifically, in this example, suggested action 500 includes the text "Buy Webster Classic Grill on Amazon". This text may be based on the reference above. Figure 4 The description uses a predefined template for generation. As mentioned above, the template may include corresponding placeholders for verb 502, semantic entity 504, and computer application 506. The computing system can be configured to perform the action in response to receiving user touch input requesting the action. For example, a user can click button 501, slide a slider, or otherwise interact with suggested action 500 to request the computing system to perform the action.
[0103] An explanation 510 can be displayed regarding the suggested action 500, describing information about obtaining contextual data, including, for example, the time and location of the computing device when the contextual data was obtained and / or the source of the contextual data. In this example, explanation 510 states that the contextual data was "saved at Home Depot" at "8:38 AM". Explanation 510 can indicate the source of the contextual data used to generate the suggested action 500, for example, by displaying icon 512. In this example, icon 512 can indicate that the contextual data was obtained from an image (e.g., a photo or screenshot). Explanation 510 can encourage users to be more comfortable or trusting of the operation of the artificial intelligence system, which can make the artificial intelligence system more useful to the user. The computing system can be configured to detect user touch input requesting additional information about explanation 510. In this example, in response to receiving user touch input for "details" 514, the computing system can display additional explanatory information about obtaining the contextual data.
[0104] Figure 5B Additional example suggested action 520 according to aspects of this disclosure is described. In this example, suggested action 520 may include shopping for the item "Buy a BKR water bottle on Amazon" using a shopping app. This text may be based on the above references. Figure 4 The description uses a predefined template to generate the action. As described above, the template may include a verb 522, a semantic entity 524, and corresponding placeholders for the computer application 526. The suggested action 520 may include a button 525 for performing the suggested action 520 and an explanation 530. In this example, the explanation 530 indicates the location and time when the context data was saved (e.g., "saved at home · 8:06 AM"). The explanation 530 may indicate the source of the context data used to generate the suggested action 520 by displaying an icon 532. In this example, the icon 532 may indicate that the context data was obtained from audio (e.g., voice memos or external audio). As indicated above, the explanation 530 may encourage users to be more comfortable or trusting of the AI system, making the AI system more useful to the user. The computing system may be configured to detect user touch input requesting additional information about the explanation 530. In this example, in response to receiving user touch input for "details" 534, the computing system may display additional explanatory information about obtaining the context data.
[0105] Figure 5CAdditional example suggested action 540 according to aspects of this disclosure is depicted. This example suggested action 540 describes an available action that can be performed by a music streaming application. In this example, suggested action 540 may include listening to a specific song 544 by a specific artist 546, which may correspond to a stored semantic entity. A user can perform suggested action 540 by clicking button 541. Additionally, in this example, suggested action 540 may include a save button 548 for saving suggested action 540 to a later time and / or a share button 550 for sharing suggested action, for example, via social media, text message, email, etc.
[0106] An explanation 552 may be displayed regarding the suggested action 540, describing information about obtaining context data, including, for example, the time and location of the device and / or the source of the context data when it was obtained. In this example, explanation 552 states that the context data was "saved at 11:06 PM in Linda's Tavern". A portion of explanation 552, such as the location "Linda's Tavern", may include links, for example, to further information about that location (e.g., web search or map application search).
[0107] Explanation 552 can indicate the source of the contextual data used to generate suggested action 540 by displaying icon 554. In this example, icon 554 could indicate that the contextual data was obtained from external music (e.g., detected by the microphone of the computing device). Explanation 552 can encourage the user to be more comfortable or trusting of the operation of the artificial intelligence system, which can make the artificial intelligence system more useful to the user. The computing system can be configured to detect user touch input requesting additional information about explanation 552. In this example, in response to receiving user touch input for "details" 556, the computing system can display additional explanatory information about obtaining the contextual data.
[0108] Figure 6 A computing device 600 is depicted with an example panel 602 illustrating various suggested actions 604, 606, 608 displayed in a lock screen, according to aspects of this disclosure. The lock screen can be displayed when the computing device 600 is locked and authentication is required to access the main menu or perform other operations.
[0109] After collecting context data during the first time interval, multiple suggested actions 604, 606, and 608 can be displayed together during the second time interval. More specifically, at least one additional suggested action 606 or 608 can be displayed in the user interface along with suggested action 604. The additional suggested actions 606 and 608 can be distinct from suggested action 604. The additional suggested actions 606 and 608 can also be generated and stored by the artificial intelligence system based on distinguishable semantic entities (e.g., during the first time interval). For example, the artificial intelligence system can obtain additional context data that is distinct from the context data and input multiple additional model inputs, including the additional context data, into multiple machine learning models. The artificial intelligence system can receive data describing the multiple additional suggested actions as additional outputs of the multiple machine learning models, such as as referenced above. Figure 2 and 3 Therefore, an artificial intelligence system can use (multiple) machine learning models to store multiple semantic entities within a first time interval (e.g., as the user continues their day), and then display multiple suggested actions during a second time interval (e.g., at the end of the day).
[0110] Additional buttons can also be displayed in panel 602 for the user to control or manipulate suggested actions 604, 606, 608 and / or adjust the settings of the artificial intelligence system. As an example, a settings icon 610 can be displayed in panel 602. In response to a user touch action on the settings icon 610, the computing system can display a settings panel, as shown in the following reference. Figure 11 Users can adjust the settings of the artificial intelligence system using the settings panel.
[0111] As another example, a search icon 612 can be displayed in panel 602. In response to a user touch action on the search icon 612, the user can search for suggested actions that are not currently displayed in panel 602.
[0112] As another example, a "View All" button 614 can be displayed in panel 602. In response to a user touch action on the "View All" button 614, the user can view additional suggested actions that are not currently displayed in panel 602.
[0113] Figure 7 A computing device 700 is depicted displaying an example notification panel 702 with suggested actions 704 including notifications 706 and 708. The notification panel 702 can be displayed automatically or in response to user input requesting the display of the notification panel 702.
[0114] Figure 8A computing device 800 according to aspects of this disclosure is depicted in a first state, in which multiple category labels 802, 804, 806, 808, 810, 812, and 814 corresponding to categorized suggested actions are displayed. Multiple suggested actions 816 and 818 may be displayed in a panel 820 having multiple category labels 802, 804, 806, 808, 810, and 812. An artificial intelligence system can categorize suggested actions 816, 818, 822, and 824 with respect to multiple categories corresponding to the multiple category labels 802, 804, 806, 808, 810, 812, and 814. Category labels 802, 804, 806, 808, 810, 812, and 814 may be displayed in a user interface. Category labels 802, 804, 806, 808, 810, 812, and 814 can describe two or more of multiple categories, allowing users to navigate between categories of suggested actions 816, 818, 822, and 824 using category labels 802, 804, 806, 808, 810, 812, and 814 (e.g., in separate panels or pages). For example, a computing system can detect a user touch action related to a category label 814 and display suggested actions corresponding to the selected category label 814, as shown in the following example combination. Figure 9 As stated above.
[0115] Figure 9 This describes one of the many category labels, category label 814, which has been selected. Figure 8 The computing device 800 displays suggested actions 822, 824 corresponding to the selected category label 814, according to aspects of this disclosure. More specifically, in response to detecting a user touch action, the computing system can display suggested actions 822, 824 regarding the classification of the selected category label 814. Therefore, the artificial intelligence system can classify the suggested actions 816, 818, 822, 824 and provide the user with an intuitive way to navigate between the categories of suggested actions 816, 818, 822, 824 for selection.
[0116] In some implementations, the AI system can (e.g., based on user data) rank, sort, categorize, and prioritize suggested actions 816, 818, 822, and 824. User data may include user preferences, calendar data (e.g., the user's plans or schedules), and / or other information about the user or computing device. For example, the AI system can rank suggested actions 816, 818, 822, and 824 and arrange them within the user interface based on the ranking. Therefore, the AI system can prioritize suggested actions 816, 818, 822, and 824 and selectively display a set of the most important and / or relevant suggested actions 816, 818, 822, and 824 to the user in the user interface.
[0117] Figure 10 A computing system 1000 is described in accordance with aspects of this disclosure to display a suggestion action 1002, wherein a semantic entity 1004 has been selected and a search is performed on a panel 1006 that is being covered in response to the selection of the semantic entity.
[0118] Figure 11 A computing system 1100 is depicted displaying a settings panel 1102, where a user can select a default computer application 1104 for a type of suggested action (e.g., a shopping suggestion). The settings panel 1102 allows the user to adjust other settings associated with the artificial intelligence system.
[0119] Example Method
[0120] Figure 12 A flowchart is depicted for an example method 1200 for identifying information of interest, storing that information, and providing suggested actions to a user of a computing system at a later time based on the stored information. Although for illustrative and discussion purposes, Figure 12 The steps are described in a specific order, but the method of this disclosure is not limited to the specific order or arrangement shown. The steps of method 1200 may be omitted, rearranged, combined and / or adapted in a variety of ways without departing from the scope of this disclosure.
[0121] In (1202), the computing system can acquire contextual data during the first time interval. As the computing device performs tasks throughout the day, the artificial intelligence system can capture information of interest during the first time interval. For example, when a user navigates between multiple computing applications and / or switches between different tasks or activities, the artificial intelligence system can identify and store semantic entities.
[0122] In (1204), the computing system can input model inputs, including context data, into one or more machine learning models, as shown in the reference above. Figure 2 and3 As stated above.
[0123] In (1206), the computing system can receive model output describing one or more semantic entities referenced by context data, as the output of one or more machine learning models, such as those mentioned above. Figure 2 and 3 As stated above.
[0124] In (1208), the computing system can store the model output in a tangible, non-transitory computer-readable medium.
[0125] In (1210), the computing system can provide suggested actions for one or more semantic entities described by the model output, to be displayed in the user interface during a second time interval following a first time interval. For example, the computing system can display the suggested actions at a convenient time for the user to view them (e.g., after get off work, after dinner, at a regularly scheduled time interval, etc.). As an example implementation, the first time interval can be defined as the duration of a call associated with a specific business. Contextual data may include payment dates, payment amounts, or other information discussed during the call.
[0126] As another example, suggested actions can be displayed in response to an event (e.g., a second time interval can begin in response to an event). For instance, a computational system could provide suggested actions that include scheduling a calendar event for the user at a time interval (e.g., 7 days) prior to the event. In some implementations, the duration between obtaining contextual data and providing the suggested action can be learned based on user interactions (e.g., previous suggested actions, the application associated with the suggested action to be provided, etc.). As a further example, a promotion of an item or a price reduction could allow a computational system to provide suggested actions about the item based on contextual data including user interactions with the item at an earlier time.
[0127] In some implementations, suggested actions can be provided hours, days, weeks, or even months after the context data is obtained. For example, suggested actions can be provided at least four hours (or eight hours, or longer) after the context data is obtained, making the original event that triggered the acquisition of the context data less fresh in the user's mind. In this way, the suggested actions may be more useful as a "reminder" to the user.
[0128] Additional disclosure
[0129] The technologies discussed in this paper involve servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from these systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and partitions of tasks and functions among or within components. For example, the processes discussed in this paper can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0130] While the subject matter has been described in detail with reference to several specific example embodiments, each example is provided by way of explanation and not as a limitation of this disclosure. Those skilled in the art will readily recognize changes, variations, and equivalents to these embodiments upon understanding the foregoing. Therefore, this disclosure does not exclude modifications, variations, and / or additions to the subject matter, which will be apparent to those skilled in the art. For example, features shown or described as part of one embodiment may be used with another embodiment to produce yet another embodiment. Therefore, this disclosure is intended to cover such changes, variations, and equivalents.
Claims
1. A computing system, the computing system comprising: Artificial intelligence systems with memory-based context storage, The artificial intelligence system includes one or more machine learning models, and The artificial intelligence system is configured to perform operations, including: During the first time interval: The artificial intelligence system receives model input, including contextual data. The artificial intelligence system processes the model input using one or more machine learning models to generate a model output that describes one or more semantic entities referenced by context data; Store the model output in memory; and During the second time interval following the first time interval: Based on the model output stored in memory, suggest actions for one or more semantic entities described by the model output.
2. The computing system according to claim 1, wherein, The context data includes visual or audio information displayed, played, or processed by the computing system.
3. The computing system according to claim 2, wherein, The context data includes screenshots.
4. The computing system according to claim 2, wherein, The one or more semantic entities include one or more words or phrases identified in the displayed visual information.
5. The computing system according to claim 1, wherein, The context data includes information describing the user's surrounding environment during the first time interval.
6. The computing system according to claim 1, wherein, The artificial intelligence system automatically acquires contextual data from one or more sources without requiring specific user actions to initiate the data acquisition.
7. The computing system according to claim 1, wherein, The artificial intelligence system automatically retrieves contextual data from one or more sources in response to a user switching from a first application or task to a second application or task.
8. The computing system according to claim 1, wherein, The artificial intelligence system automatically acquires contextual data based on one or more user-defined instructions or preferences.
9. The computing system according to claim 1, wherein, The artificial intelligence system operates to store multiple model outputs in a memory over time, and the artificial intelligence system operates to rank, sort, or classify the multiple model outputs to manage the amount of data stored in the memory.
10. The computing system according to claim 1, wherein, The artificial intelligence system is configured to automatically remove or overwrite stored model outputs based on one or more rules related to the age of the model output.
11. The computing system according to claim 1, wherein, The suggested action includes one or more of the following: communication action, information retrieval action, reservation action, or information storage action.
12. The computing system according to claim 1, wherein, The artificial intelligence system interfaces with one or more computer applications to perform suggested actions, wherein the suggested actions are transmitted to one or more computer applications via a predefined application programming interface.
13. The computing system according to claim 1, wherein, The artificial intelligence system provides suggested actions based on user activity during a second time interval, for display in context-relevant times.
14. The computing system according to claim 1, wherein, The artificial intelligence system displays a suggested action in response to an event, wherein the artificial intelligence system displays the suggested action at a predetermined time interval prior to the event.
15. The computing system according to claim 14, wherein, The event is an identified change to at least one of the one or more semantic entities.
16. The computing system according to claim 1, wherein, The second time interval is at least one day later than the first time interval.
17. The computing system according to claim 1, wherein, The second time interval is determined based on previous user interactions with the previously suggested action.
18. The computing system according to claim 1, wherein, The second time interval is determined based on the semantic identity identified in the first time interval.
19. A non-transitory computer-readable medium comprising: An artificial intelligence system with memory-based context storage, wherein the artificial intelligence system includes one or more machine learning models, and wherein the artificial intelligence system is configured to perform operations including: During the first time interval: The artificial intelligence system receives model input, including contextual data. The artificial intelligence system processes the model input using one or more machine learning models to generate a model output that describes one or more semantic entities referenced by context data; Store the model output in memory; and During the second time interval following the first time interval: Based on the model output stored in memory, suggest actions for one or more semantic entities described by the model output.
20. A computer-implemented method for providing suggested actions to a user, the computer-implemented method comprising: During the first time interval: A computing system including an artificial intelligence system with a memory-based context storage receives model input including context data, the artificial intelligence system including one or more machine learning models; The artificial intelligence system processes the model input using one or more machine learning models to generate a model output that describes one or more semantic entities referenced by context data; Store the model output in memory; and During the second time interval following the first time interval: Based on the model output stored in memory, suggest actions for one or more semantic entities described by the model output.