Activity agents using episodic memory

A computing system with machine-learned models integrates health coaching agents and data sources to provide personalized, context-aware coaching, addressing inefficiencies and enhancing user experience.

WO2026084703A1PCT designated stage Publication Date: 2026-04-23GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
GOOGLE LLC
Filing Date
2024-10-16
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Current health coaching technologies lack a holistic system-level solution that integrates multiple agents and data sources, fail to manage complex data interactions, and do not support advanced conversational capabilities, leading to inefficiencies and lack of personalization.

Method used

Implementing a computing system with machine-learned models to extract context information from dialogue and biometric data, store it episodically, and provide personalized coaching by coordinating multiple health coaching agents, ensuring consistent and context-aware interactions.

Benefits of technology

Enhances personalization and effectiveness of health coaching by integrating diverse data sources, managing notifications intelligently, and providing contextually appropriate recommendations, thus improving user experience and conserving computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024051580_23042026_PF_FP_ABST
    Figure US2024051580_23042026_PF_FP_ABST
Patent Text Reader

Abstract

A computing system includes one or more memories to store instructions and one or more processors to execute the instructions to perform operations, including: extracting, via a machine-learned model, context information from dialogue information received from a user via an activity agent; receiving first biometric information associated with the user relating to a first biometric activity; storing, as a first episode among a plurality of episodes in the one or more memories, the context information, the first biometric activity, and the first biometric information; receiving a query associated with the context information or the first biometric activity; searching the plurality of episodes stored in the one or more memories according to the query; and when the first episode includes information responsive to the query, providing, for presentation to the user via the activity agent, an output related to the first biometric activity and the context information.
Need to check novelty before this filing date? Find Prior Art

Description

ACTIVITY AGENTS USING EPISODIC MEMORYFIELD

[0001] This disclosure relates generally to machine learning processes and machine-learned devices and systems. More particularly, the disclosure relates to implementing a plurality of machine-learned models to implement and improve activity (e.g., biometric activity) systems. For example, an activity application or system (platform) can use one or more machine- learned models to implement activity7agents across a plurality of applications to provide a personalized health coaching experience to one or more users.BACKGROUND

[0002] A computer can receive input(s). The computer can execute instructions to process the input(s) to generate output(s) using a parameterized model. The computer can obtain feedback on its performance in generating the outputs with the model. The computer can generate feedback by evaluating its performance. The computer can receive feedback from an external source. The computer can update parameters of the model based on the feedback to improve its performance. In this manner, the computer can iteratively “learn” to generate the desired outputs. The resulting model is often referred to as a machine-learned model.

[0003] Current technologies in health coaching are typically limited to specific applications or devices and do not offer a holistic system-level solution that can integrate multiple agents and data sources. These existing solutions often lack the capability to manage complex data interactions and do not support advanced conversational capabilities.

[0004] Further, existing health coaching technologies generally lack the capability to integrate and utilize user data comprehensively. They often rely on simplistic data handling methods that do not support advanced personalization or cross-platform accessibility.Current notification systems in health applications are also limited, typically offering static and non-adaptive messaging functionalities.SUMMARY

[0005] Aspects and advantages of embodiments of the disclosure will be set forth in part in the following description, or can be learned from the description, or can be learned through practice of the embodiments.

[0006] Example aspects of the disclosure provide an example computing system that includes one or more processors and one or more example non-transitory computer-readable mediastoring instructions that are executable by the one or more processors to cause the computing system to perform example operations. In some implementations, the example operations can include extracting, via one or more machine-learned models, context information from dialogue information received from a user via an activity agent which interacts with the user through dialogue operations; receiving first biometric information associated with the user relating to a first biometric activity, the first biometric information being captured by one or more sensors; storing, as a first episode among a plurality of episodes in the one or more memories, the context information, data associated with the first biometric activity, and the first biometric information; receiving a query associated with at least one of the context information or the first biometric activity'; searching the plurality of episodes stored in the one or more memories according to the query; and when the first episode includes information which is responsive to the query, providing, for presentation to the user via the activity agent, one or more outputs related to the first biometric activity and the context information.

[0007] In some implementations, each of the plurality of episodes is associated with a time interval that is less than a predetermined duration of time.

[0008] In some implementations, extracting, via the one or more machine-learned models, the context information from the dialogue information comprises identifying one or more topics in the dialogue information and mapping the one or more topics to at least one of a network model or a graph model.

[0009] In some implementations, searching the plurality of episodes stored in the one or more memories according to the query comprises traversing the at least one of the network model or the graph model to determine one or more topics mapped to the at least one of the network model or the graph model which correspond to at least one topic in the query.

[0010] In some implementations, when the query is associated with the context information, information included in the plurality of episodes stored in the one or more memories is searched for the context information; and when the first episode includes information corresponding to the context information, providing, for presentation to the user via the activity agent, one or more outputs related to the first biometric activity and the context information.

[0011] In some implementations, the context information relates to an emotional state of the user, searching the plurality of episodes includes searching for information corresponding to the emotional state of the user: and when the first episode includes information correspondingto the emotional state of the user, the one or more outputs provided for presentation to the user via the activity agent indicate the first biometric activity is associated with the emotional state of the user.[00121 In some implementations, when the query is associated with the first biometric activity, information included in the plurality of episodes stored in the one or more memories is searched for the data associated with the first biometric activity; and when the first episode includes the data associated with the first biometric activity, providing, for presentation to the user via the activity agent, one or more outputs related to the first biometric activity and the context information.

[0013] In some implementations, emotional state information associated with the user is mapped to the first biometric activity and stored as part of the first episode among the plurality of episodes in the one or more memories, and the operations further comprise: when the first episode includes the data associated with the first biometric activity, retrieving from the first episode the emotional state information associated with the user mapped to the first biometric activity; and the one or more outputs provided for presentation to the user via the activity' agent indicate an emotional state of the user based on the emotional state information associated with the user mapped to the first biometric activity.

[0014] In some implementations, the first biometric information indicates physiological information associated with the user measured while the user performs the first biometric activity, and the physiological information associated with the user is mapped to the data associated with the first biometric activity and stored as part of the first episode among the plurality of episodes in the one or more memories.

[0015] In some implementations, when the query is associated with the physiological information, information included in the plurality of episodes stored in the one or more memories is searched for the physiological information; when the first episode includes the physiological information, retrieving from the first episode the first biometric activity; and the one or more outputs provided for presentation to the user via the activity agent indicate the physiological information is associated with the first biometric activity.

[0016] In some implementations, when the query is associated with the first biometric activity, information included in the plurality of episodes stored in the one or more memories is searched for the first biometric activity; and when the first episode includes the first biometric activity, retrieving from the first episode the physiological information associatedwith the user mapped to the first biometric activity7; and the one or more outputs provided for presentation to the user via the activity’ agent indicate the first biometric activity is associated with the physiological information.

[0017] In some implementations, the operations further include receiving second biometric information associated with the user relating to a second biometric activity, the second biometric information being captured by one or more sensors, and the second biometric information and data associated with the second biometric activity' are mapped to the first biometric activity and stored as part of the first episode among the plurality of episodes in the one or more memories.

[0018] In some implementations, when the query is associated with the first biometric activity and the second biometric activity, information included in the plurality of episodes stored in the one or more memories is searched for data associated with the first biometric activity and the second biometric activity; when the first episode includes the data associated with the first biometric activity’ and the second biometric activity, retrieving, from the first episode, context information associated with the user which is mapped to the first biometric activity7and the second biometric activity; and the one or more outputs provided for presentation to the user via the activity' agent indicate the context information associated with the user which is mapped to the first biometric activity and the second biometric activity.

[0019] Example aspects of the disclosure provide an example computer-implemented method. In some implementations, the example computer-implemented method can include: extracting, via one or more machine-learned models of a computing system comprising one or more processors, context information from dialogue information received from a user via an activity’ agent; receiving, by the computing system, first biometric information associated with the user relating to a first biometric activity’, the first biometric information being captured by one or more sensors; storing, as a first episode among a plurality of episodes in one or more memories of the computing system, the context information, data associated with the first biometric activity, and the first biometric information; receiving, by the computing system, a query' associated with at least one of the context information or the first biometric activity; searching, by the computing system, the plurality’ of episodes stored in the one or more memories according to the query; and when the first episode includes information which is responsive to the query, providing, for presentation to the user via the activity7agent, one or more outputs related to the first biometric activity' and the context information.

[0020] In some implementations, each of the plurality of episodes is associated with a time interval that is less than a predetermined duration of time.

[0021] In some implementations, extracting, via the one or more machine-learned models, the context information from the dialogue information comprises identifying one or more topics in the dialogue information and mapping the one or more topics to at least one of a network model or a graph model.

[0022] In some implementations, the computer-implemented method includes searching the plurality of episodes stored in the one or more memories according to the query comprises traversing the at least one of the network model or the graph model to determine one or more topics mapped to the at least one of the network model or the graph model which correspond to at least one topic in the query.

[0023] In some implementations, when the query is associated with the context information, information included in the plurality of episodes stored in the one or more memories is searched for the context information; and when the first episode includes information corresponding to the context information, providing, for presentation to the user via the activity agent, one or more outputs related to the first biometric activity and the context information.

[0024] In some implementations, the context information relates to an emotional state of the user, searching the plurality of episodes includes searching for information corresponding to the emotional state of the user; and when the first episode includes information corresponding to the emotional state of the user, the one or more outputs provided for presentation to the user via the activity agent indicate the first biometric activity is associated with the emotional state of the user.

[0025] The computer-implemented method may execute any of the operations of the computing system as described herein.

[0026] Example aspects of the disclosure provide one or more example non-transitory computer-readable media storing instructions that are executable by one or more processors to cause a computing system to perform example operations. In some implementations, the example operations can include extracting, via one or more machine-learned models, context information from dialogue information received from a user via an activity agent; receiving first biometric information associated with the user relating to a first biometric activity, the first biometric information being captured by one or more sensors; storing, as a first episodeamong a plurality of episodes in the non-transitory computer readable medium, the context information, data associated with the first biometric activity, and the first biometric information; receiving a query associated with at least one of the context information or the first biometric activity; searching the plurality of episodes stored in the non-transitory computer readable medium according to the query'; and when the first episode includes information which is responsive to the query’, providing, for presentation to the user via the activity agent, one or more outputs related to the first biometric activity and the context information.

[0027] The non-transitory computer-readable medium may store additional instructions to execute other aspects and operations of the computing systems, devices, and computer- implemented methods as described herein.

[0028] Other example aspects of the disclosure are directed to other systems, methods, apparatuses, tangible non-transitory computer-readable media, and devices for performing functions described herein. These and other features, aspects, and advantages of various implementations will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate implementations of the disclosure and, together with the description, help explain the related principles.BRIEF DESCRIPTION OF THE DRAWINGS

[0029] FIG. 1 A is an example system, according to one or more example embodiments of the disclosure;

[0030] FIG. IB is an example block diagram of a computing system, according to one or more example embodiments of the disclosure;

[0031] FIG. 2 illustrates a llow diagram of an example, non-limiting computer-implemented method, according to one or more example embodiments of the disclosure;

[0032] FIG. 3 illustrates an example block diagram of a system including a biometric activity application, according to one or more example embodiments of the disclosure;

[0033] FIGS. 4A-4B are example user interfaces of a biometric activity’ application, according to one or more example embodiments of the disclosure;

[0034] FIGS. 5A-5D are example implementations of a computing system for generating dialogue and recommendations, according to one or more example embodiments of the disclosure;

[0035] FIG. 6 illustrates a flow diagram of an example, non-limiting computer-implemented method, according to one or more example embodiments of the disclosure;

[0036] FIG. 7 illustrates an example block diagram of a system including a biometric activity application, according to one or more example embodiments of the disclosure;

[0037] FIG. 8 is an example user interface of a biometric activity application, according to one or more example embodiments of the disclosure;

[0038] FIG. 9 illustrates a flow diagram of an example, non-limiting computer-implemented method, according to one or more example embodiments of the disclosure;

[0039] FIG. 10 illustrates an example block diagram of a system including a computing platform, according to one or more example embodiments of the disclosure;

[0040] FIG. 11 is a flow chart diagram illustrating an example method for training a machine-learned model according to example implementations of aspects of the disclosure;

[0041] FIG. 12 is a block diagram of an example processing flow for using machine-learned model(s) to process input(s) to generate output(s) according to example implementations of aspects of the disclosure;

[0042] FIG. 13 is a block diagram of an example sequence processing model according to example implementations of aspects of the disclosure;

[0043] FIG. 14 is a block diagram of an example technique for populating an example input sequence for processing by a sequence processing model according to example implementations of aspects of the disclosure;

[0044] FIG. 1 is a block diagram of an example model development platform according to example implementations of aspects of the disclosure;

[0045] FIG. 16 is a block diagram of an example training workflow for training a machine- learned model according to example implementations of aspects of the disclosure;

[0046] FIG. 17 is a block diagram of an inference system for operating one or more machine- learned model(s) to perform inference according to example implementations of aspects of the disclosure;

[0047] FIG. 18 is a block diagram of an example networked computing system according to example implementations of aspects of the disclosure;

[0048] FIG. 19 is a block diagram of an example computing device according to example implementations of aspects of the disclosure; and

[0049] FIG. 20 is a block diagram of an example computing device according to example implementations of aspects of the disclosure.DETAILED DESCRIPTION

[0050] Reference now will be made to embodiments of the disclosure, one or more examples of which are illustrated in the drawings, wherein like reference characters denote like elements. Each example is provided by way of explanation of the disclosure and is not intended to limit the disclosure. In fact, it will be apparent to those skilled in the art that various modifications and variations can be made to disclosure without departing from the scope or spirit of the disclosure. For instance, features illustrated or described as part of one embodiment can be used with another embodiment to yield a still further embodiment. Thus, it is intended that the disclosure covers such modifications and variations as come within the scope of the appended claims and their equivalents.

[0051] Existing computing systems often fail to provide personalized and context-aware coaching due to the lack of a unified data management system. These systems can suffer from fragmentation and inefficiency in handling user data across different health coaching agents.

[0052] Existing technologies, such as generic fitness applications and health trackers, provide basic health monitoring and generic, non-contextual advice or coaching that does not adapt to individual user needs. These systems often operate in isolation without considering other relevant user data, resulting in less effective health coaching. Current health coaching systems often lack personalization and fail to integrate multiple data sources effectively.

[0053] According to examples of the disclosure, an activity agent (health coach) can communicate with users in a consistent and cohesive manner to ensure that content is personalized, reliable, effective, and safe. According to examples of the disclosure, the activity agent can communicate with other activity agents in a uniform and consistent manner to ensure that content is personalized, reliable, effective, and safe.

[0054] According to examples of the disclosure, a computing platform can coordinate between activity agents in a uniform and consistent manner to ensure communicationbetween and from the activity agents is consistent and implements uniform prompts, evaluation, and fine-tuning for tone, style, strategy, and safely.

[0055] According to examples of the disclosure, the computing systems and methods described herein provide a sophisticated notification system that intelligently manages when and how users receive health-related messages. Further, according to examples of the disclosure, the computing systems and methods described herein provide an enhanced level of personalization and integrate information from among a plurality of sources of information.

[0056] Aspects of the disclosure relate to an innovative computing system that leverages machine-learned models (e.g., large language models (LLMs)) to extract, abstract, store, and service user-specific information from combined sources of natural language and sensor data. The computing systems described herein include advanced mechanisms for extracting and labeling user data, such as health topics and user interactions, which are used to personalize the health coaching experience.

[0057] The computing systems described herein are configured to convert conversational coaching interactions (e.g., dialogue operations) into structured health data that can be used to populate logs and track progress. In some implementations, the computing systems include efficient data storage and retrieval systems that convert unstructured or inconsistently structured user data into a structured, consistent format that can be easily accessed across different services and devices (e.g., across different activity agents).

[0058] In some implementations, the computing systems described herein extract user data to tailor health coaching operations when checking in with a user. In some implementations, insession coaching techniques dynamically pull data from relevant episodic memories associated with a user to enhance the interaction.

[0059] In some implementations, expert machine-learned models (e.g., fine-tuned LLMs) can be utilized to validate outputs to ensure the accuracy and relevance of the information provided during coaching interactions.

[0060] The computing systems described herein can manage personalized user notifications by implementing a layer that transforms coaching conversations and user data into rules and controls for notifications, ensuring that messages are timely and contextually appropriate. In some implementations, the one or more machine-learned models (e.g., the LLMs) can summarize and prioritize groups of notifications based on the user's current context, such as a biometric activity of the user or location of the user.

[0061] The computing systems described herein can manage coordination between a plurality of activity agents (e.g.. health coaching agents) that utilize large language models (LLMs) or other machine learning models and health sensor data to deliver personalized coaching messages to users. The computing systems incorporate advanced model types and training histories to enhance the effectiveness of the health coaching. The computing systems described herein can integrate biometric sensor data with other application programming interfaces such as calendars and maps to provide context-aware coaching. The computing systems can utilize embeddings to make details like calendar events retrievable, enhancing the personalization of coaching.

[0062] In some implementations, different activity agents focus on different aspects of a user’s health (e.g., sleep, exercise, nutrition, etc.) to provide comprehensive coaching.

[0063] The computing systems described herein provide structured interactions based on instructions provided in a prompt (e.g., "celebrate success" and "give options"), which guide the flow and content of coaching dialogue (messages).

[0064] The computing systems described herein include mechanisms to manage the flow of interactions over multiple turns (e.g., based on feedback received from the user), ensuring a coherent and contextually appropriate conversation.

[0065] According to examples of the disclosure, the computing systems and methods described herein provide a unified system that can effectively coordinate multiple health coaching agents (activity agents) and data sources. Existing systems often operate in silos, leading to inefficiencies and potential data privacy concerns. The computing systems and methods described herein can integrate various components and ensure smooth and secure operations across different biometric activity applications (e.g., health coaching applications).

[0066] Aspects of the disclosure are directed to a comprehensive system-level framework that facilitates interactions among various applications, databases, and computing devices to support activity agents (e.g., health coaching agents).

[0067] The computing systems described herein are configured to handle and integrate diverse data sources while ensuring user privacy and data security. For example, the computing systems described herein can utilize APIs to enable seamless communication and data exchange between different activity7agents and systems. The computing systems described herein can also implement Retrieval Augmented Generation (RAG) methods for more nuanced and context-aware conversations between users and activity agents. The computing systems described herein can implement a structured system to manageinteractions and data flow among multiple activity agents, ensuring that each activity agent can function optimally without compromising the overall system integrity.

[0068] One or more technical benefits of the disclosure include the implementation of machine-learned models which generate dialogue for communications with a user and provide recommendations relating to a biometric activity7based on biometric information associated with the user and dialogue information received from the user. The computing systems and methods described herein improve the quality and personalization of dialogue between the user and the computing systems by extracting context information from the dialogue information and inferring information about the user (e.g., a user’s motivation level, sentiment, emotional state, etc ). Therefore, a technical effect achieved by the constituted by the computing systems and methods described herein includes improving a user experience based on more concise training evaluations and / or training recommendations. Such improvement in terms of conciseness is achieved, for example, based on avoiding conflicting operations betw een the plurality of activity7agents and / or avoiding conflicting or inconsistent information that is utilized by an activity agent in providing an output (e.g., a recommendation) to a user. Such improvement in terms of conciseness can also be achieved, for example, by taking into account context information extracted from the dialogue information. The computing systems and methods described herein also improve the quality and personalization of recommendations to the user by extracting context information from the dialogue information and inferring information about the user (e.g.. a user’s motivation level, sentiment, emotional state, etc.) and further taking into account biometric information associated w ith the user. The machine-learned models described herein can conserve computing resources including processing power, memory7, network resources (e.g., bandwidth), etc., by interacting with the user in a more personalized manner and by providing accurate and personalized recommendations, reducing the need for additional requests by the user for revised recommendations, and saving time and computing resources by not requiring the user to input additional prompts or edit existing prompts and thus avoiding the need for processing prompts and generating further inferences. The machine-learned models described herein can conserve computing resources including processing power, memory, network resources (e.g., bandwidth), etc., by implementing a computing platform which coordinates and shares information in an efficient and reliable manner, for example, by providing context information associated with a user that is relevant to some activity agents while refraining from providing the context information to other activity agents to which the computing platform determines is not relevant. The machine-learned models describedherein can conserve computing resources including processing power, memory, network resources (e.g.. bandwidth), etc., by a computing system implementing a plurality of machine-learned models in parallel, where a first machine-learned model is configured to interact with the user through dialogue operations and a second machine-learned model is configured to simultaneously perform classification operations by extracting topics, sentiments, emotions, etc. from dialogue information that can be written to memory (e.g., an episodic memory) and is configured to also retrieve information from the memory (e.g., the episodic memory ). The first machine-learned model can be configured to utilize the information retrieved by the second machine-learned model for performing reasoning and dialogue operations. Further, the second machine-learned model can require less processing power than the first machine-learned model (e.g., the first machine-learned model may utilize billions of parameters while the second machine-learned model may utilize millions of parameters, i.e., ten time less parameters). Therefore, the dialogue operations can be performed more quickly, energy (battery power) can be conserved, TPU resources can be conserved, etc. Further, the smaller model (e.g., the second machine-learned model) can be stored or hosted on a device that has less processing power, therefore avoiding the need to transmit certain information to a remote device for performing the operations of the second machine-learned model and enhancing security of the information provided as an input to the second machine-learned model. Further, in some implementations the machine-learned models described herein can be embodied by pre-existing machine-learned models that are capable of processing prompts as described herein to generate the dialogue and recommendations for interacting with the user. For example, enabling the reuse of a preexisting machine-learned model with the new techniques described herein, can save or conserve storage on a computing device and / or time for training because it is not necessary’ to train and store a new model.

[0069] Therefore, aspects of the disclosure provide technical effects, benefits, and / or improvements in computing technology and the technology of recommendation and dialogue generation systems and machine-learned models, via one or more computing devices (e.g., a user computing device, a server computing system, and combinations thereof) which implement machine-learned models, as described herein.

[0070] Referring now to the drawings, FIG. 1 A is an example system according to one or more example embodiments of the disclosure. FIG. 1 A illustrates an example of a system 1100 which includes a computing device 100, an external computing device 200, a servercomputing system 300, and external content 500, which may be in communication with one another over a network 400. For example, the computing device 100 and the external computing device 200 can include any of a personal computer, a smartphone, a tablet computer, a laptop, a global positioning sendee device, a smartwatch, and the like. In some implementations, the external computing device 200 can include a third-part}' computing system that may store information about a user that can be used by the activity application described herein for providing recommendations to a user. The network 400 may include any type of communications network including a wired or wireless network, or a combination thereof. The network 400 may include a local area network (LAN), wireless local area network (WLAN), wide area network (WAN), personal area network (PAN), virtual private network (VPN), or the like. For example, wireless communication between elements of the example embodiments may be performed via a wireless LAN, Wi-Fi, Bluetooth, ZigBee, WiFi direct (WFD), ultra wideband (UWB), infrared data association (IrDA), Bluetooth low energy (BLE), near field communication (NFC), a radio frequency (RF) signal, and the like. For example, w ired communication between elements of the example embodiments may be performed via a pair cable, a coaxial cable, an optical fiber cable, an Ethernet cable, and the like. Communication over the network 400 can use a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).

[0071] As will be explained in more detail below, in some implementations the computing device 100, external computing device 200, and / or server computing system 300 may form part of an application system which can provide a tool for an activity recommendation system (e.g., biometric activity recommendation system) by which, via machine-learned models described herein provide content and recommendations concerning a biometric activity to a user. Further, in some implementations an application system as described herein can coordinate between activity agents provided at the computing device 100, external computing device 200, and / or server computing system 300 to provide, via the machine-learned models described herein, uniform and consistent content and recommendations concerning the biometric activity' to the user. Still further, in some implementations the application system (or an activity agent) as described herein can be configured to store episodic information concerning a biometric activity associated w ith a user which can be accessed by the application system (or the activity agent) provided at the computing device 100, external computing device 200, and / or server computing system 300 to provide, via the machine-learned models described herein, content and recommendations concerning a biometric activity to the user.

[0072] In some example embodiments, the server computing system 300 may obtain data from one or more of a sensor data store 340. a biometric activity data store 350, a content data store 360, and a machine-learned model data store 370, to implement various operations and aspects of the application systems as disclosed herein. The sensor data store 340, biometric activi data store 350, content data store 360, and machine-learned model data store 370 may be integrally provided with the server computing system 300 (e.g., as part of the one or more memory devices 320 of the server computing system 300) or may be separately (e.g., remotely) provided. Further, sensor data store 340, biometric activity data store 350, content data store 360, and machine-learned model data store 370 can be combined as a single data store (database) or may include a plurality of respective data stores. Data stored in one data store (e.g.. the biometric activity data store 350) may overlap with some data stored in another data store (e.g., sensor data store 340). In some implementations, one data store (e.g., the machine-learned model data store 370) may reference data that is stored in another data store (e.g., the sensor data store 340).

[0073] In some implementations, the sensor data store 340 can store information relating to information collected via one or more sensors. For example, the sensor data store 340 can store sensor data related to various biometrics, including biometrics associated with an electrocardiogram (ECG), photoplethysmography (PPG) information, heart rate, heart rate recovery, pulse information, body mass index information, heart rate variability', blood pressure, oxygen saturation, body temperature, sleep quality, physical activities (e.g., number of steps walked, number of miles cycled, number of laps swam. etc.), and the like. For example, the sensor data store 340 can store sensor data related to other information including imagery and / or videos captured by image sensors, lighting information, weather information, noise information, movement or motion information, etc.

[0074] In some implementations, the information stored in the sensor data store 340 can be associated with and / or stored according to a particular user or a plurality of users, according to a particular sensor category, particular biometric activity, sensor context, according to a particular time, location, content type, etc. In some implementations, the information stored in the sensor data store 340 can be associated yvith and / or stored according to a particular environment (e.g., outdoor, indoor, etc.). In some implementations, the information stored in the sensor data store 340 can be associated yvith and / or stored according to a particular entitythat is associated with the collection of the sensor data (e.g., an entity that collected the sensor data, a user that is associated with the sensor data, an entity to which the collected sensor data is transmitted, etc.). For example, sensor data may be stored or retrieved from the sensor data store 340 that is relevant to a dialogue between a user and an activity agent for providing content, generating dialogue for communication with a user, or providing a recommendation to a user.

[0075] In some implementations, the biometric activity data store 350 can store information relating to biometric activities that are performed by a user. The biometric activity data store 350 can store information relating to particular biometric activities (e.g., sleep activities, exercise activities, nutrition activities, etc.) and may include goals associated with the biometric activity and the user. The information relating to the particular biometric activities can include a time engaged in performing the biometric activity, a metric or score relating to the performance of the biometric activity (e.g., a distance travelled, a number of laps swam, miles bicycled, repetitions and / or sets performed, calories eaten, hours slept, etc ). The information relating to the particular biometric activities can also include or be associated with biometric information collected during the biometric activity (e.g., a heart rate measured while engaged in performing the biometric activity, a stress level, a detected sleep state such as REM, Nl, N2, and N3 stages, calories burned, etc.).

[0076] In some implementations, the information stored in the biometric activity data store 350 can be associated with and / or stored according to a particular user or a plurality of users, according to a particular biometric activity category, biometric activity context, according to a particular time, location, etc. In some implementations, the information stored in the biometric activity data store 350 can be associated with and / or stored according to a particular environment (e.g., outdoor, indoor, etc.). In some implementations, the information stored in the biometric activity data store 350 can be associated with and / or stored according to a particular entity that is associated with the performance of the biometric activity (e.g.. an entity that monitored the biometric activity, a user that is associated with the biometric activity, an entity to which the biometric activity data is transmitted, etc.). For example, biometric activity data may be stored or retrieved from the biometric activity data store 350 that is relevant to a dialogue between a user and an activity7agent for providing content, generating dialogue for communication with a user, or providing a recommendation to a user.

[0077] In some implementations, the content data store 360 can store data associated with content. For example, the content can include images, videos, textual descriptions (e.g., static or commonly used phrases, recommendations, dialogue, etc.). In some implementations, the information stored in the content data store 360 can be associated with and / or stored according to a particular user or a plurality of users, according to a particular content category, content genre, content context, time, location, content type, content environment, etc. In some implementations, the information stored in the content data store 360 can be associated with and / or stored according to a particular entity that is associated with the content (e.g., an entity that requests the content to be generated, an entity that is to receive the content, an entity that appears in the content, etc.). For example, machine-learned models described herein can reference or retrieve content from the content data store 360 when generating dialogue, when generating recommendations, etc. The content which is referenced or retrieved from the content data store 360 may be associated with a location or preferences of the user who is to receive the content (e.g., an image of a favorite food of the user may be referenced and provided to the user by the machine-learned models described herein to generate and recommend anutntional plan).

[0078] Machine-learned model data store 370 can store machine-learned models which can be retrieved and implemented by the server computing system 300 for generating distilled or fine-tuned machine-learned models (e.g., distilled or fine-tuned generative machine-learned models) that, in some implementations, can also be provided to the computing device 100. Machine-learned model data store 370 can also store distilled or fine-tuned machine-learned models (e.g.. distilled or fine-tuned generative machine-learned models) which can be retrieved and implemented by the computing device 100. In some implementations, the computing device 100 can retrieve and implement machine-learned models which are large parameter models that have not been fine-tuned or distilled. The machine-learned models (including large parameter models and distilled or fine-tuned models) stored at the machine- learned model data store 370 can include generative machine-learned models respectively associated with different types of applications, types of items, etc., that may be implemented across a variety of domains (e.g., healthcare, gaming, engineering / science, entertainment, travel, retail, etc.). The machine-learned models may include large language models and general, multimodal models (e.g., Gemini). The machine-learned models may include text- to-text large language models, text-to-image large language models, etc. The machine- learned models may include language models which have been trained using reinforcementlearning from human feedback. The machine-learned models may include generative artificial intelligence (Al) models which may implement generative adversarial networks (GANs), transformers, variational autoencoders (VAEs), neural radiance fields (NeRFs), and the like.

[0079] External content 500 can be any form of external content including news articles, webpages, image files, video files, audio files, written descriptions, ratings, game content, social media content, photographs, commercial offers, transportation method, weather conditions, sensor data obtained by various sensors, or other suitable external content. The computing device 100, external computing device 200, and server computing system 300 can access external content 500 over network 400. External content 500 can be searched by computing device 100, external computing device 200, and server computing system 300 according to known searching methods and search results can be ranked according to relevance, popularity, or other suitable attributes, including location-specific filtering or promotion.

[0080] FIG. IB is an example block diagram of a computing system, according to one or more example embodiments of the disclosure. Referring now to FIG. IB, example block diagrams of a system 1200 including a computing device 100, external computing device 200, and server computing system 300 according to one or more example embodiments of the disclosure will now be described. Although computing device 100 is represented in detail in FIG. IB, features of the computing device 100 described herein are also applicable to the external computing device 200.

[0081] The computing device 100 may include one or more processors 110, one or more memory7devices 120, an application system 130, a position determination device 140, an input device 150, a display device 160, an output device 170, a capture device 180, and one or more sensors 190. The server computing system 300 may include one or more processors 310, one or more memory devices 320, and an application system 330.

[0082] For example, the one or more processors 110, 310 can be any suitable processing device that can be included in a computing device 100 or server computing system 300. For example, the one or more processors 110, 310 may include one or more of a processor, processor cores, a controller and an arithmetic logic unit, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an image processor, a microcomputer, a field programmable array, a programmable logic unit, an application-specific integrated circuit (ASIC), a microprocessor, a microcontroller, etc., and combinations thereof, including any other device capable of responding to and executing instructions in a defined manner. The one or more processors 110, 310 can be a single processor or a plurality of processors that are operatively connected, for example in parallel.

[0083] The one or more memory' devices 120, 320 can include one or more non-transitory computer-readable storage mediums, including a Read Only Memory' (ROM), Programmable Read Only Memory (PROM), Erasable Programmable Read Only Memory (EPROM), and flash memory, a USB drive, a volatile memory device including a Random Access Memory (RAM), a hard disk, floppy disks, a Blu-ray disk, or optical media such as CD ROM discs and DVDs, and combinations thereof. However, examples of the one or more memory devices 120, 320 are not limited to the above description, and the one or more memorydevices 120, 320 may be realized by other various devices and structures as would be understood by those skilled in the art.

[0084] For example, the one or more memory devices 120 can also include data 122 and instructions 124 that can be retrieved, manipulated, created, or stored by the one or more processors 110. In some example embodiments, such data can be accessed and used as input to implement biometric activity application 132, and to execute the instructions to perform various operations as described according to examples of the disclosure.

[0085] For example, the one or more memory devices 320 can also include data 322 and instructions 324 that can be retrieved, manipulated, created, or stored by the one or more processors 310. In some example embodiments, such data can be accessed and used as input to implement biometric activity application 332, and to execute the instructions to perform various operations as described according to examples of the disclosure.

[0086] In some example embodiments, the computing device 100 includes an application system 130. For example, the application system 130 may include the biometric activity application 132. The application system 130 can include various other applications including health applications, search applications, gaming applications, document applications, text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, map applications, social media applications, navigation applications, etc.

[0087] According to examples of the disclosure, the biometric activity application 132 maybe executed by the computing device 100 to invoke an activity agent that implements one ormore machine-learned models to interact or communicates with a user (e.g., via dialogue operations) regarding biometric activities that a user performs (e.g.. exercise, sleep, nutrition, etc.). In some implementations, the activity agent can implement the one or more machine- learned models to provide feedback to the user regarding a biometric activity, provide recommendations to the user, update the user with whether certain biometric goals are being met, ask open-ended questions, etc. In some implementations, the activity agent is configured to coordinate with a computing platform (e.g.. server computing system 300) and / or other activity agents provided at external computing devices, to generate the feedback, recommendations, status updates, questions, etc. Therefore, information provided by the activity agent can be consistent and uniform across multiple devices and activity agents provided at other computing devices. In some implementations, the biometric activity application 132 may be part of another application (e.g., a search application, health application, gaming application, etc.) or may be a standalone application. The biometric activity application 132 may be configured to be dynamically interactive according to various user inputs. Example implementations of the biometric activity application 132 are described herein, however the disclosure is not limited to these examples as various modifications may be made to the embodiments described herein.

[0088] In some examples, one or more aspects of the biometric activity’ application 132 may be implemented by the biometric activity application 332 of the server computing system 300 which may be remotely located, to implement the functions and operations of the activity agent described herein, via one or more machine-learned models. In some examples, one or more aspects of the biometric activity application 332 may be implemented by the biometric activity’ application 132 of the computing device 100, to implement the functions and operations of the activity agent described herein, via one or more machine-learned models.

[0089] In some example embodiments, the computing device 100 includes a position determination device 140. Position determination device 140 can determine a current geographic location of the computing device 100 and communicate the geographic location to the server computing system 300 over network 400. The position determination device 140 can be any device or circuitry for analyzing the position of the computing device 100. For example, the position determination device 140 can determine actual or relative position by using a satellite navigation positioning system (e.g. a GPS system, a Galileo positioning system, the GLObal Navigation satellite system (GLONASS), the BeiDou Satellite Navigation and Positioning system), an inertial navigation system, a dead reckoning system,based on an IP address, by using triangulation and / or proximity to cellular towers or WiFi hotspots, and / or other suitable techniques for determining a position of the computing device 100. For example, in some implementations the biometric activity application 132 may be configured to utilize position information determined by the position determination device 140 in connection with generating the recommendations or feedback to a user regarding a biometric activity’ (e.g., a position of a user may be used by the biometric activity application 132 to determine weather information when generating an exercise recommendation, may be used to determine whether a user is travelling when determining whether a user’s diet or sleep may be affected due to the travel, etc.).

[0090] The computing device 100 may include an input device 150 configured to receive an input from a user and may include, for example, one or more of a keyboard (e.g., a physical keyboard, virtual keyboard, etc.), a mouse, a joystick, a button, a switch, an electronic pen or stylus, a gesture recognition sensor (e.g., to recognize gestures of a user including movements of a body part), an input sound device or speech recognition sensor (e.g., a microphone to receive a voice input such as a voice command or a voice query ), a track ball, a remote controller, a portable (e.g., a cellular or smart) phone, a tablet PC, a pedal or footswitch, a virtual-reality device, and so on. The input device 150 may also be embodied by a touch- sensitive display having a touchscreen capability, for example. For example, the input device 150 may be configured to receive an input from a user associated with the input device 150 for executing the biometric activity' application 132, for providing an input prompt to the biometric activity’ application 132. for providing feedback to the biometric activity application 132, for communicating with other users, for accepting or declining suggestions or recommendations provided by the computing device 100 with respect to a biometric activity7, etc.

[0091] The computing device 100 may include a display device 160 which displays information viewable by the user (e.g., a user interface screen). For example, the displaydevice 160 may be anon-touch sensitive display or a touch-sensitive display. The display device 160 may include a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, active matrix organic light emitting diode (AMOLED), flexible display, 3D display, a plasma display panel (PDP), a cathode ray tube (CRT) display, and the like, for example. However, the disclosure is not limited to these example displays and may include other types of displays. The display device 160 can be used by the application system 130 provided at the computing device 100 to displayinformation to a user relating to the content to be generated, to display information to a user relating to the content which has been generated, to display a user interface to a user for providing an input prompt to the biometric activity application 132, for providing information relating to a rationale for the generation of the content, etc. The display device 160 can be configured to provide, for presentation to a user, one or more user interface screens having user interface elements which are selectable by the user for providing information to the biometric activity application 132 (e.g., for providing feedback, for accepting or declining a recommendation, etc.). The display device 160 may be configured to provide feedback or instructions (guidance) to a user regarding a biometric activity7so that the biometric activity application 132 can obtain information for generating a recommendation.

[0092] The computing device 100 may include an output device 170 to provide an output to the user and may include, for example, one or more of an audio device (e.g., one or more speakers), a haptic device to provide haptic feedback to a user (e.g., a vibration device), a light source (e.g., one or more light sources such as LEDs which provide visual feedback to a user), a thermal feedback system, and the like. For example, the output device 170 may provide information relating to the operations for generating and outputting dialogue operations between the user and the computing device 100. The output device 170 may provide an output including content in response to receiving an input prompt, an output confirming receipt of the input prompt, an output relating to generating a recommendation, an output for providing feedback or instructions (guidance) to a user regarding an item, an output relating to achieving a biometric goal, etc.

[0093] The computing device 100 may include a capture device 180 that is capable of capturing media content, according to various examples of the disclosure. For example, the capture device 180 can include an image capturer 182 (e.g., a camera) which is configured to capture images (e.g., photos, video, and the like). For example, the image capturer 182 can include one or more cameras having an imaging sensor (e.g., a complementary metal-oxide- semiconductor (CMOS) or charge-coupled device (CCD)). For example, the capture device 180 can include a sound capturer 184 (e.g., a microphone) which is configured to capture sound or audio (e.g., an audio recording). The media content captured by the capture device 180 may be transmitted to one or more of the server computing system 300, sensor data store 340, biometric activity data store 350, content data store 360, and machine-learned model data store 370, for example, via network 400. For example, in some implementations, content which is captured by the capture device 180 may be provided as an input to one ormore machine-learned models for various tasks associated with a recommendation and / or dialogue generation system (application system 130) and biometric activity application 132, described herein.

[0094] The computing device 100 may include one or more sensors 190. For example, the one or more sensors 190 may include an inertial measurement unit which includes one or more accelerometers and / or one or more gyroscopes. The one or more accelerometers and one or more gyroscopes may be used to capture motion information with respect to the computing device 100. The motion information obtained via the inertial measurement unit may be associated with the user when the computing device 100 is worn or carried by the user. For example, the one or more sensors 190 may include one or more optical sensors (e.g., one or more photoplethysmography (PPG) sensors, one or more electrocardiogram sensors, etc.). The one or more optical sensors may be configured to provide information about a heart rate of the user, heart rate variability (HRV) information, blood oxygen saturation (SpO2) levels, and the like. The one or more sensors 190 may also include other sensors such as a magnetometer, GPS sensor, proximity7sensor, Hall effect sensor, galvanic skin sensors, force sensors, temperature sensors, pressure sensors, noise sensors, and the like. For example, in some implementations, content which is captured by the one or more sensors 190 may be provided as an input to one or more machine-learned models for various tasks associated with the recommendation and / or dialogue generation system (application system 130) and biometric activity7application 132, described herein. For example, weather conditions (e.g., temperature, wind, precipitation, etc.) measured by various weather sensors of the computing device 100 may be referenced by one or more machine-learned models when generating a recommendation for a user relating to a biometric activity.

[0095] In accordance with example embodiments of the disclosure, the server computing system 300 can include one or more processors 310 and one or more memory' devices 320 as described herein. The server computing system 300 may also include an application system 330 which is similar to the application system 130 described herein.

[0096] For example, the application system 330 may include the biometric activity application 332 which performs functions similar to those described herein with respect to biometric activity application 132. In some implementations, one or more machine-learned models (e.g., generative machine-learned models, large language models, etc.) associated with the biometric activity application 332 may be configured to generate feedback, recommendations, open-ended questions, etc., as described according to examples of thedisclosure (e.g., as described with respect to biometric activity application 132). In some implementations, the biometric activity application 332 may be part of another application (e.g., a search application, health application, gaming application, etc.) or may be a standalone application.

[0097] For example, one or more machine-learned models (e.g., generative machine-learned models, large language models, etc.) associated with the application system 330 (e.g., biometric activity application 332) may be configured to perform a first action (e.g., receive and transmit dialogue information and sensor data to the computing device 100), while the computing device 100 (e.g., biometric activity application 132) may be configured to perform a second action (e.g., execute or invoke an activity agent to implement one or more machine- learned models to generate a recommendation according to a plurality of inputs including the dialogue information, the sensor data, etc.). For example, one or more machine-learned models (e.g.. generative machine-learned models, large language models, etc.) associated with the application system 130 (e.g., biometric activity application 132) may be configured to perform a first action (e.g., generate dialogue information for communicating with a user), while the server computing system 300 (e.g., biometric activity application 332) may be configured to perform a second action (e.g., execute or invoke an activity agent to implement one or more machine-learned models to generate a recommendation according to a plurality of inputs including the dialogue information, the sensor data, etc.).

[0098] Examples of the disclosure are directed to computer implemented methods for recommendation, content, and / or dialogue generation systems including implementing one or more machine-learned models to generate recommendation and content relating to a biometric activity associated with a user and in association with an activity agent that interacts with the user.

[0099] The flow diagram of FIG. 2 illustrates a method 21 0 for generating an output based on dialogue information obtained via dialogue operations and sensor data obtained via one or more sensors, by implementing one or more machine-learned models. Although shown in a particular sequence or order, unless otherwise specified, the order of the processes can be modified. Thus, the illustrated embodiments should be understood only as examples, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel. Additionally, one or more processes can be omitted in various embodiments. Thus, not all processes are required in every embodiment. Other process flows are possible.

[0100] The operations of FIG. 2 will be explained with reference to FIG. 3. FIG. 3 illustrates an example block diagram or architecture of a computing system (recommendation, content, and / or dialogue generation system) 3100 including a biometric activity application 3110, according to one or more example embodiments of the disclosure.

[0101] Referring to FIG. 2, at operation 21 10 the method 2100 includes a computing system executing an activity agent to receive dialogue information from a user via dialogue operations. As described herein, the computing system may be embodied as computing device 100, server computing system 300. or combinations thereof. For example, in the computing system 3100 of FIG. 3, in some implementations a biometric activity application 3110 includes an activity' agent 3120 which may be configured to interact with the user and to receive dialogue information from the user via dialogue operations. The activity agent 3120 may also be referred to as a virtual assistant, digital assistant, Al assistant, chat agent, etc. The activity agent 3120 may be associated with a particular persona, for example a health coach, motivational trainer, etc. In some implementations, the activity agent 3120 may be implemented as an application that is designed to engage in conversations with a user (e.g., a human user) through text and / or voice. In various embodiments of the disclosure, the activity agent 3120 may be configured to engage in conversations with the user via dialogue operations with respect to one or more biometric activities and provide content or feedback that is personalized, reliable, effective, and safe. In various embodiments of the disclosure, the activity agent 3120 may be configured as an agent for a particular ty pe of activity, such as a sleep agent, a nutrition agent, and / or an exercise agent. There may also be separate activityagents for each type of exercise (e.g., a weightlifting agent, a bicycling agent, a running agent, etc.). As illustrated in FIG. 3, a user input 3102 may be received by the biometric activity application 3110 (and particularly by the activity' agent 3120), for example, via the input device 150. The user input 3102 may be in the form of a text input, voice input, the selection of a user interface element, etc., and be part of a dialogue exchange with the activity agent 3120 with respect to a biometric activity’ associated with the user. In some implementations, the activity' agent 3120 includes an engagement engine 3122 which may be configured to implement one or more machine-learned models that can interact with the user using a particular persona, for example, having a particular style and / or behavior, following certain psychological rules and utilizing reasoning capabilities to exchange information with the user. The engagement engine 3122 may further be configured to connect to the context extractor 3124 and the recommendation engine 3126 of the activity agent 3120.

[0102] As an example implementation, FIG. 4A illustrates a first user interface 4100 which displays various dialogue operations between a user and the computing system (e.g., activity agent 3120 such as a sleep agent), by which a user can indicate various information concerning the quality of the user’s sleep. At a first dialogue operation 4110, the activity agent 3120 provides a summary of the user’s sleep (e.g., 6 hours and 50 minutes of sleep) which may be collected via one or more biometric sensors (e.g., through a wearable computing device, cameras, etc.). At a second dialogue operation 4120, the user can provide an input via selection of a user interface element (e.g., a slider bar) indicating a subjective opinion regarding the quality' of sleep of the user (e.g., satisfied). At a third dialogue operation 4130, the user can provide another input via selection of another user interface element indicating whether the user had trouble falling asleep (e.g.. a lot of trouble). At a fourth dialogue operation 4140, the user can provide a further input via selection of a further user interface element indicating whether the user had trouble staying asleep (e.g., some trouble). Further dialogue operations to allow the user to provide further context regarding the user’s sleep can be input via selection of the user interface element 4150 (e.g., "‘Provide more details”).

[0103] As another example implementation, FIG. 4B illustrates a second user interface 4200 which displays various dialogue operations between a user and the computing system (e.g., activity' agent 3120 such as a sleep agent), by which the user can engage in further conversation and provide additional information concerning the quality of the user’s sleep. At a first dialogue operation 4210. the activity agent 3120 provides an open-ended question asking the user if anything affected the user’s sleep the night prior. For example, the content of the first dialogue operation 4210 may be generated via the one or more machine-learned models 3130 in response to the content of a prior dialogue operation (e.g., the selection of a particular user interface element such as the user interface element indicating the user had trouble falling asleep). At a second dialogue operation 4220, the user can provide an input responsive to the query from the computing system at first dialogue operation 4210, indicating further details explaining why the user had trouble falling asleep.

[0104] At operation 2120 the method 2100 includes the computing system implementing one or more machine-learned models to extract context information from the dialogue information received by the activity agent 3120 (e.g., via the engagement engine 3122). For example, in the computing system 3100 of FIG. 3, the context extractor 3124 may be configured to extract information from the dialogue information (e.g., dialogue from the user)to determine a context associated with the dialogue. For example, the context extractor 3124 may be configured to determine or identify one or more entities from the context information including names of people, places, things, etc., objects, time information, to determine or identify a topic of the dialogue, and / or to determine or identify a goal or purpose (e.g., intent) of the dialogue. In some implementations, other information including emotional or sentiment information associated with the user can be determined from the dialogue. Referring to the example user interfaces of FIGS. 4A-4B. the context extractor 3124 may be configured to determine, based on the dialogue information received from the user via the second dialogue operation 4220, context information including terms or phrases which represent one or more topics or themes, one or more entities, one or more intents, one or more events, one or more sentiments or emotions, etc. In some implementations, prior dialogue operations can also be utilized to further determine or identify the context information. In some implementations, the context extractor 3124 may be configured to employ natural language extraction to extract structured information from unstructured natural language text. As shown in FIG. 4B, the context extractor 3124 may be configured to implement natural language processing (extraction) techniques with respect to the dialogue information (e.g.. from prior dialogue operations) to determine topics related to poor sleep quality which include “noise,” “racing thoughts,” and “using technology.”

[0105] At operation 2130 the method 2100 includes the computing system receiving biometric information associated with the user with respect to a first biometric activity. For example, in the computing system 3100 of FIG. 3, the biometric activity application 3110 may be configured to receive biometric information 3104 associated with the user and a first biometric activity which may be part of biometric activity information 3106. In some implementations, the biometric activity application 3110 may also receive external information 3108 which may include sensor data that is not necessarily biometric information associated with the user, but may include other information (e.g.. related to the dialogue and / or biometric activity) such as environmental information (e.g., temperature information, weather information, noise information, calendar information, location information, etc.).The external information 3108 may also external information that is provided via one or more external databases that can be used as part of a retrieval-augmented generation (RAG) framework by the biometric activity application 3110.

[0106] The biometric information 3104 may be captured by one or more sensors of the computing system itself and / or may be captured by one or more sensors of an externalcomputing device. Similarly, other sensor data may be captured by one or more sensors of the computing system itself and / or may be captured by one or more sensors of an external computing device. The biometric information 3104 can include information relating to biometrics associated with an ECG, PPG, heart rate, heart rate recovery, pulse information, BMI, heart rate variability, blood pressure, oxygen saturation, body temperature, sleep quality, physical activities (e.g., number of steps walked, number of miles cycled, number of laps swam. etc.), and the like. Further, the computing system (biometric activity application 3110) may be configured to generate or display information associated with a measured biometric, including an electrocardiogram, a photoplethysmogram, heart rate, heart rate recovery, blood pressure, oxygen saturation, respiration rate, body temperature, physical activity, a sleep metric, electrical conductance, and the like. In the example of FIG. 4A, the first dialogue operation 4110 indicates the user slept 6 hours and 50 minutes, which may be a metric that was measured via one or more sensors of the computing system or of an external computing device (e.g., a camera, a wearable computing device, etc.).

[0107] At operation 2140 the method 2100 includes the computing system implementing the one or more machine-learned models to determine, based on the biometric information and the context information, one or more outputs related to the first biometric acti vity and the user. For example, in the computing system 3100 of FIG. 3, the one or more machine- learned models 3130 of the biometric activity application 3110 may be configured to generate an output 3150 based on the context information obtained at operation 2120 and the biometric information obtained at operation 2130. In some implementations, the output may include dialogue for a dialogue operation between the computing system and the user. In some implementations, the dialogue can include recommendations regarding an action for the user to take which are generated based on the dialogue information and the biometric information 3104. In some implementations, the activity agent 3120 may have been trained to have the persona of a health coach, motivational trainer, etc., and the dialogue generated via the one or more machine-learned models 3130 may characteristically be empathetic, provide encouragement to the user, celebrate achievements with respect to goals met or progress made, provide feedback and suggestions for improving performance and achieving goals, ask open-ended questions, take into account the mood and sentiment of the user, and forego making critical and / or judgmental statements. In some implementations, the one or more machine-learned models 3130 may be configured to utilize a prompt based on constitutional rules (e.g., rules defining a style, tone, psychological guidelines, safety guidelines, etc.) andsurface specific rules (e.g., rules defining criteria for celebrating success, for activating goals, implementing decision trees, interpreting an activity-, etc.). Style rules may define guidelines associated with the length of the text, use of the data, how to break up text, ensure readability, provide a diversity of messages and personalization. Tone rules may help users stay motivated and ensure the dialogue communications are supportive and friendly.Psychological rules can provide that the dialogue communications are non-judgmental, empower the user, leverage SMART goals, and reference health benefits. Safety rules can ensure the one or more machine-learned models 3130 avoid providing medical advice and / or information which could be unsafe for the user.

[0108] In some implementations, the one or more machine-learned models 3130 may be configured to provide an output which is limited to be equal to or less than a predetermined amount (e.g.. less than or equal to 50 words, between 50 words and 80 words, etc.). In some implementations, the activity agent 3120 may include the recommendation engine 3126 which implements the one or more machine-learned models 3130 to generate the recommendations.

[0109] In some implementations, the biometric activity application 3110 may further include one or more expert machine-learned models 3140 which have been trained based on feedback provided by subject matter experts with respect to a particular activity (e.g., a fitness trainer who specializes in training people in a particular exercises (e.g., cycling, weightlifting. running, etc ), a sleep expert, a nutrition expert, etc. The one or more expert machine-learned models 3140 can include fine-tuned models which are trained to verify or validate an output of the one or more machine-learned models 3130, and / or the one or more expert machine- learned models 3140 can be selectively implemented to generate a part of the output 3150 that is associated with a topic or subject matter area that the one or more expert machine- learned models 3140 have been trained on. In some implementations, the one or more machine-learned models 3130 may be trained to generate the dialogue in a particular manner (e.g., with a particular tone, style, persona) while the one or more expert machine-learned models 3140 may be configured to generate the recommendation with respect to a particular suggestion pertaining to the biometric activity7(e.g., generating a recommendation to engage in a mindfulness activity before bed to help a user go to sleep faster).

[0110] For example, the one or more machine-learned models 3130 may be configured to implement a heuristic method or algorithmic method to generate a recommendation. In the example of FIG. 4B, the one or more machine-learned models 3130 may determine arecommendation 4260 to the user to engage in or view mindfulness content to improve sleep quality when the sleep quality has been determined to be poor (e.g., based on the received biometric information) and based on the context information extracted from the dialogue information (e.g., one of the topics of the user’s dialogue during a conversation about the user’s sleep included having racing thoughts). Methods other than that shown in FIG. 4B may be implemented by the biometric activity application 3110 (e.g., the one or more machined earned models 3130 and / or one or more expert machine-learned models 3140) to generate a prediction output regarding the recommendation and / or dialogue. For example, the one or more machine-learned models 3130 and / or one or more expert machine-learned models 3140 may be configured to predict a recommendation of mindfulness content based on various inputs such as the user having a poor sleep quality and context information indicating the user having racing thoughts. Further, the one or more machine-learned models 3130 and / or one or more expert machine-learned models 3140 may be configured to receive, as inputs, information beyond the biometric information and context information. For example, the one or more machine-learned models 3130 and / or one or more expert machine- learned models 3140 may further be configured to receive, as inputs, external information 3108 including sensor information other than biometric information, geographic information, calendar or event information, etc. In some implementations, the external information 3108 can include information that is provided by third-parties and / or other activity agents. For example, a nutrition activity agent may provide information regarding the user’s food intake before the user went to sleep to the activity agent 3120 (which may be a sleep activity agent), one or more machine-learned models 3130, and / or one or more expert machine-learned models 3140. The activity agent 3120, one or more machine-learned models 3130, and / or one or more expert machine-learned models 3140 may take into account the affect that the user’s food intake may have had on the sleep quality of the user when determining dialogue content and / or a recommendation to provide to the user.[01 1 1] For example, the one or more machine-learned models 3130 and / or one or more expert machine-learned models 3140 may further be configured to implement a retrieval- augmented generation (RAG) framework to retrieve relevant information from an external data source (e.g.. an external database) using a query generated by the one or more machine- learned models 3130 and / or one or more expert machine-learned models 3140, and the retrieved relevant information can be provided as an input to the one or more machine- learned models 3130 and / or one or more expert machine-learned models 3140 for generatingthe dialogue content and / or recommendation. The information stored in the external data source may be stored as a vector in a vector database or the biometric activity application 3110 may be configured to convert the retrieved relevant information to a vector format, to allow for fast and accurate retrieval and searching based on semantic similarity.

[0112] In some implementations, the biometric activity application 3110 (e.g., the activity agent 3120) may be configured to convert a format of the user input 3102, biometric information 3104, biometric activity information 3106, and external information 3108 (including information provided by other activity agents) into a structured data format, such that the different information (which may be received from different sources) may be in a common and consistent or uniform format for processing by the one or more machine-learned models 3130 and / or the one or more expert machine-learned models 3140. In some implementations, the biometric activity application 3110 (e.g., the activity agent 3120) may be configured to implement a plurality of machine-learned models in parallel (simultaneously) when interacting with a user, where a first machine-learned model may be configured to interact with the user through dialogue operations and a second machine- learned model may be configured to simultaneously perform classification operations based on dialogue information between the user and the activity agent 3120, by extracting topics, sentiments, emotions, etc. from the dialogue information and which can be written to a memory (e.g., an episodic memory' 7170 as described with respect to FIG. 7) and is configured to also retrieve information from the memory' (e g., the episodic memory 7170 as described with respect to FIG. 7). For example, the second machine-learned model can implement one or more classifiers to classify information from the context information. For example, the one or more classifiers can include at least one of a topic classifier to classify one or more topics included in the context information, an emotional classifier to classify one or more emotional states associated with the user indicated by the context information, or a decision classifier to classify one or more decisions indicated by the context information. The first machine-learned model can be configured to utilize the information retrieved by the second machine-learned model for performing communications and dialogue operations in generating content (e.g., dialogue operations, recommendations, etc.) to be provided to the user. Further, the second machine-learned model can require less processing power than the first machine-learned model (e.g., the first machine-learned model may utilize billions of parameters while the second machine-learned model may' utilize millions of parameters, for example, about ten times less parameters). Therefore, the dialogue operations can beperformed or generated more quickly, energy (battery power) can be conserved, TPU resources can be conserved, etc. Further, the smaller model (e.g.. the second machine- learned model) can be stored or hosted on a device (e.g., computing device 100) that has less processing power, therefore avoiding the need to transmit certain information to a remote device (e.g., server computing system 300) for performing the operations of the second machine-learned model and thereby enhancing security of the information provided as an input to the second machine-learned model. As an example, a user may indicate that they “slept horribly because the cat kept on crawling on their face” and the second machine- learned model can classify the dialogue information (e.g., topics such as “trouble with sleep” “cat,” etc.) and retrieve relevant information including advice or recommendation information that is connected to the classified information (e.g.. from the episodic memory 9170) and / or to make decisions (e.g., via a decision tree). The first machine-learned model may perform dialogue operations while the second machine-learned model performs classification and retrieval operations, and then the first machine-learned model can subsequently utilize the retrieved information for generating further dialogue operations and / or content (e.g., recommendations). Accordingly, less computation power may be expended by using a lighter model and disaggregating the tasks using different models, as well as increasing performance speed.

[0113] At operation 2150 the method 2100 includes the computing system providing, for presentation to the user, the one or more outputs related to the first biometric activity and the user. For example, in the computing system 3100 of FIG. 3, the biometric activity application 3110 (e.g.. the activity agent 3120) may provide, for presentation to the user, the one or more outputs (e.g., via the display device 160 and / or output device 170). In some implementations, the output 3150 may be provided in a dialogue exchange as illustrated in FIG. 4B, such as in third dialogue operation 4230 where the computing system provides a dialogue empathizing with the user (“That can be frustrating”) regarding the sleep quality of the user’s sleep (the first biometric activity) while also providing a recommendation to the user (“Often you can do xxx”) for improving the qualify of their sleep (e.g., engaging in mindfulness content). In some implementations, the biometric activity application 3110 (e.g., the activity agent 3120) may provide the user with the ability to provide feedback or to provide additional information for the recommendation to possibly be revised. For example, in FIG. 4B the user can indicate agreement or disagreement with a recommendation via aselectable user interface element 4240 and / or provide further information via a user interface element 4270 (e.g., via a text box and / or a voice input).

[0114] Various prompts may be provided internally to the one or more machine-learned models 3130 according to one or more examples of the disclosure, in generating the one or more outputs (e g., dialogue information, recommendations, etc.). For example, the internal prompt can be provided by the activity agent 3120 or by a computing platform (e.g., computing platform 1020) that manages a plurality of activity agents. In some implementations, the internal prompt can be provided to the one or more machine-learned models 3130 for generating one or more outputs which summarize the status of where a user is relative to their goals. For example, the internal prompt can instruct the one or more machine-learned models to take on the persona of an expert fitness coach and to provide motivating and / or encouraging statements to the user relating to the biometric goals of the user. For example, the internal prompt can instruct the one or more machine-learned models 3130 to determine (e.g., infer or predict) the user’s motivation level and to generate different dialogue communications and / or recommendations based on the user’s motivation level. For example, for an unmotivated user dialogue operations may be generated to encourage the user based on the importance of exercise and the recommendation may adjust the biometric goals to be more achievable. If the user is motivated, the dialogue operations may be generated to confirm or acknowledge the user’s achievements and the recommendation may adjust the biometric goals to be more challenging but attainable. The internal prompt can also instruct the one or more machine-learned models 3130 to limit the output to be less than a predetermined number of words (e.g.. less than 100 words) or to be of a particular length within a certain range (e.g., between 50 and 80 words). The internal prompt can also instruct the one or more machine-learned models 3130 to personalize the dialogue exchange by referring to personal information associated with the user (e.g., the user’s name, biometric achievements of the user, a time of day, location, etc.). The internal prompt can also instruct the one or more machine-learned models 3130 to find areas for improvement, celebrate successes, provide insights associated with the user based on a user’s current fitness level (e.g., compared to a prior fitness level), provide recommendations based on the user's biometric achievements thus far and a time remaining to meet the biometric goals, and to avoid making judgmental or critical statements. Such an internal prompt can be associated with exercise, nutrition, sleep, and / or other biometric activities.

[0115] In some implementations, an internal prompt can be provided to the one or more machine-learned models 3130 to utilize a structured format that can be applied by the one or more machine-learned models 3130 for generating motivating statements, according to examples of the disclosure. For example, such an internal prompt can be provided by the activity agent 3120 or by a computing platform (e.g., computing platform 1020) that manages a plurality of activity agents. The structured format can provide a uniform and consistent input structure that can be applied by various activity agents and / or biometric activity applications for describing the status of the user’s progress toward completing one or more biometric goals. For example, the input structure can include the user’s name, age, gender, the number of days to achieve the biometric goal and the current number of days spent attempting to achieve the biometric goal, and / or a summary of each biometric activity with respect to the biometric goal (e.g., number of minutes achieved for high intensity cardio out of the number of minutes associated with the biometric goal for the high intensity cardio, and a difference between the achievement and the biometric goal). For example, a JSON format may be utilized with predetermined parameters (e g., keys).

[0116] In some implementations, an internal prompt can be provided to the one or more machine-learned models 3130 for generating one or more outputs which summarize the user’s accomplishments with respect to a biometric goal. For example, such an internal prompt can be provided by the activity agent 3120 or by a computing platform (e.g., computing platform 1020) that manages a plurality of activity agents. For example, the internal prompt can instruct the one or more machine-learned models 3130 to take on the persona of an expert fitness coach and to review the user’s biometric achievements with respect to biometric goals for one or more biometric activities. For example, the internal prompt can instruct the one or more machine-learned models 3130 to determine whether the user met or exceeded their biometric goals, and to generate different dialogue communications and / or recommendations based on whether the user met or exceeded their biometric goals. For example, if the user did not meet their biometric goal, dialogue operations may be generated to encourage the user based on the health benefits that could be achieved by completing the biometric goal, and the recommendation may adjust the biometric goals to be more achievable. If the user has achieved their biometric goals, dialogue operations may be generated to confirm or acknowledge the user’s achievements and to highlight the health benefits that were obtained by achieving the biometric goals, and the recommendation may adjust the biometric goals so that the user can advance beyond theirgoals. The internal prompt can also instruct the one or more machine-learned models 3130 to limit the output to be less than a predetermined number of words (e.g., less than 100 words) or to be of a particular length within a certain range (e.g., between 50 and 80 words), to confine the goal period to a certain timeframe (e.g., 7 days), etc. The internal prompt can also instruct the one or more machine-learned models 3130 to personalize the dialogue exchange by referring to personal information associated with the user (e.g., the user’s name, biometric achievements of the user, a time of day. location, etc.). The internal prompt can also instruct the one or more machine-learned models 3130 to find areas for improvement, celebrate successes, provide insights associated with the user based on a user’s current fitness level (e.g., compared to a prior fitness level), provide recommendations based on the user’s biometric achievements thus far and a time remaining to meet the biometric goals, and to avoid making judgmental or critical statements. Such an internal prompt can be associated with exercise, nutrition, sleep, and / or other biometric activities.

[0117] In some implementations, an internal prompt can be provided to the one or more machine-learned models 3130 to utilize another structured format that can be applied by the one or more machine-learned models 3130 for generating motivating statements, according to examples of the disclosure. For example, such an internal prompt can be provided by the activity agent 3120 or by a computing platform (e.g., computing platform 1020) that manages a plurality of activity agents. The structured format can provide a uniform and consistent input structure that can be applied by various activity agents and / or biometric activity' applications for analyzing the user’s biometric achievements and / or the status of the user’s progress toward completing one or more biometric goals. For example, the input structure can include the user’s name, age, gender, the prior week’s achievements, and a summan of each biometric activity with respect to the biometric goal (e.g., number of minutes achieved for high intensity cardio out of the number of minutes associated with the biometric goal for the high intensity cardio, and a difference between the achievement and the biometric goal). For example, a JSON format may be utilized with predetermined parameters (e.g., keys).

[0118] When a user interacts with the activity agent 3120, the biometric activity application 3110 may be configured to ensure that information (data) transmitted to or received from the user is secure, so as to prevent information from being leaked or misused. In some implementations, the biometric activity application 3110 may be configured to implement end-to-end encryption for securing communications during data transfer between the user and the biometric activity application 3110, to ensure that the information is accessible only to itsintended recipient(s). Other security measures may additionally, or alternatively, be implemented by the biometric activity application 3110 to secure sensitive information including storing sensitive information with non-sensitive placeholders (e.g., tokens), anonymizing information by removing personally identifiable information from the data before transmitting the information, masking sensitive information, utilizing multi-factor authentication methods, etc.

[0119] In some implementations, the biometric activity application 3110 may be configured to store information associated with the user (e.g.. dialogue information, context information, biometric information, etc.) in an encrypted manner such that stored data is converted into an unreadable format that can only be decrypted by an entity or computing system having the proper encryption or cryptographic key. Data can be encrypted at various levels for multiple layers of security (e.g.. at the database level including at a row, column, field, level, etc., at an application level, at a file level, at a disk or storage system level, etc.).

[0120] Further, the biometric activity application 3110 may be configured to provide a user wi th controls allowing the user to make an election as to both if and when systems, programs, or features described herein may enable collection of user information (e.g., information about a user’s social network, social actions, biometric activities, calendar information, dialogue communications, profession, a user’s preferences, a user’s current location, etc.), and if the user is sent content or communications from a server computing system or computing platform. In addition, certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user’s identity may be treated so that no personally identifiable information can be determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined. Thus, the biometric activity application 3110 may be configured to provide a user with controls over what information is collected about the user, how that information is used, and what information is provided to the user. Further, the biometric activity application 3110 may be configured to provide a user with controls over what information is stored and the ability to delete information about the user which has been stored.

[0121] FIG. 5A illustrates an overview process flow 5100 of the biometric activity application including an activity agent, according to examples of the disclosure. Referring to FIG. 5 A, the activity agent 5110 can receive, as an input, first state information 5130 whichrepresents a current situation or condition of the environment 5120 that the activity agent 5110 is interacting with. The first state information 5130 is essentially a snapshot of the environment at a particular time (e.g., time t), which the activity agent 51 10 perceives to make decision or take action 5150. The first state information 5130 can include various features that describe relevant aspects of the environment and can include geographic or location information, sensor data, biometric information, biometric activity information, user information, dialogue information, etc. The activity agent 5110 can also receive, as an input, first reward information 5140 which can include a value that the activity agent 5110 receives after taking the action 5150 in a particular state. The first reward information 5140 can provide or indicate feedback on how successful the action was in relation to achieving the activity agent’s 5110 goal. For example, if the user does not implement or carry out the activity agent’s recommendation or the user is confused or non-responsive to a dialogue operation sent by the activity' agent 5110, the reward value may be low. If the user does implement or carry out the activity agent’s recommendation or the user positively responds to the dialogue operation sent by the activity agent 5110, the reward value may be high. When the reward value is low, the activity agent 5110 may be configured to adjust certain parameters or values, adjust policies, modify actions, etc., in an effort to maximize the cumulative reward over time. For example, in a next state at a subsequent time (e.g., t+1), second state information 5160 may be provided to the activity agent 5110 together with second reward information 5170 which can provide or indicate feedback on how successful the action (output based on the first state information 5130 and first reward information 5140) was in relation to achieving the activity7agent’s 5110 goal.

[0122] FIG. 5B illustrates another overview process flow 5200 of the biometric activity application including an activity' agent, according to examples of the disclosure. Referring to FIG. 5B, at operation 5210 the biometric activity application (e.g., the biometric activity application 3110 including the activity agent 3120) can receive as an input, at operation 5210, sensor information output by one or more sensors 5212, and can analyze, at operation 5214, the user data associated with the sensor information. For example, the user data can include biometric information such as a blood pressure that is obtained from PPG sensors, sleep information indicating a duration of sleep, a number of steps walked, etc.

[0123] At operation 5220, the biometric activity application (e.g., the biometric activity application 3110 including the activity agent 3120) can receive the user data and generate, at operation 5222, an insight relating to the user data. Previous methods generate an insight andretrieve a recommendation that is prestored beforehand based on the sensor data, according to a heuristic method. For example, if the result of analyzing the sensor data is that the user slept four hours, the insight may be that the user did not sleep long enough and the recommendation may be to recommend to the user to sleep in a quiet and dark environment. Different from the previous method, according to examples of the disclosure the insight and recommendation is generated by one or more machine-learned models (e.g., one or more machine-learned models 3130) according to the sensor information (including biometric information) as well as according to feedback provided by the user and / or dialogue information from the user which provides subjective information that is unique or personal to the user. For example, at operation 5222 the one or more machine-learned models may be configured to generate an insight relating to the biometric information and / or to a biometric activity and at operation 5224 may be configured to provide a suggestion or recommendation. At operation 5230 the user can receive the suggestion or recommendation. At operation 5240 the user can take action or not take action (e.g., provide a response via dialogue operations, adjust or modify sleep habits according to the recommendation, adjust or modify nutrition habits according to the recommendation, follow an exercise recommendation, etc.). At operation 5250, the biometric activity application (e.g., the biometric activity application 3110 including the activity' agent 3120) may be configured to analyze the user data based on the action taken by the user. The user data can be stored in one or more memories 5270 which can comprise objective data 5272 as well as subject experience information 5274. At operation 5260 the user can provide feedback regarding the user data obtained at operation 5250. At operation 5280 the biometric activity7application (e.g., the biometric activity' application 3110 including the activity agent 3120) may be configured to analyze the feedback and provide the results of the analysis to the one or more machine-learned models as an input for generating an insight (e.g., a revised insight, a new insight, etc.) and for providing the recommendation (e.g., a revised recommendation, a new recommendation, etc.). The results of the analysis of the user's feedback can also be stored in the one or more memories 5270. In some implementations, the biometric activity application (e.g., the biometric activity application 3110 including the activity agent 3120) may be configured to retrieve stored feedback information from the one or more memories 5270 for generating the insight.

[0124] FIG. 5C illustrates an example process flow 5300 of the biometric activity application including an activity agent, according to examples of the disclosure. Referring to FIG. 5C,user data 5310 (and optionally additional inputs 5340) can be provided to one or more machine-learned models 5320 to generate one or more outputs 5330. The user data 5310 can include, for example, information about the user including the user’s name, age, gender, activities the user engages in, current weather associated with a location of the user, etc. The user data 5310 can also include biometric information about the user associated with one or more biometric activities. For example, in the example of FIG. 5C the user data 5310 includes information relating to goal of the user and actual achievements of the user (which may be measured by one or more sensors). In the example, the user achieved 30 minutes of high intensity cardio (with a goal of 60 minutes), 140 minutes of moderate intensity cardio (with a goal of 120 minutes), and engaged in four strength training sessions (with a goal of three strength training sessions).

[0125] In some implementations, additional inputs 5340 can be provided to the one or more machine-learned models 5320 for generating the one or more outputs 5330. For example, the additional inputs 5340 can include information relating to other biometric activities (e.g., sleep information, nutrition information, etc.). In some implementations, the additional inputs 5340 can be obtained from other activity agents (e.g., a sleep activity agent, a nutrition activity agent, etc.). As indicated in FIG. 5C, the one or more machine-learned models 5320 may be configured to generate the one or more outputs 5330 based on the user data 5310 and the additional inputs 5340, according to various prompts that are given to the one or more machine-learned models 5320. For example, the prompts can include instructions to provide an empathetic summary, to provide suggestions for meeting their goals, to ask open-ended questions, to celebrate successes, to compare the biometric activity goals with the achievements, to assess the mood and sentiment of the user, to determine a user’s personal motivation to engage in a particular activity7, to assess perceived barriers and frustrations, etc.

[0126] As indicated in FIG. 5C, the one or more outputs 5330 can include dialogue encouraging personal feedback (e.g., an open-ended question asking the user how they feel), dialogue for engaging in a two-way conversation with the user relating to goals and achievements, dialogue or recommendations relating various health dimensions of the user (e.g., interrelating sleep metrics associated with the user and eating habits and / or exercise habits), recommendations regarding goals that are appropriate for the user given the objective data (e.g., achievements) and subjective impressions of the user, etc. The dialogue operations can be in the form of natural language 5350, for example. For example, the one or more machine-learned models may be configured to employ natural language processing tocommunicate with the user via the dialogue operations and can also provide updated user data to the user (e.g., computing device 100) in a structured format (e.g., structured new user knowledge 5360). That is, the one or more outputs can further include updates to the user data on the user data 5310 and the additional inputs 5340, according to various prompts that are given to the one or more machine-learned models 5320. For example, the biometric activity application (e.g.. the biometric activity application 3110 including the activity agent 3120) can be configured to update information relating to user preferences, user mood and sentiment, perceived barriers and aspirations, and a user’s motivation to engage. As an example, the one or more machine-learned models 5320 may be configured to generate an output soliciting feedback regarding how a user feels after running a long distance, and the user may provide feedback such as '’happy.” The biometric activity application (e.g., the biometric activity application 3110 including the activity agent 3120) may be configured to associate the user’s feeling of “happy” with the user running a long distance, and can be configured to store the association in one or more memories (e.g., one or more memories 5270). Further, the association can be stored as part of a graph or network model that is searchable and interrelates various associations. For example, if the user queries the biometric activity application (e.g., the biometric activity application 3110 including the activity7agent 3120) regarding what activities make them feel happy, the biometric activity7application can be configured to search the graph or network model to determine the particular biometric activities which are associated with making the user feel happy (e.g., running a long distance, swimming in a lake, playing tennis, etc.). For example, if the user queries the biometric activity' application (e.g., the biometric activity' application 3110 including the activity agent 3120) regarding what activities are most effective at burning calories without making the user feel sore the next day, the biometric activity application can be configured to search the graph or network model to determine the particular biometric activities which are associated with the request and output those particular biometric activities.

[0127] FIG. 5D is an example user interface 5400 that can be provided to (or displayed at) a computing device (e.g., computing device 100) according to examples of the disclosure. In FIG. 5D, a first dialogue operation 5410 from the biometric activity application (e.g., the biometric activity application 3110 including the activity agent 3120) illustrates various outputs that include a summary of the user’s biometric activities (e.g., 24 minutes of high intensity7cardio, 91 minutes of low intensity cardio) and a suggestion to perform certainbiometric activities to meet the biometric goals of the user (e.g., focusing on strength training and an aerobic exercise session over the next three days). For example, the biometric information relating to the biometric activities of the user and the biometric goals relating to the biometric activities, can be obtained from the sensor data store 340, the biometric activity data store 350, from sensors 190, can be stored locally or remotely, etc. Further, the first dialogue operation 5410 includes encouraging feedback to the user (e.g., “Sara, you're doing great”), which can be generated via the one or more machine-learned models 5320.

[0128] At the second dialogue operation 5420, in response to the suggestions provided by the biometric activity application (e.g., the biometric activity application 3110 including the activity agent 3120), the user provides feedback indicating that the user has an injury and cannot lift.

[0129] At the third dialogue operation 5430, in response to the feedback provided by the user, the one or more machine-learned models 5320 revises the recommendation to exclude strength training while adding an extra high intensity cardio session. In some implementations, the one or more machine-learned models of the biometric activity application (e.g., the biometric activity application 3110 including the activity agent 3120) can be configured to store the information included in the dialogue exchange comprising a plurality of dialogue operations as an episode in memory. For example, the information can include the user’s feedback, biometric information, biometric goals, recommendation, user data, revised recommendation, etc. In some implementations, the one or more machine- learned models of the biometric activity application (e.g., the biometric activity application 3110 including the activity agent 3120) can be configured to store the information included in the dialogue exchange comprising a plurality of dialogue operations as an episode in memory. For example, the information can include the user’s feedback, biometric information, biometric goals, recommendation, user data, revised recommendation, etc. The information stored as part of the episode may also be mapped to certain features, for example, in a graph or network model. For example, the term “injury’” can be associated with the user’s health and / or strength training. In some implementations, the one or more machine- learned models of the biometric activity application (e.g., the biometric activity application 3110 including the activity agent 3120) can be configured to retrieve the episode in memory7in a subsequent dialogue exchange with the user, for example, for making subsequent recommendations, for asking open-ended questions, etc. For example, the one or more machine-learned models of the biometric activity application (e.g., the biometric activityapplication 3110 including the activity agent 3120) can refer to the stored episode and ask the user whether their injury has healed or ask whether the user is ready for strength training, or explain that a recommendation is based on an assumption the user remains injured.

[0130] Examples of the disclosure are directed to computer implemented methods for recommendation, content, and / or dialogue generation systems including implementing one or more episodic memories which can be used to store episode information relating to context information extracted by one or more machine-learned models from dialogue information received via an activity’ agent and relating to a first biometric activity associated with the user.

[0131] The flow diagram of FIG. 6 illustrates a method 6100 for generating an output based on information included in an episode store in one or more memories. Although shown in a particular sequence or order, unless otherwise specified, the order of the processes can be modified. Thus, the illustrated embodiments should be understood only as examples, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel. Additionally, one or more processes can be omitted in various embodiments. Thus, not all processes are required in every embodiment. Other process flows are possible.

[0132] The operations of FIG. 6 will be explained with reference to FIG. 7. FIG. 6 illustrates an example block diagram or architecture of a computing system (recommendation, content, and / or dialogue generation system) 7100 including a biometric activity application 7110, according to one or more example embodiments of the disclosure.

[0133] Referring to FIG. 6, at operation 6110 the method 2100 includes a computing system extracting, via one or more machine-learned models, context information from dialogue information received from a user via an activity7agent. As described herein, the computing system may be embodied as computing device 100, sen- er computing system 300, or combinations thereof. For example, in the computing system 7100 of FIG. 7, in some implementations a biometric activity7application 7110 includes an activity agent 7120 which may be configured to interact with the user and to receive dialogue information from the user via dialogue operations, one or more machine-learned models 7130, and one or more expert machine-learned models 7140. The activity agent 7120 may include the engagement engine 7122, context extractor 7124, and recommendation engine 7126. The features and operations of the biometric activity application 7110 including the activity' agent 7120, one or moremachine-learned models 7130. and one or more expert machine-learned models 7140 correspond to the biometric activity application 3110 including the activity agent 3120, one or more machine-learned models 3130, and one or more expert machine-learned models 3140 of FIG. 3 and a repeated description of these features and operations will not be repeated for the sake of brevity.

[0134] As illustrated in FIG. 7, the biometric activity application 7110 may receive as inputs one or more of a user input 7102, biometric information 7104, biometric activity information 7106, and external information 7108. The features and operations of the user input 7102, biometric information 7104, biometric activity information 7106, and external information 7108 correspond to the user input 3102, biometric information 3104, biometric activity information 3106, and external information 3108 of FIG. 3 and a repeated description of these features and operations will not be repeated for the sake of brevity.

[0135] For example, in the computing system 7100 of FIG. 7, the context extractor 7124 maybe configured to extract information from the dialogue information (e.g., dialogue from the user) to determine a context (context information) associated with the dialogue. For example, the context extractor 7124 may be configured to determine or identify one or more entities from the context information including names of people, places, things, etc., objects, time information, to determine or identify a topic of the dialogue, and / or to determine or identify a goal or purpose (e.g., intent) of the dialogue. In some implementations, other information including emotional or sentiment information associated with the user can be determined from the dialogue.

[0136] FIG. 8 is an example user interface which includes dialogue operations between the user and computing system, according to examples of the disclosure. Referring to the example user interface 8000 of FIG. 8, in a first dialogue operation 8010 the user provides dialogue information indicating a workout felt “really good7’ while noting certain aspects of the workout that were deficient as well as information relating to pain felt during the workout. The context extractor 7124 may be configured to determine, based on the dialogue information received from the user via the first dialogue operation 8010, context information including terms or phrases which represent one or more topics or themes, one or more entities, one or more intents, one or more events, one or more sentiments or emotions, etc. In some implementations, prior dialogue operations can also be utilized to further determine or identify- the context information. In some implementations, the context extractor 7124 maybe configured to employ natural language extraction to extract structured information fromunstructured natural language text. For example, based on the first dialogue operation 8010 in FIG. 8, the context extractor 7124 may be configured to implement one or more machine- learned models with respect to the dialogue information to determine topics and / or keywords related to the workout (first biometric activity) such as “felt really good,” “more variety,” “pain in shoulder,” and “plank-walkouts.”

[0137] At operation 6120 the method 6100 includes the computing system receiving biometric information associated with the user with respect to the first biometric activity. For example, in the computing system 7100 of FIG. 7, the biometric activity application 7110 may be configured to receive biometric information 7104 associated with the user and the first biometric activity which may be part of biometric activity information 7106. In some implementations, the biometric activity application 7110 may also receive external information 7108 which may include sensor data that is not necessarily biometric information associated with the user, but may include other information (e.g.. related to the dialogue and / or biometric activity) such as environmental information (e.g., time information, temperature information, weather information, noise information, calendar information, location information, etc.). The biometric information 7104 may be captured by one or more sensors of the computing system itself and / or may be captured by one or more sensors of an external computing device. Similarly, other sensor data may be captured by one or more sensors of the computing system itself and / or may be captured by one or more sensors of an external computing device. The biometric information 7104 can include information relating to biometrics associated with an ECG. PPG. heart rate, heart rate recovery, pulse information, BML heart rate variability, blood pressure, oxygen saturation, body temperature, sleep quality, physical activities (e.g., number of repetitions performed, number of sets completed, number of steps w alked, number of miles cycled, number of laps swam, etc.), and the like. Further, the computing system (biometric activity application 7110) may be configured to generate or display information associated with a measured biometric, including an electrocardiogram, a photoplethysmogram, heart rate, heart rate recovery, blood pressure, oxygen saturation, respiration rate, body temperature, physical activity, a sleep metric, electrical conductance, and the like. In the example of FIG. 8, the first dialogue operation 8010 indicates the user performed a workout which included doing plank- walkouts. The biometric information may include metrics relating to the w orkout including the exercises performed, repetitions and / or sets completed, movements performed during the workout, the duration of time spent working out, etc. The metrics may be measured via one or moresensors of the computing system or of an external computing device (e.g., a camera, a wearable computing device, etc.).

[0138] At operation 6130 the method 6100 includes the computing system storing, as a first episode among a plurality of episodes in the one or more memories, the context information, the first biometric activity (data associated with the first biometric activity), and the first biometric information. For example, in the computing system 7100 of FIG. 7, the biometric activity application 7110 may be configured to store the context information, the first biometric activity (data associated with the first biometric activity). and the first biometric information in one or more memories, for example, in an episodic memory 7170 which includes a plurality of episodes 7172 (e.g., a first episode, a second episode, etc.). In some implementations, each of the plurality of episodes may be associated with a time interval that is less than a predetermined duration of time. In some implementations, each of the plurality of episodes may be associated with a dialogue session in which the user interacts with the computing system before exiting or terminating the biometric activity' application 7110 or after a predetermined number of dialogue operations. For example, the biometric activity application 7110 may be configured to store the context information, the first biometric activity (data associated with the first biometric activity), and the first biometric information that is obtained over a period of time during a dialogue session that is less than a threshold duration of time (e.g., less than one minute). For example, the biometric activity application 7110 may be configured to store the context information that is obtained during a dialogue session that is less than a threshold duration of time (e.g., less than one minute) and during the performance of the first biometric activity. For example, in the example of FIG. 8. the biometric activity application 7110 may be configured to store as part of the first episode of the plurality of episodes 7172, the context information obtained during the dialogue session (e.g., after the first dialogue operation 8010 or after the second dialogue operation 8020) and the first biometric information that is obtained during the performance of the first biometric activity (e g., the workout). In this manner, particular information associated with a particular biometric activity such as a user’s sentiment or information about a user’s health, can be stored in a segmented manner that can easily be searched. For example, at operation 6110 the extracting, via the one or more machine-learned models, the context information from the dialogue information can include identifying one or more topics in the dialogue information and mapping the one or more topics to at least one of a network model or a graph model. The netw ork model or graph model can be searched to determine relationshipsbetween various features. For example, in the example of FIG. 8, if the user's workout is focused on abdominal muscles, the network model or graph model can associate or map the user’s emotion of “feeling really good” with an ab workout. For example, in the example of FIG. 8, the network model or graph model can associate or map the user’s preference of liking “variety” with biometric activities and / or with particularly an ab workout. For example, in the example of FIG. 8, the network model or graph model can associate or map the user’s feedback of feeling pain in her shoulder with plank walkouts.

[0139] At operation 6140 the method 6100 includes the computing system receiving a query associated with at least one of the context information or the first biometric activity. At operation 6150 the method 6100 includes the computing system searching the plurality of episodes stored in the one or more memories according to the query'. For example, in the computing system 7100 of FIG. 7. the biometric activity’ application 7110 (e.g., the activity’ agent 7120) may be configured to receive the query as part of the user input 7102. The query may be a request to the biometric activity application 7110 (e.g., the activity agent 7120) regarding a recommendation for a biometric activity' or other coaching advice. In the example of FIG. 8, at the second dialogue operation 8020 the biometric activity' application 7110 (e.g., the activity agent 7120) indicates that the feedback provided by the user will be considered in future operations (e.g., recommendations). For example, if the user provides a query’ for workouts that make them feel “good,” the biometric activity application 7110 (e.g., the activity agent 7120) may be configured to search the plurality of episodes 7172 which correspond to or include context information that corresponds to the query. In the example of FIG. 10, the biometric activity application 7110 (e.g.. the activity agent 7120) may also be configured to generate queries (e.g., questions) for the user to ask based on the content (e.g., context) of the dialogue operations and to provide the queries as selectable user interface elements 8030. In some implementations, if the user selects one of the selectable user interface elements 8030. the biometric activity’ application 7110 (e.g., the activity’ agent 7120) may be configured to search the plurality of episodes 7172 which correspond to or include context information that corresponds to the query' in determining an output (e.g., a recommendation, answer to the question, dialogue operations, etc.). For example, if the user provides a query for workouts that make them feel “good,” the biometric activity application 7110 (e.g., the activity agent 7120) may be configured to search the plurality of episodes 7172 which correspond to or include context information that corresponds to the query. In some implementations, searching the plurality of episodes 7172 stored in the one or morememories (e.g., the episodic memory' 7170) according to the query may include traversing the at least one of the network model or the graph model to determine the one or more topics mapped to the at least one of the network model or the graph model which correspond to at least one topic in the query. For example, the first episode may store episodic information relating to the first biometric activity (e.g., the abs workout) which may be mapped to context information such as "‘feeling really good.” The biometric activity application 7110 (e.g., the activity agent 7120) may be configured to search the plurality of episodes and retrieve the first biometric activity from the first episode as an activity that makes the user feel good and which corresponds to the topic in the query.

[0140] At operation 6160 the method 6100 includes, when the first episode includes information yvhich is responsive to the query-, providing, by the computing system and for presentation to the user, one or more outputs related to the first biometric activity and the context information. For example, in the computing system 7100 of FIG. 7, the biometric activity application 7110 (e g., the activity agent 7120) may provide, for presentation to the user, the one or more outputs (e.g., via the display device 160 and / or output device 170). In some implementations, the output 7150 may be provided in a dialogue exchange where the computing system provides a dialogue indicating the activity that makes the user feel good ("’You usually feel really good after doing an abs workout”). In some implementations, the biometric activity application 7110 (e.g., the activity agent 7120) may also provide further commentary based on the user’s prior feedback regarding the first biometric activity (e.g., “Here’s my plan for an abs workout - I’ve added more variety and avoided exercises involving your shoulder since you felt some pain in your last abs workout”).

[0141] As example implementations, the query can include context information (e.g., an emotional state of a user, mood of the user, sentiment of the user, etc.), biometric information (e.g., physiological information of the user), biometric activity information, etc. When the query is associated with the context information, information included in the plurality' of episodes 7172 stored in the one or more memories is searched for the context information. When the first episode includes information corresponding to the context information, the biometric activity application 7110 (e.g., the activity agent 7120) can provide for presentation to the user one or more outputs related to the first biometric activity and the context information. For example, if the context information in the dialogue information includes an emotional state of the user (e.g., “what makes me feel good?”) the plurality of episodes can be searched for information corresponding to the emotional state of the user. When the firstepisode includes information corresponding to the emotional state of the user, the one or more outputs provided for presentation to the user can indicate the first biometric activity- is associated with the emotional state of the user (e.g., “doing an abs workout will leave you feeling great!”).

[0142] When the query is associated with the first biometric activity, information included in the plurality7of episodes stored in the one or more memories is searched for the first biometric activity. For example, if the user queries “how do ab workouts make me feel?”, the biometric activity application 7110 (e.g.. the activity agent 7120) can be configured to search the episodic memory77170 (e.g., the plurality of episodes 7172) to determine episodes which include ab workouts. When the first episode includes information corresponding to the first biometric activity-, the biometric activity application 7110 (e.g., the activity agent 7120) can be configured to provide, for presentation to the user, one or more outputs related to the first biometric activity and the context information. For example, the biometric activity application 7110 (e.g., the activity agent 7120) may determine from the first episode that ab workouts make the user feel really good (e.g., based on emotional state information associated with the user being mapped to the first biometric activity) and an output can be provided in accordance with that determination. For example, the one or more outputs provided for presentation to the user can indicate the emotional state of the user based on the emotional state information associated with the user mapped to the first biometric activity. In some implementations, if a plurality of episodes include ab workouts, the biometric activityapplication 7110 (e.g., the activity agent 7120) may be configured to determine an overall sentiment relating to the ab workouts (e.g., based on an average of sentiments or emotional states that are fed back to the biometric activity application 7110 (e.g., the activity agent 7120) and mapped to the ab workouts).

[0143] When the query is associated with physiological information, information included in the plurality- of episodes stored in the one or more memories is searched for the physiological information. For example, if the user queries “I want to bum the most calories in the fastest time possible”, the biometric activity application 7110 (e.g., the activity agent 7120) can be configured to search the episodic memory77170 (e.g., the plurality of episodes 7172) to determine episodes in which calories are burned at the highest rate. When the first episode includes information corresponding to the physiological information, the biometric activityapplication 7110 (e.g., the activity agent 7120) can be configured to provide, for presentation to the user, one or more outputs indicate the physiological information which is associatedwith the first biometric activity. For example, the biometric activity' application 7110 (e.g., the activity agent 7120) may determine from the first episode that a high intensity interval training workout bums the most calories in a given time (e.g., based on physiological information that is associated with or mapped to the high intensity interval training workout) relative to other biometric activities that are stored in the episodic memory 7170 and an output can be provided in accordance with that determination. In some implementations, if a plurality of episodes (or plurality of biometric activities) meet criteria for the physiological information, the biometric activity application 7110 (e g., the activity agent 7120) may be configured to present a plurality' of biometric activities for selection by the user.

[0144] When the query is associated with a biometric activity', information included in the plurality' of episodes stored in the one or more memories is searched for the biometric activity. For example, if the user queries ’’How effective is running long distance for burning calories”, the biometric activity application 7110 (e.g., the activity agent 7120) can be configured to search the episodic memory 7170 (e.g., the plurality of episodes 7172) to determine episodes having corresponding to the first biometric activity' (e.g., running long distances, such as over 5 miles). When the first episode includes the first biometric activity, the biometric activity application 7110 (e.g., the activity agent 7120) can be configured to retrieve from the first episode the physiological information associated with the user mapped to the first biometric activity'. For example, the biometric activity' application 7110 (e.g., the activity agent 7120) can be configured to provide, for presentation to the user, one or more outputs which indicate the first biometric activity is associated with the physiological information. For example, the biometric activity application 7110 (e.g., the activity agent 7120) may determine from the first episode (or from a plurality of episodes) that a long run by the user bums a certain number of calories in a given time (e.g., based on physiological information that is associated with or mapped to long runs by the user and stored in the episodic memory 7170) and an output can be provided in accordance with that determination.

[0145] In some implementations, the computing system may further receive second biometric information associated with the user relating to a second biometric activity, the second biometric information being captured by one or more sensors, and the second biometric information and second biometric activity' are mapped to the first biometric activity' and stored as part of the first episode among the plurality of episodes 7172 in the one or more memories (e.g., the episodic memory 7170). For example, biometric information associatedwith a user consuming food (a nutrition biometric activity) may be mapped to a sleep biometric activity and stored as part of the first episode.

[0146] When a query- is associated with the first biometric activity and the second biometric activity, the biometric activity application 7110 (e.g., the activity agent 7120) can be configured to search information included in the plurality of episodes stored in the one or more memories for the first biometric activity and the second biometric activity. When the first episode includes the first biometric activity and the second biometric activity, the biometric activity application 7110 (e.g., the activity agent 7120) can be configured to retrieve, from the first episode, context information associated with the user which is mapped to the first biometric activity and the second biometric activity. For example, if the user indicates via dialogue operations that they have a hard time going to sleep after eating a late night meal, the contextual information (e.g.. poor sleep quality or difficulty sleeping) may be mapped to the combination of a sleep activity and nutrition activity. The one or more outputs provided for presentation to the user can indicate the context information associated with the user which is mapped to the first biometric activity and the second biometric activity. For example, if the user asks what effect eating before bedtime has on their sleep, the biometric activity application 7110 (e.g.. the activity agent 7120) can be configured to retrieve, from the first episode (or other episodes which satisfy the query criteria), context information associated with the user which is mapped to the sleep activity and the nutrition activity and can indicate that the user generally has poor sleep quality when the user eats within a certain time period before going to bed. In some implementations, the biometric activity application 7110 (e.g.. the activity agent 7120) can determine other metrics that may be provided to the user in connection with the one or more outputs (e.g., reporting a certain average number of minutes to get to sleep after eating close to bedtime based on sensor data).

[0147] Examples of the disclosure are directed to computer implemented methods for recommendation, content, and / or dialogue generation systems including implementing a computing platform to coordinate operations of a plurality of activity agents.

[0148] The flow diagram of FIG. 9 illustrates a method 9100 implemented by a computing platform for coordinating operations of a plurality of activity agents. Although shown in a particular sequence or order, unless otherwise specified, the order of the processes can be modified. Thus, the illustrated embodiments should be understood only as examples, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel. Additionally, one or more processes can be omitted in variousembodiments. Thus, not all processes are required in every embodiment. Other process flows are possible.

[0149] The operations of FIG. 9 will be explained with reference to FIG. 10. FIG. 9 illustrates an example block diagram or architecture of a computing system (recommendation, content, and / or dialogue generation system) 1000 including a computing platform 1020 and a plurality of activity agents 1228, according to one or more example embodiments of the disclosure. FIG. 10 illustrates an example block diagram of a system including a computing platform, according to one or more example embodiments of the disclosure;

[0150] Referring to FIG. 9, at operation 9110 the method 9100 includes a computing platform providing, to a plurality of activity agents including at least a first activity’ agent and a second activity agent, a common framework for interacting with a user via dialogue operations. As described herein, the computing platform may be embodied as server computing system 300, however in some implementations the computing platform may be implemented via computing device 100, or combinations thereof. For example, in the computing system 1000 of FIG. 10, in some implementations the computing platform 1020 may correspond to the server computing system 300 of FIG. IB, and therefore a description of each of the features of the server computing system 300 (e.g., one or more processors, one or more memories, application system, etc.) will not be repeated for the sake of brevity.

[0151] As illustrated in FIG. 10. the computing platform 1020 can receive input information 1002 which may include a user input, biometric information, biometric activity information, external information, episodic information, etc. The computing platform 1020 can store an activity agent framework 1022 which can be provided to the plurality of activity agents 1030 (e.g., a first activity agent 1032, a second activity agent 1034. an nth activity agent 1036. etc.). The activity agent framework 1022 corresponds to a common framework for interacting with a user via dialogue operations. Therefore, each of the plurality’ of activity agents 1030 can communicate and interact with a user in a uniform and consistent manner. For example, the common framework can include a plurality of rules and a plurality of expectations for communicating with the user and for providing recommendations, that are uniform among the plurality of activity agents. Example rules which may be common to the plurality of activity’ agents 1030 may include general strategies to utilize when speaking with the user (e.g., identity’ and celebrate successes, suggest areas for improvement, provide recommendations where the user can make their own choices, set goals which are SMART(specific, measurable, achievable, relevant, and time-bound), keeping dialogue to a predetermined length, personalize the messaging, never make judgmental and critical statements, etc.). The rules can include style rules that define guidelines associated with the length of the text, use of the data, how to break up text, ensure readability, provide a diversity of messages and personalization. The rules can include tone rules that help users stay motivated and ensure the dialogue communications are supportive and friendly. The rules can include psychological rules that provide that the dialogue communications are non- judgmental, empower the user, leverage SMART goals, and reference health benefits. The rules can include safety' rules that ensure the activity agents (e.g., the one or more machine- learned models of the activity agent) avoid providing medical advice and / or information which could be unsafe for the user. For example, the common framework can include an evaluation framework to evaluate biometric achievements of the user against biometric goals of the user, that is uniform and consistent among the plurality of activity agents 1030. The evaluation framework can include certain features which are common to each of the plurality of activity agents 1030 (e.g.. a structured format that includes the user's name, age, gender, etc.). The evaluation framework can also include common metrics and a structured format for evaluating the progress of a user by analyzing actual achievement compared to a particular goal with respect to a particular biometric activity (e.g., comparing a current week against a prior week, tracking progress for a current week, etc.). The evaluation framework can also include providing one or more machine-learned models 1024 that can be utilized for auto-evaluating the performance of the activity agents (e g., the one or more machine-learned models of the activity agents). For example, the computing platform 1020 may be configured to implement one or more expert machine-learned models which are configured to validate one or more outputs determined by one or more machine-learned models associated with the first activity agent 1032 (or other activity agents), where the one or more outputs are related to a biometric activity and the user. For example, the one or more expert machine-learned models can help prevent or avoid the activity agents from providing medical advice and / or information yvhich could be unsafe for the user.

[0152] In some implementations, the plurality of activity agents 1030 can each have a different persona that is customized or configured according to the activity associated with the activity agent. For example, the plurality of activity agents can include at least a sleep activity7agent that provides coaching and motivational advice to the user regarding the user’s sleep, an exercise activity agent that provides coaching and motivational advice to the userregarding one or more exercises that the user performs, and a nutrition activity agent that provides coaching and motivational advice to the user regarding the diet of the user. In some implementations, activity agents may be provided for specific exercises (e.g., a swimming activity agent, a cycling activity agent, a running activity agent, etc.).

[0153] Referring back to FIG. 9, at operation 9120 the method 9100 includes the computing platform coordinating, between the plurality of activity agents, dialogue information associated with the user received via the dialogue operations via respective activity agents among the plurality’ of activity agents. For example, the computing platform 1020 may be configured to coordinate between the plurality of activity agents 1030 the dialogue information associated with the user by providing context information extracted from dialogue information received from the user via the first activity' agent 1032, to the second activity agent 1034 and / or to other activity’ agents. As an example, if the first activity agent 1032 (e.g.. a weightlifting activity agent) receives dialogue information from the user and extracts context information indicating that the user has suffered a shoulder injury, the computing platform 1020 may be configured to receive that context information and share the context information with other activity' agents (e.g., a second activity agent such as a swimming activity agent). For example, the swimming activity agent may then adjust or modify biometric goals and / or biometric activities associated with the user in light of receiving the context information, and may also adjust or modify dialogue operations to empathize wi th the user in light of their injury. The computing platform 1020 may be configured to implement one or more classifiers (e.g.. via the one or more machine-learned models 1024) to classify information from the context information. For example, the one or more classifiers can include at least one of a topic classifier to classify one or more topics included in the context information, an emotional classifier to classify’ one or more emotional states associated with the user indicated by the context information, or a decision classifier to classify one or more decisions of the user indicated by the context information. The classified context information can also be shared with the other activity agents to relieve a processing burden of the other activity’ agents.

[0154] The computing platform 1020 may be configured to share the context information to only those activity agents to which the context information is relevant, thereby conserving computing resources. For example, if the shoulder injury would not affect the performance of a particular biometric activity associated with an activity agent, that activity agent may not receive the context information from the computing platform 1020. The computing platform1020 may also be configured to share the context information to only those activity agents which are permitted to receive the context information, thereby improving security of the information.[01551 At operation 9130 the method 9100 includes the computing platform coordinating, between the plurality' of activity agents, biometric information associated with the user relating to one or more biometric activities, the biometric information being captured by one or more sensors. For example, the computing platform 1020 may be configured to coordinate between the plurality of activity agents 1030 the biometric information associated with the user relating to the one or more biometric activities by providing first biometric information captured by the one or more sensors and input to the first activity agent 1032 while the user performs a first biometric activity7, to the second activity' agent 1034 and / or to other activity' agents. As an example, if the first activity agent 1032 (e.g.. a sleep activity agent) receives biometric information from the user indicating the user has slept only three hours, the computing platform 1020 may be configured to receive that biometric information and share the biometric information with other activity agents (e.g., a second activity' agent such as a running activity agent). For example, the running activity agent may then adjust or modify biometric goals and / or biometric activities associated with the user in light of receiving the biometric information, and may also adjust or modify dialogue operations to empathize with the user in light of their lack of sleep. The computing platform 1020 may be configured to share the biometric information to only those activity agents to which the biometric information is relevant, thereby conserving computing resources. The computing platform 1020 may also be configured to share the biometric information to only those activity agents which are permitted to receive the biometric information, thereby improving security' of the information.

[0156] When computing platform 1020 interacts with the plurality of activity' agents 1030 and / or a user, the computing platform 1020 may be configured to ensure that information (data) transmitted to or received from the activity agents and / or user is secure, so as to prevent information from being leaked or misused. In some implementations, the computing platform 1020 may be configured to implement end-to-end encryption for securing communications during data transfer between the activity agents and / or user and the computing platform 1020, to ensure that the information is accessible only to its intended recipient(s). Other security measures may additionally, or alternatively, be implemented by the computing platform 1020 to secure sensitive information including storing sensitiveinformation with non-sensitive placeholders (e.g., tokens), anonymizing information by removing personally identifiable information from the data before transmitting the information, masking sensitive information, utilizing multi-factor authentication methods, etc.

[0157] In some implementations, the computing platform 1020 may be configured to store information associated with the user and / or activity agents (e.g., dialogue information, context information, biometric information, etc.) in an encrypted manner such that stored data is converted into an unreadable format that can only be decrypted by an entity or computing system having the proper encryption or cryptographic key. Data can be encrypted at various levels for multiple layers of security (e.g., at the database level including at a row, column, field, level, etc., at an application level, at a file level, at a disk or storage system level, etc.).

[0158] Further, the computing platform 1020 may be configured to provide a user with controls allowing the user to make an election as to both if and when systems, programs, or features described herein may enable collection of user information (e.g.. information about a user’s social network, social actions, biometric activities, calendar information, dialogue communications, profession, a user’s preferences, a user’s current location, etc.), and if the user is sent content or communications from an activity agent, server computing system, or computing platform. In addition, certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user’s identity' may be treated so that no personally identifiable information can be determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined. Thus, the computing platform 1020 may be configured to provide a user with controls over what information is collected about the user, how that information is used, and what information is provided to the user. Further, the computing platform 1020 may be configured to provide a user with controls over what information is stored and the ability to delete information about the user which has been stored.

[0159] At operation 9140 the method 9100 includes the computing platform performing one or more deconflicting operations between the plurality of activity agents to ensure a first biometric activity associated with the first activity agent does not conflict with a second biometric activity associated with the second activity agent. For example, in the computing system 1000 of FIG. 10, the computing platform 1020 may be configured to perform one or more deconflicting operations between the plurality of activity agents 1030 to ensure a firstbiometric activity' associated with the first activity agent 1032 does not conflict with a second biometric activity associated with the second activity agent 1034 and / or with biometric activities of other activity agents. For example, the computing platform 1020 may be configured to ensure a time to perform the first biometric activity to achieve a first biometric goal generated by the first activity agent 1032 does not conflict with a second biometric activity to achieve a second biometric goal generated by the second activity agent 1034. As an example, if the first activity agent 1032 (e.g.. a running activity agent) indicates to the computing platform 1020 that the user is to run a midnight marathon, the computing platform 1020 may ensure that the second activity agent 1034 (e.g., a sleep activity agent) does not recommend a plan for the user to go to sleep at 11 pm.

[0160] For example, the computing platform 1020 may be configured to determine whether performing the first biometric activity to achieve a first biometric goal reduces a likelihood of the user completing the second biometric activity to achieve a second biometnc goal by more than a predetermined amount. In some implementations, the one or more machine-learned models 1024 of the computing platform 1020 may be configured to predict or determine whether if a user performs a first biometric activity (e g., running a marathon) to achieve a first biometric goal (e.g.. running a certain number of miles in the week) will reduce the likelihood of the user completing another (second) biometric activity (e.g., strength training) to achieve a second biometric goal (e.g., performing a certain number of strength training sessions in a week), by more than a predetermined amount. For example, if the one or more machine-learned models 1024 of the computing platform 1020 determine it is unlikely that the user will complete the second biometric activity (or achieve the second biometric goal), the computing platform 1020 may be configured to notify the respective activity agents so that the activity agents can adjust or modify' the biometric goals or recommendations associated with the biometric activity, accordingly. In addition, the computing platform 1020 and / or respective activity agents may be configured to alert or notify the user of the conflict via dialogue operations so that the user is aware of the potential conflict in achieving different biometric goals.

[0161] In some implementations, the computing platform 1020 may be configured to receive feedback from the user with respect to one or more outputs determined by one or more machine-learned models associated with the first activity agent, where the one or more outputs are related to a biometric activity and the user. The computing platform 1020 may be configured to communicate the feedback to at least one other activity agent among theplurality of activity agents. As an example, in the first dialogue operation 8010 of FIG. 8, the user provided feedback indicating they preferred variety’ to a first activity agent (e.g., a strength training activity agent). This feedback may be provided to the computing platform 1020 which may' be configured to communicate the feedback to other activity' agents which can incorporate such feedback in providing recommendations to the user and generating plans (e.g., a workout plan, nutrition plan, etc.). Therefore, consistent and appropriate recommendations may be provided to the user across a plurality’ of activity agents, helping to ensure that future recommendations are likely to be accepted by the user and executed.

[0162] FIG. 11 depicts a flowchart of a method 1110 for training one or more machine- learned models according to aspects of the disclosure. For instance, an example machine- learned model can include one or more of a LLM, a generative machine-learned model, etc. For example, the one or more machine-learned models may be configured to implement the operations of the biometric activity applications as descnbed herein.

[0163] FIG. 11 is a flow chart diagram illustrating an example method for training a machine-learned model according to example implementations of aspects of the disclosure. One or more portion(s) of example method 1110 can be implemented by a computing system that includes one or more computing devices such as, for example, computing systems described with reference to the other drawings. Each respective portion of example method 1110 can be performed by any (or any combination) of one or more computing devices. Moreover, one or more portion(s) of example method 1110 can be implemented on the hardware components of the device(s) described herein, for example, to train one or more systems or models. FIG. 11 depicts elements performed in a particular order for purposes of illustration and discussion. Those of ordinary skill in the art, using the disclosures provided herein, will understand that the elements of any of the methods discussed herein can be adapted, rearranged, expanded, omitted, combined, or modified in various ways without deviating from the scope of the disclosure. FIG. 11 is described with reference to elements / terms described with respect to other systems and drawings for exemplary illustrated purposes and is not meant to be limiting. One or more portions of example method 1 110 can be performed additionally, or alternatively, by other systems.

[0164] At 11 12, example method 1110 can include obtaining a training instance. A set of training data can include a plurality of training instances divided between multiple datasets (e.g., a training dataset, a validation dataset, or testing dataset). A training instance can be labeled or unlabeled. Although referred to in example method 1110 as a “training” instance.it is to be understood that runtime inferences can form training instances when a model is trained using an evaluation of the model's performance on that runtime instance (e.g., online training / leaming). Example data types for the training instance and various tasks associated therewith are described throughout the disclosure.

[0165] At 11 14, example method 1 110 can include processing, using one or more machine- learned models, the training instance to generate an output. The output can be directly obtained from the one or more machine-learned models or can be a dow nstream result of a chain of processing operations that includes an output of the one or more machine-learned models.

[0166] At 11 16, example method 1110 can include receiving an evaluation signal associated with the output. The evaluation signal can be obtained using a loss function. Various determinations of loss can be used, such as mean squared error, likelihood loss, cross entropy loss, hinge loss, contrastive loss, or various other loss functions. The evaluation signal can be computed using known ground-truth labels (e.g., supervised learning), predicted or estimated labels (e.g., semi- or self-supervised learning), or without labels (e.g., unsupervised learning). The evaluation signal can be a reward (e.g., for reinforcement learning). The reward can be computed using a machine-learned reward model configured to generate rewards based on output(s) received. The reward can be computed using feedback data describing human feedback on the output(s).

[0167] At 1118, example method 1110 can include updating the machine-learned model using the evaluation signal. For example, values for parameters of the machine-learned model(s) can be learned, in some embodiments, using various training or learning techniques, such as, for example, backwards propagation. For example, the evaluation signal can be backpropagated from the output (or another source of the evaluation signal) through the machine-learned model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the evaluation signal with respect to the parameter value(s)). For example, system(s) containing one or more machine-learned models can be trained in an end-to-end manner. Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations. In some implementations, performing backwards propagation of errors can include performing truncated backpropagation through time. Example method 1110 can include implementing a number of generalization techniques (e.g., weight decays, dropouts, etc.) to improve the generalization capability' of the models being trained.

[0168] In some implementations, example method 1110 can be implemented for training a machine-learned model from an initialized state to a fully trained state (e.g., when the model exhibits a desired performance profile, such as based on accuracy, precision, recall, etc.).

[0169] In some implementations, example method 1110 can be implemented for particular stages of a training procedure. For instance, in some implementations, example method 1 11 can be implemented for pre-training a machine-learned model. Pre-training can include, for instance, large-scale training over potentially noisy data to achieve a broad base of performance levels across a variety of tasks / data types. In some implementations, example method 1110 can be implemented for fine-tuning a machine-learned model. Fine-tuning can include, for instance, smaller-scale training on higher-quality (e.g., labeled, curated, etc.) data. Fine-tuning can affect all or a portion of the parameters of a machine-learned model. For example, various portions of the machine-learned model can be ‘"frozen’' for certain training stages. For example, parameters associated with an embedding space can be “frozen” during fine-tuning (e.g., to retain information learned from a broader domain(s) than present in the fine-tuning dataset(s)). An example fine-tuning approach includes reinforcement learning. Reinforcement learning can be based on user feedback on model performance during use.Example Machine-Learned Models

[0170] FIG. 12 is a block diagram of an example processing flow for using machine-learned model(s) 1 to process input(s) 2 to generate output(s) 3.

[0171] Machine-learned model(s) 1 can be or include one or multiple machine-learned models or model components. Example machine-learned models can include neural networks (e.g., deep neural networks). Example machine-learned models can include nonlinear models or linear models. Example machine-learned models can use other architectures in lieu of or in addition to neural networks. Example machine-learned models can include decision tree based models, support vector machines, hidden Markov models, Bayesian networks, linear regression models, k-means clustering models, etc.

[0172] Example neural networks can include feed-forward neural networks, recurrent neural networks (RNNs), including long short-term memory (LSTM) based recurrent neural networks, convolutional neural networks (CNNs), diffusion models, generative-adversarial networks, or other forms of neural networks. Example neural networks can be deep neural networks. Some example machine-learned models can leverage an attention mechanism suchas self-atention. For example, some example machine-learned models can include multiheaded self-attention models.

[0173] Machine-learned model(s) 1 can include a single or multiple instances of the same model configured to operate on data from input(s) 2. Machine-learned model(s) 1 can include an ensemble of different models that can cooperatively interact to process data from input(s) 2. For example, machine-learned model(s) 1 can employ a mixture-of-experts structure. See, e.g., Zhou et al., Mixture-of-Experts with Expert Choice Routing, ARXlV:2202.09368v2 (Oct. 14, 2022).

[0174] Input(s) 2 can generally include or otherwise represent various types of data. Input(s) 2 can include one type or many different types of data. Output(s) 3 can be data of the same type(s) or of different types of data as compared to input(s) 2. Output(s) 3 can include one type or many different types of data.

[0175] Example data types for input(s) 2 or output(s) 3 include natural language text data, software code data (e.g., source code, object code, machine code, or any other form of computer-readable instructions or programming languages), machine code data (e.g.. binary code, assembly code, or other forms of machine-readable instructions that can be executed directly by a computer's central processing unit), assembly code data (e.g., low-level programming languages that use symbolic representations of machine code instructions to program a processing unit), genetic data or other chemical or biochemical data, image data, audio data, audiovisual data, haptic data, biometric data, medical data, financial data, statistical data, geographical data, astronomical data, historical data, sensor data generally (e.g., digital or analog values, such as voltage or other absolute or relative level measurement values from a real or artificial input, such as from an audio sensor, light sensor, displacement sensor, etc.), and the like. Data can be raw or processed and can be in any format or schema.

[0176] In multimodal inputs 2 or outputs 3, example combinations of data types include image data and audio data, image data and natural language data, natural language data and software code data, image data and biometric data, sensor data and medical data, etc. It is to be understood that any combination of data types in an input 2 or an output 3 can be present.

[0177] An example input 2 can include one or multiple data types, such as the example data ty pes noted above. An example output 3 can include one or multiple data types, such as the example data types noted above. The data type(s) of input 2 can be the same as or different from the data type(s) of output 3. It is to be understood that the example data types notedabove are provided for illustrative purposes only. Data types contemplated within the scope of the disclosure are not limited to those examples noted above.Example Machine-Learned Sequence Processing Models

[0178] FIG. 13 is a block diagram of an example implementation of an example machine- learned model configured to process sequences of information. For instance, an example implementation of machine-learned model(s) 1 can include machine-learned sequence processing model(s) 4. An example system can pass input(s) 2 to sequence processing model(s) 4. Sequence processing model(s) 4 can include one or more machine-learned components. Sequence processing model(s) 4 can process the data from input(s) 2 to obtain an input sequence 5. Input sequence 5 can include one or more input elements 5-1, 5-2, . . . , 5-M, etc. obtained from input(s) 2. Sequence processing model 4 can process input sequence 5 using prediction layer(s) 6 to generate an output sequence 7. Output sequence 7 can include one or more output elements 7-1, 7-2, . . . , 7-N, etc. generated based on input sequence 5. The system can generate output(s) 3 based on output sequence 7.

[0179] Sequence processing model(s) 4 can include one or multiple machine-learned model components configured to ingest, generate, or otherwise reason over sequences of information. For example, some example sequence processing models in the text domain are referred to as “Large Language Models,” or LLMs. See, e.g, PaLM 2 Technical Report, GOOGLE, https: / / ai.google / static / documents / palm2techreport.pdf (n d ). Other example sequence processing models can operate in other domains, such as image domains, see, e.g., Dosovitskiy et al., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, ARXIV:2010. 11929v2 (Jun. 3, 2021), audio domains, see, e.g., Agostinelli et al., MusicLM: Generating Music From Text, ARXIV:2301.11325V1 (Jan. 26, 2023), biochemical domains, see. e.g., Jumper et al.. Highly accurate protein structure prediction with AlphaFold, 596 Nature 583 (Aug. 26, 2021). by way of example. Sequence processing model(s) 4 can process one or multiple types of data simultaneously. Sequence processing model(s) 4 can include relatively large models (e.g., more parameters, computationally expensive, etc.), relatively small models (e.g., fewer parameters, computationally lightweight, etc ), or both.

[0180] In general, sequence processing model(s) 4 can obtain input sequence 5 using data from input(s) 2. For instance, input sequence 5 can include a representation of data from input(s) 2 in a format understood by sequence processing model(s) 4. One or more machine- learned components of sequence processing model(s) 4 can ingest the data from input(s) 2,parse the data into pieces compatible with the processing architectures of sequence processing model(s) 4 (e.g., via “tokenization”). and project the pieces into an input space associated with prediction layer(s) 6 (e.g., via '‘embedding”).

[0181] Sequence processing model(s) 4 can ingest the data from input(s) 2 and parse the data into a sequence of elements to obtain input sequence 5. For example, a portion of input data from input(s) 2 can be broken down into pieces that collectively represent the content of the portion of the input data. The pieces can provide the elements of the sequence.

[0182] Elements 5-1, 5-2, . . . , 5-M can represent, in some cases, building blocks for capturing or expressing meaningful information in a particular data domain. For instance, the elements can describe “atomic units” across one or more domains. For example, for textual input source(s), the elements can correspond to groups of one or more words or sub-word components, such as sets of one or more characters.

[0183] For example, elements 5-1, 5-2, . . . , 5-M can represent tokens obtained using a tokenizer. For instance, a tokenizer can process a given portion of an input source and output a series of tokens (e.g.. corresponding to input elements 5-1, 5-2, . . . , 5-M) that represent the portion of the input source. Various approaches to tokenization can be used. For instance, textual input source(s) can be tokenized using a byte-pair encoding (BPE) technique. See, e.g., Kudo et al., SentencePiece: A simple and language independent subword tokenizer arid detokenizer for Neural Text Processing, PROCEEDINGS OF THE 2018 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (System Demonstrations), pages 66-71 (October 31-November 4, 2018), https: / / aclanthology.org / D18-2012.pdf. Imagebased input source(s) can be tokenized by extracting and serializing patches from an image.

[0184] In general, arbitrary data types can be serialized and processed into input sequence 5. It is to be understood that element(s) 5-1, 5-2, . . . , 5-M depicted in FIG. 16 can be the tokens or can be the embedded representations thereof.

[0185] Prediction layer(s) 6 can predict one or more output elements 7-1, 7-2, . . . , 7 - V based on the input elements. Prediction layer(s) 6 can include one or more machine-learned model architectures, such as one or more layers of learned parameters that manipulate and transform the input(s) to extract higher-order meaning from, and relationships between, input element(s) 5-1, 5-2, . . . , 5-M. In this manner, for instance, example prediction layer(s) 6 can predict new output element(s) in view of the context provided by input sequence 5.

[0186] Prediction layer(s) 6 can evaluate associations between portions of input sequence 5 and a particular output element. These associations can inform a prediction of the likelihood that a particular output follows the input context. For example, consider the textual snippet, “The carpenter’s toolbox was small and heavy. It was full of .” Example prediction layer(s) 6 can identify that “It” refers back to “toolbox” by determining a relationship between the respective embeddings. Example prediction layer(s) 6 can also link “It” to the attributes of the toolbox, such as “small” and “heavy.” Based on these associations, prediction layer(s) 6 can, for instance, assign a higher probability to the word “nails” than to the word “sawdust.”

[0187] A transformer is an example architecture that can be used in prediction layer(s) 4. See, e.g., Vaswani et al., Attention Is All You Need, ARXIV: 1706.03762V7 (Aug. 2, 2023). A transformer is an example of a machine-learned model architecture that uses an attention mechanism to compute associations between items within a context window. The context window can include a sequence that contains input sequence 5 and potentially one or more output element(s) 7-1, 7-2, . . . , 7-N. A transformer block can include one or more attention layer(s) and one or more post-attention layer(s) (e g., feedforward layer(s), such as a multilayer perceptron).

[0188] Prediction layer(s) 6 can include other machine-learned model architectures in addition to or in lieu of transformer-based architectures. For example, recurrent neural networks (RNNs) and long short-term memory (LSTM) models can also be used, as well as convolutional neural networks (CNNs). In general, prediction layer(s) 6 can leverage various kinds of artificial neural networks that can understand or generate sequences of information.

[0189] Output sequence 7 can include or otherw ise represent the same or different data types as input sequence 5. For instance, input sequence 5 can represent textual data, and output sequence 7 can represent textual data. Input sequence 5 can represent image, audio, or audiovisual data, and output sequence 7 can represent textual data (e.g., describing the image, audio, or audiovisual data). It is to be understood that prediction layer(s) 6, and any other interstitial model components of sequence processing model(s) 4. can be configured to receive a variety of data types in input sequence(s) 5 and output a variety of data types in output sequence(s) 7.

[0190] Output sequence 7 can have various relationships to input sequence 5. Output sequence 7 can be a continuation of input sequence 5. Output sequence 7 can becomplementary to input sequence 5. Output sequence 7 can translate, transform, augment, or otherwise modify input sequence 5. Output sequence 7 can answer, evaluate, confirm, or otherwise respond to input sequence 5. Output sequence 7 can implement (or describe instructions for implementing) an instruction provided via input sequence 5.

[0191] Output sequence 7 can be generated autoregressively. For instance, for some applications, an output of one or more prediction layer(s) 6 can be passed through one or more output layers (e.g., softmax layer) to obtain a probability distribution over an output vocabulary (e.g.. a textual or symbolic vocabulary) conditioned on a set of input elements in a context window. In this manner, for instance, output sequence 7 can be autoregressively generated by sampling a likely next output element, adding that element to the context window, and re-generating the probability distribution based on the updated context window, and sampling a likely next output element, and so forth.

[0192] Output sequence 7 can also be generated non-autoregressively. For instance, multiple output elements of output sequence 7 can be predicted together without explicit sequential conditioning on each other. See, e.g., Saharia et al., Non- Autoregressive Machine Translation with Latent Alignments, ARXlV:2004.07437v3 (NOV. 1 , 2020).

[0193] Output sequence 7 can include one or multiple portions or elements. In an example content generation configuration, output sequence 7 can include multiple elements corresponding to multiple portions of a generated output sequence (e.g., a textual sentence, values of a discretized waveform, computer code, etc.). In an example classification configuration, output sequence 7 can include a single element associated with a classification output. For instance, an output ‘’vocabulary” can include a set of classes into which an input sequence is to be classified. For instance, a vision transformer block can pass latent state information to a multilayer perceptron that outputs a likely class value associated with an input image.

[0194] FIG. 14 is a block diagram of an example technique for populating an example input sequence 8. Input sequence 8 can include various functional elements that form part of the model infrastructure, such as an element 8-0 obtained from a task indicator 9 that signals to any model(s) that process input sequence 8 that a particular task is being performed (e.g., to help adapt a performance of the model(s) to that particular task). Input sequence 8 can include various data elements from different data modalities. For instance, an input modality 10-1 can include one modality of data. A data-to-sequence model 11-1 can process data frominput modality 10-1 to project the data into a format compatible with input sequence 8 (e.g., one or more vectors dimensioned according to the dimensions of input sequence 8) to obtain elements 8-1, 8-2, 8-3. Another input modality 10-2 can include a different modality of data. A data-to-sequence model 11-2 can project data from input modality 10-2 into a format compatible with input sequence 8 to obtain elements 8-4, 8-5, 8-6. Another input modality 10-3 can include yet another different modality of data. A data-to-sequence model 11-3 can project data from input modality 10-3 into a format compatible with input sequence 8 to obtain elements 8-7, 8-8, 8-9.

[0195] Input sequence 8 can be the same as or different from input sequence 5. Input sequence 8 can be a multimodal input sequence that contains elements that represent data from different modalities using a common dimensional representation. For instance, an embedding space can have P dimensions. Input sequence 8 can be configured to contain a plurality of elements that have P dimensions. In this manner, for instance, example implementations can facilitate information extraction and reasoning across diverse data modalities by projecting data into elements in the same embedding space for comparison, combination, or other computations therebetween.

[0196] For example, elements 8-0, . . . , 8-9 can indicate particular locations within a multidimensional embedding space. Some elements can map to a set of discrete locations in the embedding space. For instance, elements that correspond to discrete members of a predetermined vocabulary of tokens can map to discrete locations in the embedding space that are associated with those tokens. Other elements can be continuously distributed across the embedding space. For instance, some data types can be broken down into continuously defined portions (e.g., image patches) that can be described using continuously distributed locations within the embedding space.

[0197] In some implementations, the expressive power of the embedding space may not be limited to meanings associated with any particular set of tokens or other building blocks. For example, a continuous embedding space can encode a spectrum of high-order information. An individual piece of information (e.g., a token) can map to a particular point in that space: for instance, a token for the word “dog” can be projected to an embedded value that points to a particular location in the embedding space associated with canine-related information. Similarly, an image patch of an image of a dog on grass can also be projected into the embedding space. In some implementations, the projection of the image of the dog can be similar to the projection of the word “dog” while also having similarity to a projection of theword “grass,"’ while potentially being different from both. In some implementations, the projection of the image patch may not exactly align with any single projection of a single word. In some implementations, the projection of the image patch can align with a combination of the projections of the words “dog” and “grass.” In this manner, for instance, a high-order embedding space can encode information that can be independent of data modalities in which the information is expressed.

[0198] Task indicator 9 can include a model or model component configured to identify a task being performed and inject, into input sequence 8, an input value represented by element 8-0 that signals which task is being performed. For instance, the input value can be provided as a data type associated with an input modality and projected along with that input modality7(e.g., the input value can be a textual task label that is embedded along with other textual data in the input; the input value can be a pixel-based representation of a task that is embedded along with other image data in the input; etc.). The input value can be provided as a data type that differs from or is at least independent from other input(s). For instance, the input value represented by element 8-0 can be a learned within a continuous embedding space.

[0199] Input modalities 10-1, 10-2, and 10-3 can be associated with various different data types (e.g., as described above with respect to input(s) 2 and output(s) 3).

[0200] Data-to-sequence models 11-1, 11-2, and 11-3 can be the same or different from each other. Data-to-sequence models 11-1, 11-2, and 11-3 can be adapted to each respective input modality 10-1, 10-2, and 10-3. For example, a textual data-to-sequence model can subdivide a portion of input text and project the subdivisions into element(s) in input sequence 8 (e.g., elements 8-1, 8-2, 8-3, etc.). An image data-to-sequence model can subdivide an input image and project the subdivisions into element(s) in input sequence 8 (e.g., elements 8-4, 8-5, 8-6, etc.). An arbitrary datatype data-to-sequence model can subdivide an input of that arbitrary datatype and project the subdivisions into element(s) in input sequence 8 (e.g.. elements 8-7. 8-8, 8-9, etc ).

[0201] Data-to-sequence models 11-1, 11-2, and 11-3 can form part of machine-learned sequence processing model(s) 4. Data-to-sequence models 11-1, 11-2, and 11-3 can be jointly trained with or trained independently from machine-learned sequence processing model(s) 4. Data-to-sequence models 11-1, 11-2, and 11-3 can be trained end-to-end with machine-learned sequence processing model(s) 4.Example Machine-Learned Model Development Platform

[0202] FIG. 15 is a block diagram of an example model development platform 12 that can facilitate creation, adaptation, and refinement of example machine-learned models (e.g., machine-learned model(s) 1. sequence processing model(s) 4, etc.). Model development platform 12 can provide a number of different toolkits that developer systems can employ in the development of new or adapted machine-learned models.

[0203] Model development platform 12 can provide one or more model libraries 13 containing building blocks for new models. Model libraries 13 can include one or more pretrained foundational models 13-1, which can provide a backbone of processing power across various tasks. Model libraries 13 can include one or more pre-trained expert models 13-2. which can be focused on performance in particular domains of expertise. Model libraries 13 can include various model primitives 13-3, which can provide low-level architectures or components (optionally pre-trained), which can be assembled in various arrangements as desired.

[0204] Model development platform 12 can receive selections of various model components 14. Model development platform 12 can pass selected model components 14 to a workbench15 that combines selected model components 14 into a development model 16.

[0205] Workbench 15 can facilitate further refinement and adaptation of development model16 by leveraging a number of different toolkits integrated with model development platform 12. For example, workbench 15 can facilitate alignment of the development model 16 with a desired performance profile on various tasks using a model alignment toolkit 17.

[0206] Model alignment toolkit 17 can provide a number of tools for causing development model 16 to generate outputs aligned with desired behavioral characteristics. Alignment can include increasing an accuracy, precision, recall, etc. of model outputs. Alignment can include enforcing output styles, schema, or other preferential characteristics of model outputs. Alignment can be general or domain-specific. For instance, a pre-trained foundational model 13-1 can begin with an initial level of performance across multiple domains. Alignment of the pre-trained foundational model 13-1 can include improving a performance in a particular domain of information or tasks (e.g., even at the expense of performance in another domain of information or tasks).

[0207] Model alignment toolkit 17 can integrate one or more dataset(s) 17-1 for aligning development model 16. Curated dataset(s) 17-1 can include labeled or unlabeled trainingdata. Dataset(s) 17-1 can be obtained from public domain datasets. Dataset(s) 17-1 can be obtained from private datasets associated with one or more developer system(s) for the alignment of bespoke machine-learned model(s) customized for private use-cases.

[0208] Pre-training pipelines 17-2 can include a machine-learned model training workflow configured to update development model 16 over large-scale, potentially noisy datasets. For example, pre-training can leverage unsupervised learning techniques (e.g., de-noising, etc.) to process large numbers of training instances to update model parameters from an initialized state and achieve a desired baseline performance. Pre-training pipelines 17-2 can leverage unlabeled datasets in dataset(s) 17-1 to perform pre-training. Workbench 15 can implement a pre-training pipeline 17-2 to pre-train development model 16.

[0209] Fine-tuning pipelines 17-3 can include a machine-learned model training workflow configured to refine the model parameters of development model 16 with higher-quality data. Fine-tuning pipelines 17-3 can update development model 16 by conducting supervised training with labeled dataset(s) in dataset(s) 17-1. Fine-tuning pipelines 17-3 can update development model 16 by conducting reinforcement learning using reward signals from user feedback signals. Workbench 15 can implement a fine-tuning pipeline 17-3 to fine-tune development model 16.

[0210] Prompt libraries 17-4 can include sets of inputs configured to induce behavior aligned with desired performance criteria. Prompt libraries 17-4 can include few-shot prompts (e.g., inputs providing examples of desired model outputs for prepending to a desired runtime query), chain-of-thought prompts (e.g., inputs providing step-by-step reasoning within the exemplars to facilitate thorough reasoning by the model), and the like.

[0211] Example prompts can be retrieved from an available repository of prompt libraries 17- 4. Example prompts can be contributed by one or more developer systems using workbench 15.

[0212] In some implementations, pre-trained or fine-tuned models can achieve satisfactory performance without exemplars in the inputs. For instance, zero-shot prompts can include inputs that lack exemplars. Zero-shot prompts can be within a domain within a training dataset or outside of the training domain(s).

[0213] Prompt libraries 17-4 can include one or more prompt engineering tools. Prompt engineering tools can provide workflows for retrieving or learning optimized prompt values. Prompt engineering tools can facilitate directly learning prompt values (e.g., input elementvalues) based one or more training iterations. Workbench 15 can implement prompt engineering tools in development model 16.

[0214] Prompt libraries 17-4 can include pipelines for prompt generation. For example, inputs can be generated using development model 16 itself or other machine-learned models. In this manner, for instance, a first model can process information about a task and output an input for a second model to process in order to perform a step of the task. The second model can be the same as or different from the first model. Workbench 15 can implement prompt generation pipelines in development model 16.

[0215] Prompt libraries 17-4 can include pipelines for context injection. For instance, a performance of development model 16 on a particular task can improve if provided with additional context for performing the task. Prompt libraries 17-4 can include software components configured to identify desired context, retrieve the context from an external source (e.g., a database, a sensor, etc.), and add the context to the input prompt. Workbench 15 can implement context injection pipelines in development model 16.

[0216] Although various training examples described herein with respect to model development platform 12 refer to ‘"pre-training” and “fine-tuning.” it is to be understood that model alignment toolkit 17 can generally support a wide variety of training techniques adapted for training a wide variety of machine-learned models. Example training techniques can correspond to the example training method 6000 described above.

[0217] Model development platform 12 can include a model plugin toolkit 18. Model plugin toolkit 18 can include a variety of tools configured for augmenting the functionality of a machine-learned model by integrating the machine-learned model with other systems, devices, and software components. For instance, a machine-learned model can use tools to increase performance qualify where appropriate. For instance, deterministic tasks can be offloaded to dedicated tools in lieu of probabilistically performing the task with an increased risk of error. For instance, instead of autoregressively predicting the solution to a system of equations, a machine-learned model can recognize a tool to call for obtaining the solution and pass the system of equations to the appropriate tool. The tool can be a traditional system of equations solver that can operate deterministically to resolve the system of equations. The output of the tool can be returned in response to the original query. In this manner, tool use can allow some example models to focus on the strengths of machine-learned models — e.g.. understanding an intent in an unstructured request for a task — while augmenting theperformance of the model by offloading certain tasks to a more focused tool for rote application of deterministic algorithms to a well-defined problem.

[0218] Model plugin toolkit 18 can include validation tools 18-1. Validation tools 18-1 can include tools that can parse and confirm output(s) of a machine-learned model. Validation tools 18-1 can include engineered heuristics that establish certain thresholds applied to model outputs. For example, validation tools 18-1 can ground the outputs of machine-learned models to structured data sources (e.g., to mitigate ‘‘hallucinations’').

[0219] Model plugin toolkit 18 can include tooling packages 18-2 for implementing one or more tools that can include scripts or other executable code that can be executed alongside development model 16. Tooling packages 18-2 can include one or more inputs configured to cause machine-learned model(s) to implement the tools (e.g., few-shot prompts that induce a model to output tool calls in the proper syntax, etc.). Tooling packages 18-2 can include, for instance, fine-tuning training data for training a model to use a tool.

[0220] Model plugin toolkit 18 can include interfaces for calling external application programming interfaces (APIs) 18-3. For instance, in addition to or in lieu of implementing tool calls or tool code directly with development model 16, development model 16 can be aligned to output instruction that initiate API calls to send or obtain data via external systems.

[0221] Model plugin toolkit 18 can integrate with prompt libraries 17-4 to build a catalog of available tools for use with development model 16. For instance, a model can receive, in an input, a catalog of available tools, and the model can generate an output that selects a tool from the available tools and initiates a tool call for using the tool.

[0222] Model development platform 12 can include a computational optimization toolkit 19 for optimizing a computational performance of development model 16. For instance, tools for model compression 19-1 can allow development model 16 to be reduced in size while maintaining a desired level of performance. For instance, model compression 19-1 can include quantization workflows, weight pruning and sparsification techniques, etc. Tools for hardware acceleration 19-2 can facilitate the configuration of the model storage and execution formats to operate optimally on different hardware resources. For instance, hardware acceleration 19-2 can include tools for optimally sharding models for distributed processing over multiple processing units for increased bandwidth, lower unified memory requirements, etc. Tools for distillation 19-3 can provide for the training of lighter- weight models based on the knowledge encoded in development model 16. For instance,development model 16 can be a highly performant, large machine-learned model optimized using model development platform 12. To obtain a lightweight model for running in resource-constrained environments, a smaller model can be a '‘student model’’ that learns to imitate development model 16 as a “teacher model.” In this manner, for instance, the investment in learning the parameters and configurations of development model 16 can be efficiently transferred to a smaller model for more efficient inference.

[0223] Workbench 15 can implement one, multiple, or none of the toolkits implemented in model development platform 12. Workbench 15 can output an output model 20 based on development model 16. Output model 20 can be a deployment version of development model 16. Output model 20 can be a development or training checkpoint of development model 16. Output model 20 can be a distilled, compressed, or otherwise optimized version of development model 16.

[0224] FIG. 16 is a block diagram of an example training flow for training a machine-learned development model 16. One or more portion(s) of the example training flow can be implemented by a computing system that includes one or more computing devices such as, for example, computing systems described with reference to the other drawings. Each respective portion of the example training flow can be performed by any (or any combination) of one or more computing devices. Moreover, one or more portion(s) of the example training flow can be implemented on the hardware components of the device(s) described herein, for example, to train one or more systems or models. FIG. 16 depicts elements performed in a particular order for purposes of illustration and discussion. Those of ordinary skill in the art, using the disclosures provided herein, will understand that the elements of any of the methods discussed herein can be adapted, rearranged, expanded, omitted, combined, or modified in various ways without deviating from the scope of the disclosure. FIG. 16 is described with reference to elements / terms described with respect to other systems and drawings for exemplary illustrated purposes and is not meant to be limiting. One or more portions of the example training flow can be performed additionally, or alternatively, by other systems.

[0225] Initially, development model 16 can persist in an initial state as an initialized model 21. Development model 16 can be initialized with weight values. Initial weight values can be random or based on an initialization schema. Initial weight values can be based on prior pre-training for the same or for a different model.

[0226] Initialized model 21 can undergo pre-training in a pre-training stage 22. Pre-training stage 22 can be implemented using one or more pre-training pipelines 17-2 over data from dataset(s) 17-1. Pre-training can be omitted, for example, if initialized model 21 is already pre-trained (e.g., development model 16 contains, is, or is based on a pre-trained foundational model or an expert model).

[0227] Pre-trained model 23 can then be a new version of development model 16, which can persist as development model 16 or as anew development model. Pre-trained model 23 can be the initial state if development model 16 was already pre-trained. Pre-trained model 23 can undergo fine-tuning in a fine-tuning stage 24. Fine-tuning stage 24 can be implemented using one or more fine-tuning pipelines 17-3 over data from dataset(s) 17-1. Fine-tuning can be omitted, for example, if a pre-trained model as satisfactory performance, if the model was already fine-tuned, or if other tuning approaches are preferred.

[0228] Fine-tuned model 29 can then be a new version of development model 16, which can persist as development model 16 or as anew development model. Fine-tuned model 29 can be the initial state if development model 16 was already fine-tuned. Fine-tuned model 29 can undergo refinement with user feedback 26. For instance, refinement with user feedback 26 can include reinforcement learning, optionally based on human feedback from human users of fine-tuned model 25. As reinforcement learning can be a form of fine-tuning, it is to be understood that fine-tuning stage 24 can subsume the stage for refining with user feedback 26. Refinement with user feedback 26 can produce a refined model 27. Refined model 27 can be output to downstream system(s) 28 for deployment or further development.

[0229] In some implementations, computational optimization operations can be applied before, during, or after each stage. For instance, initialized model 21 can undergo computational optimization 29-1 (e.g.. using computational optimization toolkit 19) before pre-training stage 22. Pre-trained model 23 can undergo computational optimization 29-2 (e g., using computational optimization toolkit 19) before fine-tuning stage 24. Fine-tuned model 25 can undergo computational optimization 29-3 (e.g., using computational optimization toolkit 19) before refinement with user feedback 26. Refined model 27 can undergo computational optimization 29-4 (e.g., using computational optimization toolkit 19) before output to downstream system(s) 28. Computational optimization(s) 29-1, . . . , 29-4 can all be the same, all be different, or include at least some different optimization techniques.Example Machine-Learned Model Inference System

[0230] FIG. 17 is a block diagram of an inference system for operating one or more machine- learned model(s) 1 to perform inference (e.g., for training, for deployment, etc.). A model host 31 can receive machine-learned model(s) 1. Model host 31 can host one or more model instance(s) 31 -1, which can be one or multiple instances of one or multiple models. Model host 31 can host model instance(s) 31-1 using available compute resources 31-2 associated with model host 31.

[0231] Model host 31 can perform inference on behalf of one or more client(s) 32. Client(s) 32 can transmit an input request 33 to model host 31. Using input request 33, model host 31 can obtain input(s) 2 for input to machine-learned model(s) 1. Machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3. Using output(s) 3, model host 31 can return an output payload 34 for responding to input request 33 from client(s) 32. Output payload 34 can include or be based on output(s) 3.

[0232] Model host 31 can leverage various other resources and tools to augment the inference task. For instance, model host 31 can communicate with tool interfaces 35 to facilitate tool use by model instance(s) 31-1. Tool interfaces 35 can include local or remote APIs. Tool interfaces 35 can include integrated scripts or other software functionality.Model host 31 can engage online learning interface(s) 36 to facilitate ongoing improvements to machine-learned model(s) 1. For instance, online learning interface(s) 36 can be used within reinforcement learning loops to retrieve user feedback on inferences served by model host 31. Model host 31 can access runtime data source(s) 37 for augmenting input(s) 2 with additional contextual information. For instance, runtime data source(s) 37 can include a knowledge graph 37-1 that facilitates structured information retrieval for information associated with input request(s) 33 (e.g., a search engine service). Runtime data source(s) 37 can include public or private, external or local database(s) 37-2 that can store information associated with input request(s) 33 for augmenting input(s) 2. Runtime data source(s) 37 can include account data 37-3 which can be retrieved in association with a user account corresponding to a client 32 for customizing the behavior of model host 31 accordingly.

[0233] Model host 31 can be implemented by one or multiple computing devices or systems. Client(s) 2 can be implemented by one or multiple computing devices or systems, which can include computing devices or systems shared with model host 31.

[0234] For example, model host 31 can operate on a server system that provides a machinelearning service to client device(s) that operate client(s) 32 (e.g., over a local or wide-area network). Client device(s) can be end-user devices used by individuals. Client device(s) can be server systems that operate client(s) 32 to provide various functionality as a service to downstream end-user devices.

[0235] In some implementations, model host 31 can operate on a same device or system as client(s) 32. Model host 31 can be a machine-learning service that runs on-device to provide machine-learning functionality to one or multiple applications operating on a client device, w hich can include an application implementing client(s) 32. Model host 31 can be a part of a same application as client(s) 32. For instance, model host 31 can be a subroutine or method implemented by one part of an application, and client(s) 32 can be another subroutine or method that engages model host 31 to perform inference functions within the application. It is to be understood that model host 31 and chent(s) 32 can have various different configurations.

[0236] Model instan ce(s) 31-1 can include one or more machine-learned models that are available for performing inference. Model instance(s) 31-1 can include weights or other model components that are stored in persistent storage, temporarily cached, or loaded into high-speed memory. Model instance(s) 31-1 can include multiple instance(s) of the same model (e.g., for parallel execution of more requests on the same model). Model instance(s) 31-1 can include instance(s) of different model(s). Model instance(s) 31-1 can include cached intermediate states of active or inactive model(s) used to accelerate inference of those models. For instance, an inference session with a particular model may generate significant amounts of computational results that can be re-used for future inference runs (e.g., using a KV cache for transformer-based models). These computational results can be saved in association with that inference session so that session can be executed more efficiently w hen resumed.

[0237] Compute resource(s) 31-2 can include one or more processors (central processing units, graphical processing units, tensor processing units, machine-learning accelerators, etc.) connected to one or more memory devices. Compute resource(s) 31-2 can include a dynamic pool of available resources shared with other processes. Compute resource(s) 31-2 can include memory devices large enough to fit an entire model instance in a single memory instance. Compute resource(s) 31-2 can also shard model instance(s) across multiple memory devices (e.g., using data parallelization or tensor parallelization, etc.). This can bedone to increase parallelization or to execute a large model using multiple memory devices which individually might not be able to fit the entire model into memory.

[0238] Input request 33 can include data for input(s) 2. Model host 31 can process input request 33 to obtain input(s) 2. Input(s) 2 can be obtained directly from input request 33 or can be retrieved using input request 33. Input request 33 can be submitted to model host 31 via an API.

[0239] Model host 31 can perform inference over batches of input requests 33 in parallel. For instance, a model instance 31-1 can be configured with an input structure that has a batch dimension. Separate input(s) 2 can be distributed across the batch dimension (e.g.. rows of an array). The separate input(s) 2 can include completely different contexts. The separate input(s) 2 can be multiple inference steps of the same task. The separate input(s) 2 can be staggered in an input structure, such that any given inference cycle can be operating on different portions of the respective input(s) 2. In this manner, for instance, model host 31 can perform inference on the batch in parallel, such that output(s) 3 can also contain the batch dimension and return the inference results for the batched input(s) 2 in parallel. In this manner, for instance, batches of input request(s) 33 can be processed in parallel for higher throughput of output payload(s) 34.

[0240] Output payload 34 can include or be based on output(s) 3 from machine-learned model(s) 1. Model host 31 can process output(s) 3 to obtain output payload 34. This can include chaining multiple rounds of inference (e.g., iteratively, recursively, across the same model(s) or different model(s)) to arrive at a final output for a task to be returned in output payload 34. Output payload 34 can be transmitted to client(s) 32 via an API.

[0241] Online learning interface(s) 36 can facilitate reinforcement learning of machine- learned model(s) 1. Online learning interface(s) 36 can facilitate reinforcement learning with human feedback (RTHF). Online learning interface(s) 36 can facilitate federated learning of machine-learned model(s) 1.

[0242] Model host 31 can execute machine-learned model(s) 1 to perform inference for various tasks using various types of data. For example, various different input(s) 2 and output(s) 3 can be used for various different tasks. In some implementations, input(s) 2 can be or otherwise represent image data. Machine-learned model(s) 1 can process the image data to generate an output. As an example, machine-learned model(s) 1 can process the image data to generate an image recognition output (e.g., a recognition of the image data, alatent embedding of the image data, an encoded representation of the image data, a hash of the image data. etc.). As another example, machine-learned model(s) 1 can process the image data to generate an image segmentation output. As another example, machine-learned model(s) 1 can process the image data to generate an image classification output. As another example, machine-learned model(s) 1 can process the image data to generate an image data modification output (e.g., an alteration of the image data, etc.). As another example, machine-learned model(s) 1 can process the image data to generate an encoded image data output (e g., an encoded and / or compressed representation of the image data, etc.). As another example, machine-learned model(s) 1 can process the image data to generate an upscaled image data output. As another example, machine-learned model(s) 1 can process the image data to generate a prediction output.

[0243] In some implementations, the task is a computer vision task. In some cases, input(s) 2 includes pixel data for one or more images and the task is an image processing task. For example, the image processing task can be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the likelihood that the one or more images depict an object belonging to the object class. The image processing task may be object detection, where the image processing output identifies one or more regions in the one or more images and, for each region, a likelihood that region depicts an object of interest. As another example, the image processing task can be image segmentation, where the image processing output defines, for each pixel in the one or more images, a respective likelihood for each category in a predetermined set of categories. For example, the set of categories can be foreground and background. As another example, the set of categories can be object classes. As another example, the image processing task can be depth estimation, where the image processing output defines, for each pixel in the one or more images, a respective depth value. As another example, the image processing task can be motion estimation, where the network input includes multiple images, and the image processing output defines, for each pixel of one of the input images, a motion of the scene depicted at the pixel between the images in the network input.

[0244] In some implementations, input(s) 2 can be or otherwise represent natural language data. Machine-learned model(s) 1 can process the natural language data to generate an output. As an example, machine-learned model(s) 1 can process the natural language data to generate a language encoding output. As another example, machine-learned model(s) 1 can process the natural language data to generate a latent text embedding output. As anotherexample, machine-learned model(s) 1 can process the natural language data to generate a translation output. As another example, machine-learned model(s) 1 can process the natural language data to generate a classification output. As another example, machine-learned model(s) 1 can process the natural language data to generate a textual segmentation output. As another example, machine-learned model(s) 1 can process the natural language data to generate a semantic intent output. As another example, machine-learned model(s) 1 can process the natural language data to generate an upscaled text or natural language output (e.g., text or natural language data that is higher quality than the input text or natural language, etc.). As another example, machine-learned model (s) 1 can process the natural language data to generate a prediction output (e.g., one or more predicted next portions of natural language content).

[0245] In some implementations, input(s) 2 can be or otherwise represent speech data (e.g., data describing spoken natural language, such as audio data, textual data. etc.). Machine- learned model(s) 1 can process the speech data to generate an output. As an example, machine-learned model(s) 1 can process the speech data to generate a speech recognition output. As another example, machine-learned model(s) 1 can process the speech data to generate a speech translation output. As another example, machine-learned model(s) 1 can process the speech data to generate a latent embedding output. As another example, machine-learned model(s) 1 can process the speech data to generate an encoded speech output (e.g., an encoded and / or compressed representation of the speech data, etc.). As another example, machine-learned model(s) 1 can process the speech data to generate an upscaled speech output (e.g.. speech data that is higher quality than the input speech data, etc.). As another example, machine-learned model(s) 1 can process the speech data to generate a textual representation output (e.g., a textual representation of the input speech data, etc.). As another example, machine-learned model(s) 1 can process the speech data to generate a prediction output.

[0246] In some implementations, input(s) 2 can be or otherwise represent latent encoding data (e.g., a latent space representation of an input, etc.). Machine-learned model(s) 1 can process the latent encoding data to generate an output. As an example, machine-learned model(s) 1 can process the latent encoding data to generate a recognition output. As another example, machine-learned model(s) 1 can process the latent encoding data to generate a reconstruction output. As another example, machine-learned model(s) 1 can process the latent encoding data to generate a search output. As another example, machine-learnedmodel(s) 1 can process the latent encoding data to generate a reclustering output. As another example, machine-learned model(s) 1 can process the latent encoding data to generate a prediction output.

[0247] In some implementations, input(s) 2 can be or otherwise represent statistical data. Statistical data can be, represent, or otherwise include data computed and / or calculated from some other data source. Machine-learned model(s) 1 can process the statistical data to generate an output. As an example, machine-learned model(s) 1 can process the statistical data to generate a recognition output. As another example, machine-learned model(s) 1 can process the statistical data to generate a prediction output. As another example, machine- learned model(s) 1 can process the statistical data to generate a classification output. As another example, machine-learned model(s) 1 can process the statistical data to generate a segmentation output. As another example, machine-learned model(s) 1 can process the statistical data to generate a visualization output. As another example, machine-learned model(s) 1 can process the statistical data to generate a diagnostic output.

[0248] In some implementations, input(s) 2 can be or otherwise represent sensor data. Machine-learned model(s) 1 can process the sensor data to generate an output. As an example, machine-learned model(s) 1 can process the sensor data to generate a recognition output. As another example, machine-learned model(s) 1 can process the sensor data to generate a prediction output. As another example, machine-learned model(s) 1 can process the sensor data to generate a classification output. As another example, machine-learned model(s) 1 can process the sensor data to generate a segmentation output. As another example, machine-learned model(s) 1 can process the sensor data to generate a visualization output. As another example, machine-learned model(s) 1 can process the sensor data to generate a diagnostic output. As another example, machine-learned model(s) 1 can process the sensor data to generate a detection output.

[0249] In some implementations, machine-learned model(s) 1 can be configured to perform a task that includes encoding input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task may be an audio compression task. The input may include audio data and the output may comprise compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output comprises compressed visual data, and the task is a visual data compression task. In another example, the task may comprise generating an embedding for input data (e.g., input audio or visual data). In some cases, the input includes audio data representing a spokenutterance and the task is a speech recognition task. The output may comprise a text output which is mapped to the spoken utterance. In some cases, the task comprises encrypting or decrypting input data. In some cases, the task comprises a microprocessor performance task, such as branch prediction or memory address translation.

[0250] In some implementations, the task is a generative task, and machine-learned model(s) 1 can be configured to output content generated in view of input(s) 2. For instance, input(s) 2 can be or otherwise represent data of one or more modalities that encodes context for generating additional content.

[0251] In some implementations, the task can be a text completion task. Machine-learned model(s) 1 can be configured to process input(s) 2 that represent textual data and to generate output(s) 3 that represent additional textual data that completes a textual sequence that includes input(s) 2. For instance, machine-learned model(s) 1 can be configured to generate output(s) 3 to complete a sentence, paragraph, or portion of text that follows from a portion of text represented by input(s) 2.

[0252] In some implementations, the task can be an instruction following task. Machine- learned model(s) 1 can be configured to process input(s) 2 that represent instructions to perform a function and to generate output(s) 3 that advance a goal of satisfying the instruction function (e.g., at least a step of a multi-step procedure to perform the function). Output(s) 3 can represent data of the same or of a different modality as input(s) 2. For instance, input(s) 2 can represent textual data (e.g., natural language instructions for a task to be performed) and machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the instructions (e.g., natural language responses, programming language responses, machine language responses, etc.). Input(s) 2 can represent image data (e.g., image-based instructions for a task to be performed, optionally accompanied by textual instructions) and machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the instructions (e.g., natural language responses, programming language responses, machine language responses, etc.). One or more output(s) 3 can be iteratively or recursively generated to sequentially process and accomplish steps toward accomplishing the requested functionality. For instance, an initial output can be executed by an external system or be processed by machine-learned model(s) 1 to complete an initial step of performing a function. Multiple steps can be performed, with a final output being obtained that is responsive to the initial instructions.

[0253] In some implementations, the task can be a question answering task. Machine-learned model(s) 1 can be configured to process input(s) 2 that represent a question to answer and to generate output(s) 3 that advance a goal of returning an answer to the question (e.g., at least a step of a multi-step procedure to perform the function). Output(s) 3 can represent data of the same or of a different modality as input(s) 2. For instance, input(s) 2 can represent textual data (e.g., natural language instructions for a task to be performed) and machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the question (e.g., natural language responses, programming language responses, machine language responses, etc.). Input(s) 2 can represent image data (e.g., image-based instructions for a task to be performed, optionally accompanied by textual instructions) and machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the question (e.g., natural language responses, programming language responses, machine language responses, etc.). One or more output(s) 3 can be iteratively or recursively generated to sequentially process and accomplish steps toward answering the question. For instance, an initial output can be executed by an external system or be processed by machine-learned model(s) 1 to complete an initial step of obtaining an answer to the question (e g., querying a database, performing a computation, executing a script, etc.). Multiple steps can be performed, with a final output being obtained that is responsive to the question.

[0254] In some implementations, the task can be an image generation task. Machine-learned model(s) 1 can be configured to process input(s) 2 that represent context regarding a desired portion of image content. The context can include text data, image data, audio data, etc. Machine-learned model(s) 1 can be configured to generate output(s) 3 that represent image data that depicts imagery' related to the context. For instance, machine-learned model(s) 1 can be configured to generate pixel data of an image. Values for channel(s) associated with the pixels in the pixel data can be selected based on the context (e.g., based on a probability determined based on the context).

[0255] In some implementations, the task can be an audio generation task. Machine-learned model(s) 1 can be configured to process input(s) 2 that represent context regarding a desired portion of audio content. The context can include text data, image data, audio data, etc. Machine-learned model(s) 1 can be configured to generate output(s) 3 that represent audio data related to the context. For instance, machine-learned model(s) 1 can be configured to generate waveform data in the form of an image (e g., a spectrogram). Values for channel(s)associated with pixels of the image can be selected based on the context. Machine-learned model(s) 1 can be configured to generate waveform data in the form of a sequence of discrete samples of a continuous waveform. Values of the sequence can be selected based on the context (e.g., based on a probability determined based on the context).

[0256] In some implementations, the task can be a data generation task. Machine-learned model(s) 1 can be configured to process input(s) 2 that represent context regarding a desired portion of data (e.g., data from various data domains, such as sensor data, image data, multimodal data, statistical data. etc.). The desired data can be. for instance, synthetic data for training other machine-learned models. The context can include arbitrary data type(s). Machine-learned model(s) 1 can be configured to generate output(s) 3 that represent data that aligns with the desired data. For instance, machine-learned model(s) 1 can be configured to generate data values for populating a dataset. Values for the data object(s) can be selected based on the context (e.g.. based on a probability determined based on the context).Example Computing Systems and Devices

[0257] FIG. 18 is a block diagram of an example networked computing system that can perform aspects of example implementations of the disclosure. The system can include a number of computing devices and systems that are communicatively coupled over a network 49. An example computing device 50 is described to provide an example of a computing device that can perform any aspect of the disclosure (e.g., implementing model host 31, client(s) 32, or both). An example server computing system 60 is described as an example of a server computing system that can perform any aspect of the disclosure (e.g., implementing model host 31, client(s) 32, or both). Computing device 50 and server computing system(s) 60 can cooperatively interact (e.g., over network 49) to perform any aspect of the disclosure (e.g., implementing model host 31, client(s) 32, or both). Model development platform system 70 is an example system that can host or serve model development platform(s) 12 for development of machine-learned models. Third-party system(s) 80 are example system(s) with which any of computing device 50, server computing system(s) 60, or model development platform system(s) 70 can interact in the performance of various aspects of the disclosure (e.g., engaging third-party tools, accessing third-party databases or other resources, etc.).

[0258] Network 49 can be any type of communications network, such as a local area network (e.g., intranet), wide area netw ork (e.g., Internet), or some combination thereof and caninclude any number of wired or wireless links. In general, communication over network 49 can be carried via any type of wired or wireless connection, using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), or protection schemes (e.g., VPN, secure HTTP, SSL). Network 49 can also be implemented via a system bus. For instance, one or more devices or systems of FIG. 11 can be co-located with, contained by, or otherwise integrated into one or more other devices or systems.

[0259] Computing device 50 can be any type of computing device, such as, for example, a personal computing device (e.g., laptop or desktop), a mobile computing device (e.g., smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, a server computing device, a virtual machine operating on a host device, or any other type of computing device. Computing device 50 can be a client computing device. Computing device 50 can be an end-user computing device. Computing device 50 can be a computing device of a service provided that provides a service to an end user (who may use another computing device to interact with computing device 50).

[0260] Computing device 50 can include one or more processors 51 and a memory 52. Processor(s) 51 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. Memory 52 can include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory' devices, magnetic disks, etc., and combinations thereof. Memory 52 can store data 53 and instructions 54 which can be executed by processor(s) 51 to cause computing device 50 to perform operations. The operations can implement any one or multiple features described herein. The operations can implement example methods and techniques described herein.

[0261] Computing device 50 can also include one or more input components that receive user input. For example, a user input component can be a touch-sensitive component (e.g., a touch-sensitive display screen or a touch pad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component can serve to implement a virtual keyboard. Other example user input components include a microphone, camera, LIDAR, a physical keyboard or other buttons, or other means by which a user can provide user input.

[0262] Computing device 50 can store or include one or more machine-learned models 55. Machine-learned models 55 can include one or more machine-learned model(s) 1, such as a sequence processing model 4. Machine-learned models 55 can include one or multiple model instance(s) 31-1. Machine-learned model (s) 55 can be received from server computing system(s) 60, model development platform system 70, third party system(s) 80 (e.g., an application distribution platform), or developed locally on computing device 50. Machine- learned model(s) 55 can be loaded into memory 52 and used or otherwise implemented by processor(s) 51. Computing device 50 can implement multiple parallel instances of machine- learned model(s) 55.

[0263] Server computing system(s) 60 can include one or more processors 61 and a memory 62. Processor(s) 61 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. Memory 62 can include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory' devices, magnetic disks, etc., and combinations thereof. Memory 62 can store data 63 and instructions 64 which can be executed by processor(s) 61 to cause server computing system(s) 60 to perform operations. The operations can implement any one or multiple features described herein. The operations can implement example methods and techniques described herein.

[0264] In some implementations, server computing system 60 includes or is otherwise implemented by one or multiple server computing devices. In instances in which server computing system 60 includes multiple server computing devices, such server computing devices can operate according to sequential computing architectures, parallel computing architectures, or some combination thereof.

[0265] Server computing system 60 can store or otherwise include one or more machine- learned models 65. Machine-learned model (s) 65 can be the same as or different from machine-learned model(s) 55. Machine-learned models 65 can include one or more machine- learned model(s) 1, such as a sequence processing model 4. Machine-learned models 65 can include one or multiple model instance(s) 31-1. Machine-learned model(s) 65 can be received from computing device 50, model development platform system 70, third party system(s) 80, or developed locally on server computing system(s) 60. Machine-learned model(s) 65 can be loaded into memory 62 and used or otherwise implemented byprocessor(s) 61. Server computing system(s) 60 can implement multiple parallel instances of machine-learned model(s) 65.

[0266] In an example configuration, machine-learned models 65 can be included in or otherwise stored and implemented by server computing system 60 to establish a client-server relationship with computing device 50 for serving model inferences. For instance, server computing system(s) 60 can implement model host 31 on behalf of client(s) 32 on computing device 50. For instance, machine-learned models 65 can be implemented by serv er computing system 60 as a portion of a web service (e.g., remote machine-learned model hosting service, such as an online interface for performing machine-learned model operations over a network on server computing system(s) 60). For instance, server computing system(s) 60 can communicate with computing device 50 over a local intranet or internet connection. For instance, computing device 50 can be a workstation or endpoint in communication with server computing system(s) 60, with implementation of machine-learned models 65 being managed by server computing system(s) 60 to remotely perform inference (e g., for runtime or training operations), with output(s) returned (e.g., cast, streamed, etc.) to computing device 50. Machine-learned models 65 can work cooperatively or interoperatively with machine- learned models 55 on computing device 50 to perform various tasks.

[0267] Model development platform system(s) 70 can include one or more processors 71 and a memory 72. Processor(s) 71 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. Memory 72 can include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory 72 can store data 73 and instructions 74 which can be executed by processor(s) 71 to cause model development platform system(s) 70 to perform operations. The operations can implement any one or multiple features described herein. The operations can implement example methods and techniques described herein. Example operations include the functionality described herein with respect to model development platform 12. This and other functionality can be implemented by developer tool(s) 75.

[0268] Third-party' system(s) 80 can include one or more processors 81 and a memory 82. Processor(s) 81 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA. a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. Memory 82 can includeone or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory 82 can store data 83 and instructions 84 which can be executed by processor(s) 81 to cause third-party system(s) 80 to perform operations. The operations can implement any one or multiple features described herein. The operations can implement example methods and techniques described herein. Example operations include the functionality described herein wi th respect to tools and other external resources called when training or performing inference with machine-learned model(s) 1, 4, 16, 20, 55, 65, etc. (e.g., third-party resource(s) 85).

[0269] FIG. 18 illustrates one example arrangement of computing systems that can be used to implement the disclosure. Other computing system configurations can be used as well. For example, in some implementations, one or both of computing system 50 or server computing system(s) 60 can implement all or a portion of the operations of model development platform system 70. For example, computing system 50 or server computing system(s) 60 can implement developer tool(s) 75 (or extensions thereof) to develop, update / train, or refine machine-learned models 1, 4, 16, 20, 55, 65, etc. using one or more techniques described herein with respect to model alignment toolkit 17. In this manner, for instance, computing system 50 or server computing system(s) 60 can develop, update / train, or refine machine- learned models based on local datasets (e.g., for model personalization / customization, as permitted by user data preference selections).

[0270] FIG. 19 is a block diagram of an example computing device 98 that performs according to example embodiments of the disclosure. Computing device 98 can be a user computing device or a server computing device (e.g., computing device 50, server computing system(s) 60, etc.). Computing device 98 can implement model host 31. For instance, computing device 98 can include a number of applications (e.g., applications 1 through N). Each application can contain its own machine learning library and machine-learned model(s). For example, each application can include a machine-learned model. Example applications include a biometric activity application, a content generation application, a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, a social media application, a chat application, etc. As illustrated in FIG. 12, each application can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, or additional components. In some implementations, each application cancommunicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0271] FIG. 20 is a block diagram of an example computing device 99 that performs according to example embodiments of the disclosure. Computing device 99 can be the same as or different from computing device 98. Computing device 99 can be a user computing device or a server computing device (e.g., computing device 50, server computing system(s) 60, etc.). Computing device 98 can implement model host 31. For instance, computing device 99 can include a number of applications (e.g., applications 1 through N). Each application can be in communication with a central intelligence layer. Example applications include a biometric activity application, a content generation application, a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, a social media application, a chat application, etc. In some implementations, each application can communicate with the central intelligence layer (and model(s) stored therein) using an API (e.g., a common API across all applications).

[0272] The central intelligence layer can include a number of machine-learned models. For example, as illustrated in FIG. 20, a respective machine-learned model can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine-learned model. For example, in some implementations, the central intelligence layer can provide a single model for all of the applications. In some implementations, the central intelligence layer is included within or otherwise implemented by an operating system of computing device 99.

[0273] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for computing device 99. As illustrated in FIG. 20, the central device data layer can communicate with a number of other components of the computing device, such as. for example, one or more sensors, a context manager, a device state component, or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).Additional Disclosure

[0274] The technology7discussed herein makes reference to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for agreat variety of possible configurations, combinations, and divisions of tasks and functionality between and among components. For instance, processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0275] Aspects of the disclosure have been described in terms of illustrative embodiments thereof. Any and all features in the following claims can be combined or rearranged in anyway possible, including combinations of claims not explicitly enumerated in combination together, as the example claim dependencies listed herein should not be read as limiting the scope of possible combinations of features disclosed herein. Accordingly, the scope of the disclosure is by way of example rather than by way of limitation, and the subject disclosure does not preclude inclusion of such modifications, variations or additions to the disclosure as would be readily apparent to one of ordinary' skill in the art. Moreover, terms are described herein using lists of example elements joined by conjunctions such as “and,” “or,” “but,” etc. It should be understood that such conjunctions are provided for explanatory- purposes only. Clauses and other sequences of items joined by a particular conjunction such as “or,” for example, can refer to “and / or,” “at least one of’, “any combination of’ example elements listed therein, etc. Terms such as “based on” should be understood as “based at least in part on.”

[0276] Terms used herein are used to describe the example embodiments and are not intended to limit and / or restrict the disclosure. The singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. In this disclosure, terms such as "including", "having", “comprising”, and the like are used to specify features, numbers, steps, operations, elements, components, or combinations thereof, but do not preclude the presence or addition of one or more of the features, numbers, steps, operations, elements, components, or combinations thereof.

[0277] The term "and / or" includes a combination of a plurality- of related listed items or any item of the plurality of related listed items. For example, the scope of the expression or phrase "A and / or B" includes the item "A", the item "B", and the combination of items "A and B”.

[0278] In addition, the scope of the expression or phrase "at least one of A or B" is intended to include all of the following: (1) at least one of A. (2) at least one of B, and (3) at least one of A and at least one of B. Likewise, the scope of the expression or phrase "at least one of A, B, or C" is intended to include all of the following: (1) at least one of A, (2) at least one of B, (3) at least one of C, (4) at least one of A and at least one of B, (5) at least one of A and at least one of C, (6) at least one of B and at least one of C, and (7) at least one of A, at least one of B. and at least one of C.

[0279] It will be understood that, although the terms first, second, third, etc., may be used herein to describe various elements, the elements are not limited by these terms. Instead, these terms are used to distinguish one element from another element. For example, without departing from the scope of the disclosure, a first element may be termed as a second element, and a second element may be termed as a first element.

[0280] The term “can’' should be understood as referring to a possibility of a feature in various implementations and not as prescribing an ability that is necessarily present in every implementation. For example, the phrase '‘X can perform Y” should be understood as indicating that, in various implementations, X has the potential to be configured to perform Y, and not as indicating that in every instance X must always be able to perform Y. It should be understood that, in various implementations, X might be unable to perform Y and remain within the scope of the disclosure.

[0281] The term “may” should be understood as referring to a possibility of a feature in various implementations and not as prescribing an ability that is necessarily present in every implementation. For example, the phrase “X may perform Y” should be understood as indicating that, in various implementations, X has the potential to be configured to perform Y, and not as indicating that in every instance X must always be able to perform Y. It should be understood that, in various implementations. X might be unable to perform Y and remain within the scope of the disclosure.

[0282] To the extent terms including "module", and "unit," and the like are used herein, these terms may refer to, but are not limited to, a software or hardware component or device, such as a Field Programmable Gate Array (FPGA) or Application Specific Integrated Circuit (ASIC), which performs certain tasks. A module or unit may be configured to reside on an addressable storage medium and configured to execute on one or more processors. Thus, a module or unit may include, by way of example, components, such as software components,object-oriented software components, class components and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. The functionality provided for in the components and modules / units may be combined into fewer components and modules / units or further separated into additional components and modules.

[0283] Aspects of the above-described example embodiments may be recorded in non- transitory computer-readable media including program instructions to implement various operations embodied by a computer. The media may also include, alone or in combination with the program instructions, data files, data structures, and the like. Examples of non- transitory computer-readable media include magnetic media such as hard disks, floppy disks, and magnetic tape; optical media such as CD ROM disks, Blu-Ray disks, and DVDs; magneto-optical media such as optical discs; and other hardware devices that are specialty configured to store and perform program instructions, such as semiconductor memory, readonly memory (ROM), random access memory (RAM), flash memory, USB memory, and the like. Examples of program instructions include both machine code, such as produced by a compiler, and files containing higher level code that may be executed by the computer using an interpreter. The program instructions may be executed by one or more processors. The described hardware devices may be configured to act as one or more software modules in order to perform the operations of the above-described embodiments, or vice versa. In addition, a non-transitory computer-readable storage medium may be distributed among computer systems connected through a network and computer-readable codes or program instructions may be stored and executed in a decentralized manner. In addition, the non- transitory computer-readable storage media may also be embodied in at least one application specific integrated circuit (ASIC) or Field Programmable Gate Array (FPGA).

[0284] Each block of the flowchart illustrations may represent a unit, module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of order. For example, two blocks shown in succession may7in fact be executed substantially7concurrently (simultaneously) or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.

[0285] Further to the descriptions above, a user may be provided with controls allowing the user to make an election as to both if and when systems, programs, or features describedherein may enable collection of user information (e.g., information about a user’s social network, social actions, or activities, profession, a user’s preferences, or a user’s current location), and if the user is sent content or communications from a server. In addition, certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user's identity may be treated so that no personally identifiable information can be determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined. Thus, the user may have control over what information is collected about the user, how that information is used, and what information is provided to the user.

[0286] While the disclosure has been described with respect to various example embodiments, each example is provided by way of explanation, not limitation of the disclosure. Those skilled in the art, upon attaining an understanding of the foregoing, can readily produce alterations to, variations of, and equivalents to such embodiments. Accordingly, the disclosure does not preclude inclusion of such modifications, variations and / or additions to the disclosed subject matter as would be readily apparent to one of ordinary skill in the art. For example, features illustrated or described as part of one embodiment can be used with another embodiment to yield a still further embodiment. Thus, it is intended that the disclosure covers such alterations, variations, and equivalents.

Claims

WHAT IS CLAIMED IS:

1. A computing system, comprising: one or more memories configured to store instructions: and one or more processors configured to execute the instructions to perform operations, the operations comprising: extracting, via one or more machine-learned models, context information from dialogue information received from a user via an activity agent which interacts with the user through dialogue operations; receiving first biometric information associated with the user relating to a first biometric activity', the first biometric information being captured by one or more sensors; storing, as a first episode among a plurality’ of episodes in the one or more memories, the context information, data associated with the first biometric activity, and the first biometric information; receiving a query associated with at least one of the context information or the first biometric activity; searching the plurality of episodes stored in the one or more memories according to the query; and when the first episode includes information which is responsive to the query', providing, for presentation to the user via the activity agent, one or more outputs related to the first biometric activity and the context information.

2. The computing system of claim 1, wherein each of the plurality' of episodes is associated with a time interval that is less than a predetermined duration of time.

3. The computing system of claim 1, wherein extracting, via the one or more machine-learned models, the context information from the dialogue information comprises identifying one or more topics in the dialogue information and mapping the one or more topics to at least one of a network model or a graph model.

4. The computing system of claim 3, wherein searching the plurality of episodes stored in the one or more memories according to the query' comprises traversing the at least one of the network model or the graph model to determine one or more topics mapped to the at least one of the network model or the graph model which correspond to at least one topic in the query'.

5. The computing system of claim 1, wherein when the query is associated with the context information, information included in the plurality of episodes stored in the one or more memories is searched for the context information; and when the first episode includes information corresponding to the context information, providing, for presentation to the user via the activity agent, one or more outputs related to the first biometric activity and the context information.

6. The computing system of claim 5, wherein the context information relates to an emotional state of the user. searching the plurality of episodes includes searching for information corresponding to the emotional state of the user; and when the first episode includes information corresponding to the emotional state of the user, the one or more outputs provided for presentation to the user via the activity’ agent indicate the first biometric activity is associated with the emotional state of the user.

7. The computing system of claim 1, wherein when the query is associated with the first biometric activity, information included in the plurality of episodes stored in the one or more memories is searched for the data associated with the first biometric activity; and when the first episode includes the data associated with the first biometric activity’, providing, for presentation to the user via the activity agent, one or more outputs related to the first biometric activity and the context information.

8. The computing system of claim 7, wherein emotional state information associated with the user is mapped to the first biometric activity and stored as part of the first episode among the plurality of episodes in the one or more memories, and the operations further comprise: when the first episode includes the data associated yvith the first biometric activity’, retrieving from the first episode the emotional state information associated with the user mapped to the first biometric activity; andthe one or more outputs provided for presentation to the user via the activity agent indicate an emotional state of the user based on the emotional state information associated with the user mapped to the first biometric activity.

9. The computing system of claim 1, wherein the first biometric information indicates physiological information associated with the user measured while the user performs the first biometric activity, and the physiological information associated with the user is mapped to the data associated with the first biometric activity' and stored as part of the first episode among the plurality of episodes in the one or more memories.

10. The computing system of claim 9, wherein when the query is associated with the physiological information, information included in the plurality of episodes stored in the one or more memories is searched for the physiological information; when the first episode includes the physiological information, retrieving from the first episode the first biometric activity; and the one or more outputs provided for presentation to the user via the activity agent indicate the physiological information is associated with the first biometric activity.1 1 . The computing system of claim 9, wherein when the query is associated with the first biometric activity, information included in the plurality of episodes stored in the one or more memories is searched for the first biometric activity; when the first episode includes the first biometric activity, retrieving from the first episode the physiological information associated with the user mapped to the first biometric activity; and the one or more outputs provided for presentation to the user via the activity agent indicate the first biometric activity is associated with the physiological information.

12. The computing system of claim 1, wherein the operations further comprise receiving second biometric information associated with the user relating to a second biometric activity, the second biometric information being captured by the one or more sensors, andthe second biometric information and data associated with the second biometric activity are mapped to the first biometric activity and stored as part of the first episode among the plurality of episodes in the one or more memories.

13. The computing system of claim 12, wherein when the query is associated with the first biometric activity and the second biometric activity, information included in the plurality of episodes stored in the one or more memories is searched for data associated with the first biometric activity and the second biometric activity7; when the first episode includes the data associated with the first biometric activity7and the second biometric activity, retrieving, from the first episode, context information associated with the user which is mapped to the first biometric activity and the second biometric activity; and the one or more outputs provided for presentation to the user via the activity7agent indicate the context information associated with the user which is mapped to the first biometric activity and the second biometric activity.

14. A computer-implemented method, comprising: extracting, via one or more machine-learned models of a computing system comprising one or more processors, context information from dialogue information received from a user via an activity agent which interacts with the user through dialogue operations; receiving, by the computing system, first biometric information associated with the user relating to a first biometric activity, the first biometric information being captured by one or more sensors; storing, as a first episode among a plurality of episodes in one or more memories of the computing system, the context information, data associated with the first biometric activity, and the first biometric information; receiving, by the computing system, a query associated with at least one of the context information or the first biometric activity; searching, by the computing system, the plurality of episodes stored in the one or more memories according to the query; and when the first episode includes information which is responsive to the query7, providing, for presentation to the user via the activity agent, one or more outputs related to the first biometric activity and the context information.

15. The computer-implemented method of claim 14. wherein each of the plurality’ of episodes is associated with a time interval that is less than a predetermined duration of time.

16. The computer-implemented method of claim 14. wherein extracting, via the one or more machine-learned models, the context information from the dialogue information comprises identifying one or more topics in the dialogue information and mapping the one or more topics to at least one of a network model or a graph model.

17. The computer-implemented method of claim 16. wherein searching the plurality of episodes stored in the one or more memories according to the query comprises traversing the at least one of the network model or the graph model to determine one or more topics mapped to the at least one of the network model or the graph model which correspond to at least one topic in the query.

18. The computer-implemented method of claim 14, wherein when the query' is associated with the context information, information included in the plurality of episodes stored in the one or more memories is searched for the context information; and when the first episode includes information corresponding to the context information, providing, for presentation to the user via the activity’ agent, one or more outputs related to the first biometric activity and the context information.

19. The computer-implemented method of claim 18, wherein the context information relates to an emotional state of the user, searching the plurality7of episodes includes searching for information corresponding to the emotional state of the user; and when the first episode includes information corresponding to the emotional state of the user, the one or more outputs provided for presentation to the user via the activity' agent indicate the first biometric activity7is associated with the emotional state of the user.

20. A non-transitory computer readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform operations, the operations comprising: extracting, via one or more machine-learned models, context information from dialogue information received from a user via an activity7agent which interacts with the user through dialogue operations; receiving first biometric information associated with the user relating to a first biometric activity, the first biometric information being captured by one or more sensors; storing, as a first episode among a plurality7of episodes in the non-transitory7computer readable medium, the context information, data associated with the first biometric activity, and the first biometric information; receiving a query associated with at least one of the context information or the first biometric activity; searching the plurality of episodes stored in the non-transitory computer readable medium according to the query7; and when the first episode includes information which is responsive to the query, providing, for presentation to the user via the activity agent, one or more outputs related to the first biometric activity7and the context information.

Citation Information

Patent Citations

  • Method for providing health therapeutic interventions to a user

    US20210391083A1

  • Enabling user-centered and contextually relevant interaction

    US20230245651A1