Phone and messaging dynamic actions

The computing device leverages machine learning to analyze communication data, identify actions, and generate software components to automate task execution in applications, addressing the challenge of information recall and application interaction post-communication.

WO2026029779A1PCT designated stage Publication Date: 2026-02-05GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/040853
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Users face challenges in recalling information discussed during calls or messages and find it inconvenient to interact with multiple applications to update information, such as adding events or setting routes, after a communication session.

Method used

A computing device uses machine learning models to analyze user interaction data from calls or messages, dynamically identify relevant actions, generate software components, and cause applications to perform those actions without requiring user input or native application support.

Benefits of technology

Enables automated task execution based on communication data, enhancing user convenience by eliminating the need for manual note-taking and simplifying interaction with multiple applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024040853_05022026_PF_FP_ABST
    Figure US2024040853_05022026_PF_FP_ABST
Patent Text Reader

Abstract

A computing device may extract user interaction data from one of a phone call or a message. The computing device may determine, using a machine learning model, a particular action to perform based on the user interaction data. The computing device may dynamically identify, based on the particular action, a particular application from one or more applications executed by the computing device. The computing device may dynamically generate a software component that enables the computing device to cause the particular action. The computing device may cause, using the software component, the particular application to perform the particular action.
Need to check novelty before this filing date? Find Prior Art

Description

PHONE AND MESSAGING DYNAMIC ACTIONSBACKGROUND

[0001] The widespread usage of mobile devices, such as smartphones, enables users to be able to place and receive calls and messages to communicate information with other individuals and organizations. For example, users may communicate upcoming plans, personal information, and other important information via phone calls and messaging applications. In addition, users may communicate information that relates to other applications executed by a mobile device. For example, a user may communicate information about an upcoming group run that the user plans to track using a fitness application.SUMMARY

[0002] In general, the techniques of this disclosure are directed to techniques for enabling a computing device to, after receiving user permission, use information from a call or messages to identify an action for an application to perform and generate software components to enable the application to perform that action. Unlike traditional systems that require applications to support performing actions or requiring manual input from a user, techniques of this disclosure leverage machine learning models to automatically identify relevant actions and create the necessary software components in real time to cause the applications to perform the actions. This approach not only enhances user convenience by automating tasks but also provides a flexible and adaptive system capable of integrating with various applications.

[0003] A computing device, such as a smartphone, dynamically generates the necessary software components to enable the application to perform the action. For instance, the computing device may utilize machine learning models, such as a language model (e.g., a large language model (LLM)), to generate these software components. Once generated, the computing device uses the dynamically generated software component to cause the application to perform the action.

[0004] In one example, this disclosure describes a method that includes extracting, by a computing device, user interaction data from one of: a phone call or a message; determining, by the computing device and using a machine learning model, a particular action to perform based on the user interaction data; dynamically identifying, by the computing device and based on the particular action, a particular application from one or more applications capable of execution by the computing device; dynamically generating, by the computing device, asoftware component that enables the computing device to cause the particular application to perform the particular action; and causing, by the computing device and using the software component, the particular application to perform the particular action.

[0005] In another example, this disclosure describes a computing device that includes a memory and one or more processors implemented in circuitry in communication with the memory extract user interaction data from one of: a phone call or a message; determine, using a machine learning model, a particular action to perform based on the user interaction data; dynamically identify, based on the particular action, a particular application from one or more applications capable of execution by the computing device; dynamically generate a software component that enables the computing device to cause the particular application to perform the particular action; and cause, using the software component, the particular application to perform the action.

[0006] In another example, this disclosure describes a non-transitory computer-readable storage medium encoded with instructions that, when executed by one or more processors, cause the one or more processors to extract user interaction data from one of: a phone call or a message; determine, using a machine learning model, a particular action to perform based on the user interaction data; dynamically identify, based on the particular action, a particular application from one or more applications capable of execution by the computing device; dynamically generate a software component that enables the computing device to cause the particular application to perform the particular action; and cause, using the software component, the particular application to perform the action.

[0007] In yet another, this disclosure describes a computer program product comprising instructions that, when executed, cause one or more processors to extract user interaction data from one of: a phone call or a message; determine, using a machine learning model, a particular action to perform based on the user interaction data; dynamically identify, based on the particular action, a particular application from one or more applications capable of execution by the computing device; dynamically generate a software component that enables the computing device to cause the particular application to perform the particular action; and cause, using the software component, the particular application to perform the action.

[0008] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the disclosure will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF DRAWINGS

[0009] FIG. 1 is a conceptual diagram illustrating an example computing system for determining actions and dynamically generating software components, in accordance with one or more aspects of the present disclosure.

[0010] FIG. 2 is a block diagram illustrating an example computing device, in accordance with one or more aspects of the present disclosure.

[0011] FIG. 3 is a flowchart illustrating example operations performed by a computing device that determines actions and dynamically generates software components, in accordance with one or more aspects of the present disclosure.

[0012] FIG. 4 is a flowchart illustrating example operations performed by an action framework for determining actions and dynamically generating software components, in accordance with one or more aspects of the present disclosure.

[0013] FIG. 5 is a flowchart illustrating example operations performed by an action framework for determining actions and dynamically generating software components, in accordance with one or more aspects of the present disclosure.DETAILED DESCRIPTION

[0014] FIG. 1 is a conceptual diagram illustrating an example computing system 100 for determining actions and dynamically generating software components, in accordance with one or more aspects of the present disclosure. In the example of FIG. 1, computing system 100 may include computing device 102 that connects to network 140 to exchange phone calls and messages with other communication devices, such as other computing device 136.

[0015] As shown in FIG. 1, computing device 102 may represent an individual mobile or non-mobile computing device. Examples of computing device 102 include a mobile phone, a tablet computer, a laptop computer, a desktop computer, a server, a mainframe, a set-top box, a television, a wearable device (e.g., a computerized watch, computerized eyewear, computerized headphones, computerized gloves, augmented reality (AR) glasses / goggles, virtual reality (VR) glasses / goggles, artificial intelligence (Al)-enabled glasses / goggles, AI- enabled pin, etc.), a home automation device or system (e.g., an intelligent thermostat or home assistant device), a gaming system, a media player, an e-book reader, a mobile television platform, an automobile navigation or infotainment system, or any other type of mobile, non-mobile, wearable, and non-wearable computing device.

[0016] Computing device 102 may connect to network 140 to provide communication, such as calls and messages, to and from computing devices. Network 140 represents public orprivate communications network, such as WIFI, one or more wireless wide area networks (WAN) (e.g., a wireless cellular network, a satellite network, and / or a Free-Space Optical Communication network), one or more telephony network, such as one or more Public Switched Telephone Networks (PTSNs), one or more VoIP services, one or more hardwired networks (e.g., an Ethernet connection to a WAN, optical connections to other computing devices), and / or other types of networks, for transmitting data between computing systems, servers, and computing devices, such as over the Internet. Network 140 may include one or more network hubs, network switches, network routers, or any other network equipment, that are operatively inter-coupled thereby providing for the exchange of information between computing device 102 and remote computing devices, such as other computing device 150. Computing device 102 and remote computing devices, such as other computing device 150 may transmit and receive data across network 140 using any suitable communication techniques. Computing device 102 may be operatively coupled to network 130 using respective network links, such as Ethernet, WIFI, a cellular connection, satellite connection, optical connection (e.g., SONET, SDH), BLUETOOTH connection, or any other types of wired and / or wireless network connections.

[0017] Network 140 may implement any suitable technology and may include any suitable networks that enable computing device 102 to place and receive calls as well as send and receive messages to and from other computing devices, such as computing devices remote to computing device 102 (e.g., other computing device 150). In some examples, network 140 may implement an IP Multimedia Subsystem (IMS) that manages call sessions, including call routing, authentication and billing. The IMS may also act as a Session Initiation Protocol (SIP) server that uses SIP to perform call setup and teardown functions and to perform signaling and messaging protocols used for calls between computing devices. In some examples, network 140 may implement the functionality of an Evolved Packet (EPG) core or a System Architecture Evolution (SAE) core to handle communication between devices connected to network 140 and to networks external to network 140. Network 140 may connect one or more services that enable the exchange of messages between computing devices. For example, one or more servers may be connected to computing device 102 via network 140 and provide a messaging service such as text message and / or internet messaging.

[0018] Other computing device 150 represents a device that is connected to network 140 to make and receive calls, such as voice calls and / or video calls, and / or to send and receive messages. Examples of other computing device 150 include a landline telephone, a mobilephone, a tablet computer, a laptop computer, a desktop computer, a satellite phone, wearable computing device, a server, a mainframe, a set-top box, a television, a wearable device (e.g., a computerized watch, computerized eyewear, computerized headphones, computerized gloves, augmented reality (AR) glasses / goggles, virtual reality (VR) glasses / goggles, AI- enabled glasses / goggles, etc.), a home automation device or system (e.g., an intelligent thermostat or home assistant device), a gaming system, a media player, an e-book reader, a mobile television platform, an automobile navigation or infotainment system, or any other type of mobile, non-mobile, wearable, and non-wearable computing device, or any other type of device that can communicate with network 140.

[0019] As shown in FIG. 1, computing device 102 includes user interface component (UIC) 104, user interface module 106 (“UI module 106”), ML models 108, component generator 110, messaging application 112, phone application 114, applications 116A-N (hereinafter “applications 116”), and action module 120. UIC 104 of computing device 102 may function as an input device for computing device 102 and as an output device for computing device 102. UIC 104 may be implemented using various technologies. For instance, UIC 104 may function as an input device using a presence-sensitive input screen, such as a resistive touchscreen, a surface acoustic wave touchscreen, a capacitive touchscreen, a projective capacitive touchscreen, a pressure sensitive screen, an acoustic pulse recognition touchscreen, or another presence-sensitive display technology, such as radar-based presence-sensitive technology millimeter wave-based presence-sensitive technology, ultra-wideband-based presence-sensitive technology, and the like. In some examples, UIC 104 may function as an input device using one or more audio input devices, such as one or more microphones. UIC 104 may function as an output (e.g., display) device using any one or more display devices, such as a liquid crystal display (LCD), dot matrix display, light emitting diode (LED) display, microLED, miniLED, organic light-emitting diode (OLED) display, e-ink, or similar monochrome or color display capable of outputting visible information to a user of computing device 102. In some examples, UIC 104 may function as an audio output device and may include one or more speakers, one or more headsets, or any other audio output device capable of outputting audible information to a user of computing device 102.

[0020] In some examples, UIC 104 of computing device 102 may include a presencesensitive display that may receive tactile input from a user of computing device 102. UIC 104 may receive indications of the tactile input by detecting one or more gestures from a user of computing device 102 (e.g., the user touching or pointing to one or more locations of UIC 104 with a finger or a stylus pen). UIC 104 may present output to a user, for instance at apresence-sensitive display. UIC 104 may present the output as a graphical user interface (e.g., graphical user interfaces 118A-B, hereinafter “GUIs 118”), which may be associated with functionality provided by computing device 102. For example, UIC 104 may present various user interfaces of components of a computing platform, operating system (OS), applications (e.g., messaging application 112, phone application 114, applications 116, etc.), or services executing at or accessible by computing device 102 (e.g., an electronic message application, an Internet browser application, a mobile operating system, etc.). A user may interact with a respective user interface to cause computing device 102 to perform operations relating to a function.

[0021] UI module 106, ML models 108, component generator 110, messaging application 112, phone application 114, applications 116, and action module 120 may perform operations described herein using software, hardware, firmware, or a mixture of both hardware, software, and firmware residing in and executing on computing device 102 or at one or more other computing devices. In some examples, UI module 106, ML models 108, component generator 110, messaging application 112, phone application 114, applications 116, and action module 120 may be implemented as hardware, software, and / or a combination of hardware and software. Computing device 102 may execute UI module 106, ML models 108, component generator 110, messaging application 112, phone application 114, applications 116, and action module 120 with one or more processors. Computing device 102 may execute any of UI module 106, ML models 108, component generator 110, messaging application 112, phone application 114, applications 116, and action module 120 as or within a virtual machine executing on underlying hardware. UI module 106, ML models 108, component generator 110, messaging application 112, phone application 114, applications 116, and action module 120 may be implemented in various ways. For example, any of UI module 106, ML models 108, component generator 110, messaging application 112, phone application 114, applications 116, and / or action module 120 may be implemented as a downloadable or pre-installed application or “app.” In another example, any of UI module 106, ML models 108, component generator 110, messaging application 112, phone application 114, applications 116, and action module 120 may be implemented as part of an operating system of computing device 102. In some additional examples, a server may execute any of UI module 106, ML models 108, component generator 110, and / or other components. Other examples of computing device 102 that implement techniques of this disclosure may include additional components not shown in FIG. 1.

[0022] UI module 106 may interpret inputs detected at UIC 104. UI module 106 may relayinformation about the inputs detected at UIC 104 to one or more associated platforms, operating systems, applications, and / or services executing at computing device 102 to cause computing device 102 to perform a function. UI module 106 may also receive information and instructions from one or more associated platforms, operating systems, applications, and / or services executing at computing device 102 (e.g., phone application 114) for generating a GUI. In addition, UI module 106 may act as an intermediary between the one or more associated platforms, operating systems, applications, and / or services executing at computing device 102 and various output devices of computing device 102 (e.g., speakers, LED indicators, vibrators, haptic engines, etc.) to produce output (e.g., graphical, audible, tactile, etc.) with computing device 102. For example, UI module 106 may cause UIC to display a GUI generated by phone application 114.

[0023] Applications 116 may include one or more applications capable of execution by computing device 102. For example, applications 116 may include one or more mobile applications capable of and / or being executing at a smartphone such as computing device 102.

[0024] Phone application 114 may include functionality for placing and receiving calls, via network 140, to and from computing devices such as other computing device 150. Examples of phone application 114 may include a phone dialer application, a Voice over IP (VoIP) application, a messaging application with voice calling functionality, a video conferencing application, a video calling application, or any other application that includes functionality for placing and receiving calls. Phone application 114 may enable a user of computing device 102 to place calls and converse with other individuals, such as a user of other computing device 150.

[0025] Messaging application 112 may include functionality for sending and receiving messages via network 140 to and from computing devices such as other computing device 150. Examples of messaging application 112 include a text messaging application, an internet messaging application, social media application with messaging functionality, or any other application that includes functionality for sending and receiving messages. Messaging application 112 may enable a user of computing device 102 to exchange messages, such as text messages, videos clips, emoticons (e.g., emojis, animated images, etc.), and other types of messages with another individual such as a user of other computing device 150.

[0026] In the example of FIG. 1, applications such as messaging application 112, phone application 114, and / or applications 116 may send data to UI module 106 that causes UIC 104 to generate user interfaces (GUIs), such as GUIs 118 and elements thereof. In response,UI module 106 may output instructions and information to UIC 104 that cause UIC 104 to display a user interface (e.g., GUI 118A) according to the information received from the application. When handling input detected by UIC 104, UI module 106 may receive information from UIC 104 in response to inputs detected at locations of a screen of UIC 104 at which elements of the user interface are displayed. UI module 106 disseminates information about inputs detected by UIC 104 to other components of computing device 102 for interpreting the inputs and for causing computing device 102 to perform one or more functions in response to the inputs.

[0027] Computing device 102 may establish a call and / or facilitate messaging with other computing device 150, such as by placing a call to other computing device 150 and / or transmitting data of messages. Examples of calls placed and received may include a voice call such as a telephone call or a Voice over Internet Protocol (VoIP) call, a video call such as a videoconferencing call, a real-time mixed reality session, a real-time augmented reality session, a media call, or any other calls between two or more devices. Once computing device 102 has established the call session with other computing device 150, computing device 102 and other computing device 150 may be able to exchange data, such as audio data.

[0028] Computing device 102 may facilitate the exchange of message(s) between computing devices, such as part of a message thread with other computing device 150. Examples of messaging include text messaging, a chat-bot session, an online messaging session, and / or other types of messaging to other computing device 150. Computing device 102 may communicate audio and / or message data such as text, audio clips, images, and other data as part facilitating an exchange of messages with another computing device, such as other computing device 150. Similarly, computing device 102 may receive, from other computing device 150, text data (text of a message), audio data such as words and phrases spoken by a user of other computing device 150, and / or image data (e.g., videos, still images, stickers, etc.). In some examples, computing device 102 may establish messaging with another computing device, where messaging application 112 automatically deletes messages after a certain period of time (e.g., due to company policies, as part of the functionality of a social messaging service, etc.).

[0029] Following a call and / or an exchange of messages, a user of computing device 102 may find it challenging to recall what was discussed during the call or in the messages and any future actions that they may need to take. In an example, a user of computing device 102 forgets the time and location of an upcoming group run that was discussed during a phone call. In addition, a user of computing device 102 may find it inconvenient to interact withmultiple of applications 116 to update information following a phone call or messaging session. For example, a user of computing device 102 may find it tedious to interact with a calendar application to add an event, a running application to map a run, and a mapping application to set an upcoming route and departure time.

[0030] In accordance with the techniques of this disclosure, computing device 102 may, after receiving explicit user permission, analyze incoming and / or outgoing calls and messages between computing device 102 and another device, such as other computing device 150. Computing device 102 may analyze calls and messages to identify actions for applications 116 to perform based on the content of the calls and messages. For example, computing device 102 may analyze a call between a user of computing device 102 and other computing device 150 and dynamically identify one or more actions for applications 116 to perform based on the call. Computing device 102 may dynamically generate one or more software components to enable computing device 102 to cause the one or more identified applications of applications 116 to perform the actions. The use of action module 120 and component generator 110 may provide a technical solution to the technical problem of a user failing to recall information discussed in an interaction and finding it challenging to interact with applications 116 based on what was discussed.

[0031] Computing device 102 includes action module 120. Action module 120 may be a software component of computing device 102, such as a process, plugin, module, or other type of software component. In instances where a user provides explicit consent, action module 120 may analyze calls and messages to identify actions to be performed by applications such as applications 116. For example, action module 120 may extract user interaction data from a call and determine an action to be performed by application 116A. Action module 120 may extract user interaction data that is representative of user discussions (e.g., calls, messages, etc.) and that may include information regarding user plans or events (e.g., information regarding upcoming plans or events).

[0032] Action module 120 refrains from storing and / or saving the user interaction data. Action module 120 may refrain from storing user interaction data, such as audio of a call, transcripts of a call, message data, and other types of data, within memory of computing device 102. Action module 120 refrains from storing user interaction to ensure the security and confidentiality of the user interaction data and to conform with regulatory requirements (e.g., to conform to two-party recording requirements). In an example, action module 120 extracts user interaction data and uses the user interaction data to generate actions and dynamically generate software actions. Action module 120 then promptly deletes the userinteraction data to avoid retaining the user interaction data.

[0033] Action module 120 requests user approval prior to analyzing calls or messages and prior to extracting user interaction data from a call or messaging session. Action module 120 requires user approval prior to extracting user interaction data from calls or messages. For example, action module 120 may request user approval to analyze a call and extract user interaction data after the start of a phone call and prior to beginning to analyze the call or extracting any user interaction data. If the user of computing device 102 does not provide explicit permission for action module 120 to analyze the call, action module 120 does not analyze the call or extract any user interaction data. Action module 120 may cause UIC to output GUI 118A as including extraction approval element 122. Action module 120 may extract user interaction data in response to computing device 102 receiving user input consistent with a user of computing device 102 interacting with extraction approval element 122.

[0034] Action module 120 may provide an indication that a call is being recording and that Al is being used to process the call. Action module 120 may provide an indication, such as an audio message (e.g., “This call is being recorded and processed by an Al assistant”) during the call such that other members of the call are made aware of the use of action module 120. Action module 120 may provide the indication after receiving explicit user approval to analyze the call and prior to obtaining any audio data from phone application 114.

[0035] Action module 120 may obtain audio data from phone application 114 and preprocess the audio data into text. Phone application 114 may automatically provide audio data to action module 120 (after a user has approved the collection of such data) or in response to a request by action module 120. For example, action module 120 may obtain audio data from phone application 114 (e.g., audio data of the user of computing device 102 and other users speaking generated by phone application 114) and extract user interaction data from the audio data. Action module 120 may process audio data, image data, and other types of non-textual data into text. Action module 120 may process non-textual data into text using one or more techniques, such as applying speech-to-text models, providing the non-textual data to ML models that are capable of processing multimodal input and receiving text as output, and other techniques. In an example, action module 120 provides audio data to an ML model trained to process audio and generate transcripts based on the audio. Action module 120 receives a transcript of spoken language from the ML model as output.

[0036] Action module 120 may extract user interaction data from data obtained from messaging application 112 and text based on the non-textual data (e.g., non-textual data, suchas audio data and / or image data, that has been preprocessed into text). Action module 120 may extract user interaction data from messages by obtaining data from messaging application 112 regarding one or more messages and extracting the user interaction data from the data. For example, action module 120 may obtain data regarding a messaging chain between a user of computing device 102 and a user of another computing device and extract user interaction data from the data regarding the messaging chain. Action module 120 obtains data from messaging application 112 only when a user has provided explicit approval for the extraction of user interaction data. In some examples, action module 120 may extract user interaction data from messages that include multimodal elements (e.g., emoticons, animated images, audio clips, videos, etc.). Action module 120 may extract user interaction data using one or more techniques such as identifying key words spoken in the audio data, intonation of the speakers speaking, and / or other techniques.

[0037] Computing device 102 may include one or more of ML models 108. ML models 108 may include one or more types of ML models such unsupervised models, supervised models, deep-learning models, q-learning models, rules-based models, neural networks, natural language processing (NLP) models, language models such as large language models (LLMs), and / or other types of ML models. FIG. 1 illustrates only one particular example of ML models 108, and many other examples of ML models 108 may be used in accordance with techniques of this disclosure. In some examples, components of ML models 108 may be located in a singular location. In other examples, one or more components of ML models 108 may be in different locations (e.g., connected via network 140). That is, in some examples ML models 108 may be part of a conventional computing device, while in other examples, ML models 108 may be part of a distributed or “cloud” computing system. Further, ML models 108 may be located in and executed by a computing system remote from and communicatively coupled to computing device 102. In such examples, computing device 102 may send and receive messages, data, or otherwise exchange information with the remote computing system such that the remote computing system provides computing device 102 with the functionality of ML models 108. In some examples, ML models 108 may be functionally split between computing device 102 and one or more remote computing systems such that a portion of the functionality provided by ML models 108 is performed locally at computing device 102 and other portions of the functionality are performed by the one or more remote computing systems.

[0038] In some examples, ML models 108 may be configured to interpret both text and audio data obtained by action module 120. In some examples, ML models 108 may be configuredto infer any indication of a natural language user input. In other words, ML models 108 may infer capabilities from user intents. In some examples, ML models 208 may have search capabilities. In some examples, ML models 108 may convert the audio data, text data, and other types of data (e.g., image data) obtained from messaging application 112 or phone application 114 into structured text. For example, ML models 108 may convert any input or information to an extensible Markup Language (XML), or into other structured text types, such as, but not limited to, HTML, JSON, CSV, INI Files, etc. In this way, the data obtained by action module 120 can be provided to ML models 108 in a standardized format. ML models 108 may further determine the type of information to include in the structured text representation. More specifically, ML models 108 may analyze the data to identify particular keywords, dates, times, locations, relationships among words in the data, and identities of individuals in the data, among other information for inclusion in the user interaction data.

[0039] Action module 120 may use one or more of ML models 108 to extract user interaction data from messages and text generated based on non-textual data (e.g., audio data of a call). Action module 120 may extract user interaction data by providing input of the text data to one or more of ML models, such as an LLM of ML models 108, and receiving the user interaction data from the LLM. In addition, action module 120 may provide a prompt for the LLM along with the text data. Action module 120 may provide a prompt that includes one or more requirements for extraction the user interaction data, such as to provide the output in a particular format, instructions to identify keywords (e.g., names, dates, locations, etc.), instructions to identify relationships within the text data (e.g., associations between names and other data), and instructions to identify information relevant to extracting user interaction data (e.g., instructions to filter out information that is irrelevant to the extraction of user interaction data) among other requirements. In an example, action module 120 provides a transcript of a conversation between a user of computing device 102 and a user of other computing device 150 and a prompt that includes a requirement to generate user interaction data as input to an LLM of ML models 108. The LLM processes the input and provides an output of the user interaction data to action module 120. ML models 108 may output user interaction data that includes one or more types of data, such as vectors (e.g., vectorized representations of relationships within the user interaction data), lists of identified keywords (e.g., names, dates, locations, etc. identified by ML models 108), and other data.

[0040] Action module 120 may determine a particular action to perform based on user interaction data. Action module 120 may determine a particular action for an application of applications 116 to perform. In an example, action module 120 extracts user interaction datathat includes a discussion between users of a lunch scheduled for a particular restaurant in Manhattan. Action module 120 determines, based on the user interaction data, that a calendar reminder should be generated for the upcoming lunch and that the user of computing device 102 should be alerted as to when they should leave for the particular restaurant. In some examples, action module 120 may determine multiple actions for one or more of applications 116 to perform based on the user interaction data.

[0041] Action module 120 may use one or more of ML models 108 to determine the particular action. Action module 120 may use one or more of ML models 108 that are trained to determine actions to perform based on the user interaction data. Action module 120 may provide the user interaction data to ML models 108 and receive an indication of one or more actions, such as the particular action, to perform. In an example, action module 120 generates a prompt for an LLM of ML models 108 that includes a request to identify a particular action. Action module 120 provides the prompt, the user interaction data and information regarding application 116 to ML models 108 for processing. ML models 108 process the prompt, the user interaction data, and the information regarding application 116 and outputs an indication of a particular action to be performed by computing device 102.

[0042] In some examples, action module 120 may use a predetermined list of actions to determine the particular action. Action module 120 may use the predetermined list in addition to or in lieu of using ML models 108 to determine the particular action. Action module 120 may use a list from one or more sources such as developers of an operating system of computing device 102, developers of one or more applications such as applications 116, and other sources. For example, action module 120 may use a list created by a developer of action module 120 to determine the particular action. In an example, action module 120 extracts user interaction data that includes an indication that a user of computing device 102 is going to receive a replacement credit card. Action module 120 determines, using the predetermined list, that a reminder should be generated to remind the user in ten days that their replacement credit card may have arrived.

[0043] Action module 120 may identify a plurality of potential actions that applications 116 may perform. Action module 120 may identify a plurality of potential actions, where each potential action of the plurality of potential actions is performed by a corresponding application of applications 116 executed by computing device 102. Action module 120 may identify one or more potential actions for each application of applications 116. In an example, action module 120 identifies five actions that a fitness tracking application may perform including saving a route for an upcoming run and updating a food consumption tracker.Action module 120 may identify the potential actions using one or more techniques and / or sources of information. Action module may use techniques and sources of information, such as analyzing an application store description of an application, analyzing an application programming interface (API) of an application, obtaining a list of applications from the application store, leveraging a language model trained on data that includes information about one or more of applications 116, a list of actions that includes information about a respective set of functionality of each application, and / or pre-training data of a language model that includes information regarding the particular application, among other applications.

[0044] Action module 120 may dynamically identify a particular application from applications 116. Action module 120 may dynamically identify the application based on the particular action to perform. For example, action module 120 may identify a particular application of applications 116 based on the type of action of the particular action and actions that the particular application can perform. Action module 120 may use information regarding potential actions that applications 116 can perform to identify the particular application of applications 116. In an example, action module 120 compares the particular action to a list of actions that applications 116 may perform and identifies the particular application based on the comparison. Action module 120 may dynamically identify a particular application in that action module 120 may identify the particular application in real-time or near real-time and use one or more techniques to identify the particular application. In some examples, action module 120 may dynamically identify a particular action based on identifying the particular action from a plurality of potential actions.

[0045] Action module 120 may use component generator 110 to dynamically generate a software component that enables computing device 102 to cause the particular application to perform the particular action. Component generator 110 may receive an indication from action module 120 to generate one or more software components to cause an application of application 116 to perform an action. Component generator 110 may determine the type and configuration of a software component that is needed to cause an application to perform an action. Component generator 110 may use one or more techniques to determine the type and configuration of the software component, such as determining that an application programming interface (API) is required to interface with an application, that a process may be used to emulate user interaction with an application, that component generator 110 needs to generate input data and provide the input data to the application, and / or other techniques.

[0046] Component generator 110 may dynamically generate one or more softwarecomponents, such as APIs, processes, executables, nano-applications, plugins, and / or other types of software components to cause an application of applications 116 to perform an action. For instance, component generator 110 may generate software components by composing code for the software component and packaging the code into an executable. Component generator 110 may dynamically generate the software component in the sense that component generator 110 may generate the software component “on the fly” or in realtime when needed, as opposed to requiring preemptively generated software components. In addition, component generator 110 may generate the software component as custom -tailored to enable computing device 102 to cause an application to perform an action (e.g., generate a new software component to enable computing device 102 to cause the application to perform the action). Further, component generator 110 may generate software components that enable additional and / or simplified (e.g., from a user perspective) functionality for applications 116. For example, component generator 110 may generate a software component that enables the creation of a calendar event in a calendar application, where the calendar event includes information not typically available when using a “new event” interface of the calendar application, and allows a user of computing device 102 to edit the calendar event after the creation of the calendar event.

[0047] In some examples, component generator 110 may use one or more of ML models 108 to generate software components. Component generator 110 may use a model of ML models 108 such as a language model to generate software components. Component generator 110 may generate a prompt that includes a first set of instructions for ML models 108, where the first set of instructions includes a prompt, retrieved information regarding the particular application (e.g., intents, application where available, APIs of the particular application, documentation of the particular application, and / or an app store description of the particular application), and other information. Component generator 110 may provide the first set of instructions to a language model, such as an LLM, of ML models 108 as input and receive a second set of instructions as output, where the second set of instructions include code for one or more software components in addition to other information (e.g., instructions to package the code into an executable, an API, etc.). In an example, component generator 110 generates a first set of instructions for a language model of ML models 108 that includes information regarding the particular application, requirements for the software component, and other information. Component generator 110 provides the first set of instructions to the language model as input and receives, as output from the language model, a second set of instructions that include code of a nano-application and instructions for packaging the code into anexecutable.

[0048] Action module 120 may cause the particular application of application 116 to perform the particular action. Action module 120 may use the one or more software components generated by component generator 110 to cause the particular application to perform the particular action. For example, action module 120 may use an API generated by component generator 110 to interact with an application 116 and cause application 116 to perform an action. Action module 120 may cause applications 116 to perform one or more actions, such as generating a calendar reminder, updating information maintained by an application of applications 116, causing an application of applications 116 to generate one or more visual indicators for display, and / or other actions. In some examples, action module 120 may cause multiple of applications 116 to each perform one or more actions.

[0049] Action module 120 may generate GUI 118B as including a visual indicator that action module 120 has caused applications 116 to perform one or more actions. Action module 120 may generate a visual indicator and / or one or more visual elements that include a text and / or auditory summary of the actions. Action module 120 may cause UIC 104 to output the visual indicator and / or elements for display visually located within GUI 118B. In the example, of FIG. 1, GUI 118B includes action summary 124 that includes the text “Said by Assistant: I completed a couple of actions from your call. I created a calendar event to meet with John and Jane at the restaurant on Thursday at 9 PM, and updated your biking app with the new route for the group bike ride on Saturday.” Action module 120 may cause one or more components of computing device 102 such as a receiver to generate audio based on the text of action summary 124. In an example, action module 120 causes a receiver of computing device 102 to generate audio of “I completed an action based on your call. Your calendar is now updated with the time and location of your next doctor’s appointment.”

[0050] The techniques of this disclosure may provide one or more technical benefits. The generation of actions based on calls and messages may enable a user to avoid having to take notes or try to remember what was discussed during a call or message session. In addition, the dynamic identification of an application to perform a particular action may enable a computing device to identify a relevant application. Further, the dynamic generation of software components may enable a computing device to cause applications to perform actions based on a call or message without requiring user input or for the application to natively support such interaction by the computing device.

[0051] FIG. 2 is a block diagram illustrating an example computing device, in accordance with one or more aspects of the present disclosure. Computing device 202 may be similar tocomputing device 202 as illustrated in FIG. 1, include similar components, and provide similar functionality. For example, computing device 202 may be a smartphone that provides a number of different functions such as enabling phone calls and the exchange of messages.

[0052] Computing device 202 includes one or more of processors 226 that may implement functionality and / or execute instructions within computing device 202. Processors 226 may include one or more processors such as mobile processors, desktop processors, integrated processors, reduced instruction set computer (RISC) processors, application processors, display controllers, sensor hubs, and any other hardware configured to function as a processing unit. Processors 226 may execute the instructions of one or more software components of computing device 202. For example, computing device 202 may include one or more of processors 226 that are programmable processors in communication with memory of computing device 202 and that are configured to execute the instructions of one or more software components of computing device 202.

[0053] Computing device 202 includes one or more of communication units 228 that enable computing device 202 to communicate with external devices, such as other computing device 150 as illustrated in FIG. 1, via one or more wired and / or wireless networks by transmitting and / or receiving network signals on the one or more networks. Communication units 228 may include one or more communication components such as network interface cards (e.g., an Ethernet card), optical transceivers, radio frequency transceivers (e.g., satellite transceiver), GPS receivers, or any other types of devices that can send and / or receive information. Other examples of one or more communication units 228 may include short wave radios, cellular data radios, wireless network radios, as well as universal serial bus (USB) controllers. Communication units 288 may enable computing device 202 to communicate with other computing devices and / or systems via a network such as network 140 as illustrated in FIG. 1.

[0054] Computing device 202 includes power source 234 that provides power to computing device 202. Power source 234 may include one or more sources of power such as batteries, connection to an electrical grid, connection to an automotive electrical system, solar cells, and / or other sources of power. For example, power source 234 may include a lithium-ion battery and an occasional connection to an electrical grid.

[0055] Computing device 202 includes one or more of user interface components (UIC) 204 that may be hardware that functions as an input and / or output device for computing device 202. For example, UIC 204 may include a display component, which may be a screen at which information is displayed by UIC 204 and a presence-sensitive input component thatmay detect an object at and / or near the display component. UIC 204 may enable a user to interact with computing device 202 by providing input and receiving output via one or more components of UIC 204.

[0056] UIC 204 includes output components 230 and input components 232. Output components 230 include one or more components capable of generating output such as displays, speakers, visual indicators (e.g., LED indicators), haptic engines, and / or other types of components. Input components 232 include one or more components capable of generating input such as touchscreens, keyboards, mice, microphones, video cameras, motion and position trackers, accelerometers, temperature sensors, heart-rate monitors, and / or other types of components capable of generating input. In an example, a user of computing device 202 touches a location on a touch screen of input components 232 to cause computing device 202 to call another individual. Computing device 202 connects the call and outputs audio of the call via a speaker of output components 230.

[0057] Computing device 202 includes one or more of communication channels 236 (illustrated as “COMM. CHANNEL(S) 236” in FIG. 2). Communication channels 236 may include one or more software and / or hardware components that logically and / or physically interconnect one or more components of computing device 202. For example, communication channels 236 may include a hardware interconnect between processors 226 and storage components 224 that enables the communication of software component instructions from storage components 224 to processors 226 for execution.

[0058] Computing device 202 includes one or more of storage components 224. Storage components 224 may include one or more types of storage such as hard disk drives, solid state drives (e.g., SATA drives, NVMe drives, eMMC storage, etc.), magnetic tape drives, remote storage (e.g., cloud storage), and / or other types of storage. Storage components 224 may store information such as instructions and / or other data of software components of computing device 202 such as an operating system of computing device 202. For example, storage components 224 may include a non-transitory computer-readable storage medium encoded with instructions that, when executed, cause one or more of processors 226 to perform actions of one or more software components stored by storage components 224. Storage components 224 may include a computer program product that includes instructions that cause processors 226 to perform one or more actions of the instructions. For example, storage components 224 may include an external flash drive that includes the instructions of one or more software components of storage components 224.

[0059] Storage components 224 include OS 218. OS 218 may be one or more types of operating system such as a mobile, desktop, server, automotive, or other type of operating system. In some examples, OS 218 may be an operating system of a VM. OS 218 may provide an execution environment for one or more software components of computing device 202 such as UI module 206.

[0060] Storage components 224 include UI module 206. UI module 206 may be a software component of computing device 202 that interprets inputs received by input components 232 and facilitates outputs by output components 230. For example, UI module 206 may act as an intermediary between the one or more associated platforms, operating systems, applications, and / or services executing at computing device 202 and input components 232 and output components 230 of computing device 202.

[0061] Storage components 224 include messaging application 212. Messaging application 212 may be one or more types of application that facilitate the exchange of messages between computing device 202 and other computing devices / sy stems. Messaging application 212 may facilitate the exchange of messages that include text, audio clips, video clips, images (e.g., photos, GIFs, stickers, emoticons, etc.), and other types of messages and messaging content. For example, messaging application 212 may enable a user of computing device 202 to send a multimodal message that includes video and overlaid text to another individual.

[0062] Storage components 224 include phone application 214. Phone application 214 may be an application that enables a user of computing device 202 to place and receive phone calls from other individuals and computing devices / sy stems. For example, phone application 214 may enable a user of computing device 202 to place a call to another computing associated with another individual.

[0063] Storage components 224 include applications 216A-N (hereinafter “applications 216”). Applications 216 may include one or more applications executed by computing device 202 that provide various types of functionalities. For example, applications 216 may include an application that tracks various fitness metrics of a user of computing device 202 and provides running route mapping functionality (e.g., tracking the distance and locations of runs by the user). Applications 216 may be obtained from one or more sources such as application store 238.

[0064] Application store 238 may be an application or other type of executable that may enable a user of computing device 202 to obtain and install applications such as applications 216. In an example, computing device 202 determines that a user has interacted with a GUI of application store 238 and selected application 216A to be installed. Application store 238causes computing device 202 to obtain data of application 216A via communication units 228. Application store 238 causes computing device 202 to install application 216A in storage components 224.

[0065] Storage components 224 include calendar application 240. Calendar application 240 may be an application executed by computing device 202 that provides functionality relating to maintaining a calendar of one or more users of computing device 202. For example, calendar application 240 may enable one or more users of computing device 202 to create and record events and reminders.

[0066] Storage components 224 include action module 220. Action module 220 may be a software component such as process, program, plugin, executable, and / or other type of software component. Action module 220 extracts user interaction data from phone calls and message(s) communicated by computing device 202 only when having received explicit user approval. In an example, a user of computing device 202 initiates a phone call with another individual using phone application 214. Action module 220 generates a GUI that includes a visual indicator offering to create actions based on the call (e.g., approval element 122 as illustrated in FIG. 1) and causes output components 230 to output the GUI. Action module 220 may determine that the user has approved the extraction of user interaction data. In an example, input components 232 receive user input consistent with the user selecting a visual element labeled as “CREATE ACTIONS FROM CALL” during a phone call. UI module 206 interprets the user input and provides an indication to action module 220 that the user has selected the visual element.

[0067] Responsive to receiving the indication, action module 220 obtains audio data of the phone call and extracts user interaction data from the audio data. In some examples, action module 220 may preprocess the audio data into text before extracting user interaction data. Action module 220 may extract user interaction data that includes one or more types of data such as times and dates, locations, future activities that a user is going to perform, upcoming events, information about the user and other individuals, names, particular applications (e.g., a mention of using a particular application during a call), and other information. In some examples, action module 220 may use one or more ML models, such as ML models 208, to extract the user interaction data from the phone call.

[0068] Storage components 224 include one or more of ML models 208. ML models 208 may include one or more types of ML models such as supervised and / or unsupervised ML models, neural networks, models with one or more layers of perceptrons, deep learning models, Q-learning models, language models such as LLMs, and / or other types of MLmodels. In some examples, at least a portion of ML models 208 may be located and executed by a computing device / system exterior to that of computing device 202. In such examples, computing device 202 may securely communicate information to the exterior device for processing and receive the output of ML models 208 from the exterior device. ML models 208 may receive user interaction data from action module 220 and output indications of particular actions.

[0069] In some implementations action module 220 and / or ML models 208 may preprocess natural user interaction data, such as the data obtained from messaging application 212 and / or data obtained from phone application 215 among other software components. In some examples, a first set of instructions stored in an instruction storage may be preprocessed. Preprocessing techniques may include extracting one or more additional features from raw data. For example, feature extraction techniques may be applied to the obtained data or retrieved instructions to generate one or more new, additional features.

[0070] As described herein, ML models 208 may employ a large language model (LLM) that can interpret natural language user input and generate user interaction data. ML models 208 may perform various types of natural language processing (NLP) based on the indication of the natural language user input. The indication of the natural language user input (e.g., the user interaction data) and / or the retrieved first set of instructions (e.g., a prompt generated by action module 220) may be referred to herein as “input data”. For example, ML models 208 may summarize, translate, or organize the input data. ML models 208 may use recurrent neural networks (RNNs) and / or transformer models (self-attention models), such as GPT-3, GPT-4, BERT, Gemini, and T5. In some implementations, ML models 208 may perform classification, summarization, name generation, regression, clustering, anomaly detection, recommendation generation, and / or other tasks.

[0071] In some implementations, ML models 208 may perform various types of classification based on the input data. For example, ML models 208 may perform binary classification or multiclass classification. In binary classification, the output data may include a classification of the input data into one of two different classes. In multiclass classification, the output data may include a classification of the input data into one (or more) of more than two classes. The classifications may be single-label or multi-label. ML models 208 may perform discrete categorical classification in which the input data is simply classified into one or more classes or categories.

[0072] In some implementations, ML models 208 can perform classification in which ML models 208 provide, for each of one or more classes, a numerical value descriptive of adegree to which it is believed that the input data should be classified into the corresponding class. In some instances, the numerical values provided by ML models 208 can be referred to as “confidence scores” that are indicative of a respective confidence associated with classification of the input into the respective class. In some implementations, the confidence scores can be compared to one or more thresholds to render a discrete categorical prediction. In some implementations, only a certain number of classes (e.g., one) with the relatively largest confidence scores can be selected to render a discrete categorical prediction.

[0073] ML models 208 may output a probabilistic classification. For example, ML models 208 may predict, given a sample input, a probability distribution over a set of classes. Thus, rather than outputting only the most likely class to which the sample input should belong, ML models 208 can output, for each class, a probability that the sample input belongs to such class. In some implementations, the probability distribution over all possible classes can sum to one. In some implementations, a Softmax function, or other type of function or layer can be used to squash a set of real values respectively associated with the possible classes to a set of real values in the range (0, 1) that sum to one.

[0074] In some examples, the probabilities provided by the probability distribution can be compared to one or more thresholds to render a discrete categorical prediction. In some implementations, only a certain number of classes (e.g., one) with the relatively largest predicted probability can be selected to render a discrete categorical prediction.

[0075] In cases in which ML models 208 performs classification, ML models 208 may be trained using supervised learning techniques. For example, ML models 208 may be trained on a training dataset that includes training examples labeled as belonging (or not belonging) to one or more classes.

[0076] In some implementations, ML models 208 may perform regression to provide output data in the form of a continuous numeric value. The continuous numeric value may correspond to any number of different metrics or numeric representations, including, for example, currency values, scores, or other numeric representations. In examples, ML models 208 may perform linear regression, polynomial regression, or nonlinear regression. In examples, ML models 208 may perform simple regression or multiple regression. As described above, in some implementations, a Softmax function or other function or layer may be used to squash a set of real values respectively associated with two or more possible classes to a set of real values in the range (0, 1) that sum to one.

[0077] ML models 208 may perform various types of clustering. For example, ML models 208 may identify one or more clusters to which the input data most likely corresponds. MLmodels 208 may identify one or more clusters within the input data. That is, in instances in which the input data includes multiple objects, documents, or other entities, ML models 208 may sort the multiple entities included in the input data into a number of clusters. In some implementations in which ML models 208 performs clustering, ML models 208 may be trained using unsupervised learning techniques.

[0078] ML models 208 may perform anomaly detection or outlier detection. For example, ML models 208 can identify input data that does not conform to an expected pattern or other characteristic (e.g., as previously observed from previous input data). As examples, the anomaly detection can be used for fraud detection or system failure detection.

[0079] ML models 208 may, in some cases, act as an agent within an environment. For example, ML models 208 may be trained using reinforcement learning, which will be discussed in further detail below.

[0080] In some implementations, ML models 208 may include a parametric model while, in other implementations, ML models 208 may include a non-parametric model. In some implementations, ML models 208 may include a linear model while, in other implementations, ML models 208 may include a non-linear model.

[0081] As described above, ML models 208 may be or include one or more of various different types of machine-learned models. Examples of such different types of machine- learned models are provided below for illustration. One or more of the example models described below may be used (e.g., combined) to provide the output data in response to the input data. Additional models beyond the example models provided below may be used as well.

[0082] In some implementations, ML models 208 may be or include one or more classifier models such as, for example, linear classification models; quadratic classification models; etc. ML models 208 may be or include one or more regression models such as, for example, simple linear regression models; multiple linear regression models; logistic regression models; stepwise regression models; multivariate adaptive regression splines; locally estimated scatterplot smoothing models; etc. scatterplot smoothing models; etc.

[0083] In some examples, ML models 208 can be or include one or more decision tree-based models such as, for example, classification and / or regression trees; iterative dichotomiser 3 decision trees; C4.5 decision trees; chi-squared automatic interaction detection decision trees; decision stumps; conditional decision trees; etc.

[0084] ML models 208 may be or include one or more kernel machines. In some implementations, ML models 208 can be or include one or more support vector machines.ML models 208 may be or include one or more instance-based learning models such as, for example, learning vector quantization models; self- organizing map models; locally weighted learning models; etc. In some implementations, ML models 208 can be or include one or more nearest neighbor models such as, for example, k-nearest neighbor classifications models; k- nearest neighbors regression models; etc. ML models 208 can be or include one or more Bayesian models such as, for example, naive Bayes models; Gaussian naive Bayes models; multinomial naive Bayes models; averaged one-dependence estimators; Bayesian networks; Bayesian belief networks; hidden Markov models; etc.

[0085] In some implementations, ML models 208 may be or include one or more artificial neural networks (also referred to simply as neural networks). A neural network may include a group of connected nodes, which also may be referred to as neurons or perceptrons. A neural network may be organized into one or more layers. Neural networks that include multiple layers may be referred to as “deep” networks. A deep network may include an input layer, an output layer, and one or more hidden layers positioned between the input layer and the output layer. The nodes of the neural network may be connected or non-fully connected.

[0086] ML models 208 can be or include one or more feed forward neural networks. In feed forward networks, the connections between nodes do not form a cycle. For example, each connection can connect a node from an earlier layer to a node from a later layer.

[0087] In some instances, ML models 208 can be or include one or more recurrent neural networks. In some instances, at least some of the nodes of a recurrent neural network can form a cycle. Recurrent neural networks can be especially useful for processing input data that is sequential in nature. In particular, in some instances, a recurrent neural network can pass or retain information from a previous portion of the input data sequence to a subsequent portion of the input data sequence through the use of recurrent or directed cyclical node connections.

[0088] In some examples, sequential input data can include time-series data (e.g., sensor data versus time or imagery captured at different times). For example, a recurrent neural network can analyze sensor data versus time to detect or predict a swipe direction, to perform handwriting recognition, etc. Sequential input data may include words in a sentence (e.g., for natural language processing, speech detection or processing, etc.); notes in a musical composition; sequential actions taken by a user (e.g., to detect or predict sequential application usage); sequential object states; etc.

[0089] Example recurrent neural networks include long short-term (LSTM) recurrent neural networks; gated recurrent units; bi-direction recurrent neural networks; continuous timerecurrent neural networks; neural history compressors; echo state networks; Elman networks; Jordan networks; recursive neural networks; Hopfield networks; fully recurrent networks; sequence-to- sequence configurations; etc.

[0090] In some implementations, ML models 208 can be or include one or more convolutional neural networks. In some instances, a convolutional neural network can include one or more convolutional layers that perform convolutions over input data using learned filters.

[0091] Filters can also be referred to as kernels. Convolutional neural networks can be especially useful for vision problems such as when the input data includes imagery such as still images or video. However, convolutional neural networks can also be applied for natural language processing.

[0092] In some examples, ML models 208 may be or include one or more generative networks such as, for example, generative adversarial networks. Generative networks may be used to generate new data such as artificial feedback; texts; content; etc.

[0093] ML models 208 may be or include an autoencoder. In some instances, the aim of an autoencoder is to learn a representation (e.g., a lower- dimensional encoding) for a set of data, typically for the purpose of dimensionality reduction. For example, in some instances, an autoencoder can seek to encode the input data and provide the output data that reconstructs the input data from the encoding. Recently, the autoencoder concept has become more widely used for learning generative models of data. In some instances, the autoencoder can include additional losses beyond reconstructing the input data. ML models 208 may be or include one or more other forms of artificial neural networks such as, for example, deep Boltzmann machines; deep belief networks; stacked autoencoders; etc. Any of the neural networks described herein can be combined (e.g., stacked) to form more complex networks.

[0094] One or more neural networks can be used to provide an embedding based on the input data. For example, the embedding can be a representation of knowledge abstracted from the input data into one or more learned dimensions. In some instances, embeddings can be a useful source for identifying related entities. In some instances, embeddings can be extracted from the output of the network, while in other instances embeddings can be extracted from any hidden node or layer of the network (e.g., a close to final but not final layer of the network). Embeddings can be useful for performing auto suggest next video, product suggestion, entity or object recognition, etc. In some instances, embeddings may be useful inputs for downstream models. For example, embeddings can be useful to generalize input data (e.g., search queries) for a downstream model or processing system.

[0095] ML models 208 may include one or more clustering models such as, for example, k- means clustering models; k-medians clustering models; expectation maximization models; hierarchical clustering models; etc.

[0096] In some implementations, ML models 208 can perform one or more dimensionality reduction techniques such as, for example, principal component analysis; kernel principal component analysis; graph-based kernel principal component analysis; principal component regression; partial least squares regression; Sammon mapping; multidimensional scaling; projection pursuit; linear discriminant analysis; mixture discriminant analysis; quadratic discriminant analysis; generalized discriminant analysis; flexible discriminant analysis; autoencoding; etc.

[0097] In an example in which the input data does not include feature embeddings, one or more neural networks may be used to provide an embedding based on the input data. For example, the embedding may be a representation of knowledge abstracted from the input data into one or more learned dimensions. In some instances, embeddings may be a useful source for identifying related entities. In some instances, embeddings may be extracted from the output of the network, while in other instances embeddings may be extracted from any hidden node or layer of the network (e.g., a close to final but not final layer of the network). Embeddings may be useful for performing auto-suggest next video, product suggestion, entity or object recognition, etc. In some instances, embeddings are useful inputs for downstream models. For example, embeddings may be useful to generalize input data (e.g., search queries) for a downstream model or processing system.

[0098] In some implementations, ML models 208 may perform or be subjected to one or more reinforcement learning techniques such as Markov decision processes; dynamic programming; Q functions or Q-learning; value function approaches; deep Q-networks; differentiable neural computers; asynchronous advantage actor-critics; deterministic policy gradient; etc.

[0099] In some implementations, ML models 208 may be an autoregressive model. In some instances, an autoregressive model may specify that the output data depends linearly on its own previous values and on a stochastic term. In some instances, an autoregressive model may take the form of a stochastic difference equation. One example of an autoregressive model is WaveNet, which is a generative model for raw audio.

[0100] In some implementations, ML models 208 may include or form part of a multiple model ensemble. As one example, bootstrap aggregating may be performed, which may also be referred to as “bagging.” In bootstrap aggregating, a training dataset is split into a numberof subsets (e.g., through random sampling with replacement) and a plurality of models are respectively trained on the number of subsets. At inference time, respective outputs of the plurality of models may be combined (e.g., through averaging, voting, or other techniques) and used as the output of the ensemble.

[0101] One example ensemble is a random forest, which may also be referred to as a random decision forest. Random forests are an ensemble learning method for classification, regression, and other tasks. Random forests are generated by producing a plurality of decision trees at training time. In some instances, at inference time, the class that is the mode of the classes (classification) or the mean prediction (regression) of the individual trees may be used as the output of the forest. Random decision forests may correct for decision trees' tendency to overfit their training set.

[0102] Another example ensemble technique is stacking, which can, in some instances, be referred to as stacked generalization. Stacking includes training a combiner model to blend or otherwise combine the predictions of several other machine-learned models. Thus, a plurality of machine-learned models (e.g., of the same or different type) may be trained based on training data. In addition, a combiner model may be trained to take the predictions from the other machine-learned models as inputs and, in response, produce a final inference or prediction. In some instances, a single-layer logistic regression model may be used as the combiner model.

[0103] Another example of ensemble techniques is boosting. Boosting may include incrementally building an ensemble by iteratively training weak models and then adding to a final strong model. For example, in some instances, each new model may be trained to emphasize the training examples that previous models misinterpreted (e.g., misclassified). For example, a weight associated with each of such misinterpreted examples may be increased. One common implementation of boosting is AdaBoost, which may also be referred to as Adaptive Boosting. Other example boosting techniques include LPBoost; TotalBoost; BrownBoost; xgboost; MadaBoost, LogitBoost, gradient boosting; etc. Furthermore, any of the models described above (e.g., regression models and artificial neural networks) may be combined to form an ensemble. As an example, an ensemble may include a top-level machine-learned model or a heuristic function to combine and / or weight the outputs of the models that form the ensemble.

[0104] In some implementations, multiple machine-learned models (e.g., that form an ensemble may be linked and trained jointly (e.g., through backpropagation of errors sequentially through the model ensemble). However, in some implementations, only a subset(e.g., one) of the jointly trained models is used for inference.

[0105] In some implementations, ML models 208 may be used to preprocess the input data for subsequent input into another model. For example, ML models 208 may perform dimensionality reduction techniques and embeddings (e.g., matrix factorization, principal components analysis, singular value decomposition, word2vec / GLOVE, and / or related approaches); clustering; and even classification and regression for downstream consumption. Many of these techniques have been discussed above and will be further discussed below, inference.

[0106] In some implementations, ML models 208 can be used to preprocess the input data for subsequent input into another model. For example, ML models 208 can perform dimensionality reduction techniques and embeddings (e.g., matrix factorization, principal components analysis, singular value decomposition, word2vec / GLOVE, and / or related approaches); clustering; and even classification and regression for downstream consumption. ML models 208 may preprocess non-textual data (e.g., audio data, image data, multimodal data, etc.) into text data. Many of these techniques have been discussed above and will be further discussed below.

[0107] As discussed above, ML models 208 can be trained or otherwise configured to receive the input data and, in response, provide the output data. The input data can include different types, forms, or variations of input data. As examples, in various implementations, the input data can include features that describe the content (or portion of content) initially selected by the user, e.g., content of user-selected document or image, links pointing to the user selection, links within the user selection relating to other files available on device or cloud, metadata of user selection, etc. Additionally, with user permission, the input data includes the context of user usage, either obtained from an app itself or from other sources. Examples of usage context include breadth of share (sharing publicly, or with a large group, or privately, or a specific person), context of share, etc. When permitted by the user, additional input data can include the state of the device, e.g., the location of the device, the apps running on the device, etc.

[0108] In some implementations, ML models 208 can receive and use the input data in its raw form. In some implementations, the raw input data can be preprocessed. Thus, in addition or alternatively to the raw input data, ML models 208 can receive and use the preprocessed input data.

[0109] In some implementations, preprocessing the input data can include extracting one or more additional features from the raw input data. For example, feature extraction techniquescan be applied to the input data to generate one or more new, additional features. Example feature extraction techniques include edge detection; comer detection; blob detection; ridge detection; scale-invariant feature transform; motion detection; optical flow; Hough transform; etc.

[0110] In some implementations, the extracted features can include or be derived from transformations of the input data into other domains and / or dimensions. As an example, the extracted features can include or be derived from transformations of the input data into the frequency domain. For example, wavelet transformations and / or fast Fourier transforms can be performed on the input data to generate additional features.[OHl] In some implementations, the extracted features can include statistics calculated from the input data or certain portions or dimensions of the input data. Example statistics include the mode, mean, maximum, minimum, or other metrics of the input data or portions thereof.

[0112] In some implementations, as described above, the input data can be sequential in nature. In some instances, the sequential input data can be generated by sampling or otherwise segmenting a stream of input data. As one example, frames can be extracted from a video. In some implementations, sequential data can be made non- sequent! al through summarization.

[0113] As another example preprocessing technique, portions of the input data can be imputed. For example, additional synthetic input data can be generated through interpolation and / or extrapolation.

[0114] As another example preprocessing technique, some or all of the input data can be scaled, standardized, normalized, generalized, and / or regularized. Example regularization techniques include ridge regression; least absolute shrinkage and selection operator (LASSO); elastic net; least-angle regression; cross-validation; LI regularization; L2 regularization; etc. As one example, some or all of the input data can be normalized by subtracting the mean across a given dimension’s feature values from each individual feature value and then dividing by the standard deviation or other metric.

[0115] As another example preprocessing technique, some or all or the input data can be quantized or discretized. In some cases, qualitative features or variables included in the input data can be converted to quantitative features or variables. For example, one hot encoding can be performed.

[0116] In some examples, dimensionality reduction techniques can be applied to the input data prior to input into ML models 208. Several examples of dimensionality reduction techniques are provided above, including, for example, principal component analysis; kernelprincipal component analysis; graph-based kernel principal component analysis; principal component regression; partial least squares regression; Sammon mapping; multidimensional scaling; projection pursuit; linear discriminant analysis; mixture discriminant analysis; quadratic discriminant analysis; generalized discriminant analysis; flexible discriminant analysis; autoencoding; etc.

[0117] In some implementations, during training, the input data can be intentionally deformed in any number of ways to increase model robustness, generalization, or other qualities. Example techniques to deform the input data include adding noise; changing color, shade, or hue; magnification; segmentation; amplification; etc.

[0118] In response to receipt of the input data, ML models 208 can provide the output data. The output data can include different types, forms, or variations of output data. As examples, in various implementations, the output data can include content, either stored locally on the user device or in the cloud, that is relevantly shareable along with the initial content selection.

[0119] As discussed above, in some implementations, the output data can include various types of classification data (e.g., binary classification, multiclass classification, single label, multi- label, discrete classification, regressive classification, probabilistic classification, etc.) or can include various types of regressive data (e.g., linear regression, polynomial regression, nonlinear regression, simple regression, multiple regression, etc.). In other instances, the output data can include clustering data, anomaly detection data, recommendation data, or any of the other forms of output data discussed above.

[0120] In some implementations, the output data can influence downstream processes or decision making. As one example, in some implementations, the output data can be interpreted and / or acted upon by a rules-based regulator.

[0121] In some implementations, during training, the input data may be intentionally deformed in any number of ways to increase model robustness, generalization, or other qualities. Example techniques to deform the input data include adding noise; changing color, shade, or hue; magnification; segmentation; amplification; etc.

[0122] In response to receipt of the input data, ML models 208 may provide the output data. As examples, in various implementations, the output data may include content, either stored locally on the user device or in the cloud, that is relevantly shareable along with the initial content selection.

[0123] In some implementations, the output data may influence downstream processes or decision-making. As one example, in some implementations, the output data, or the second set of instructions, may be interpreted and / or acted upon by a rules-based regulator.

[0124] The techniques of the present disclosure may be implemented by or otherwise executed on one or more computing devices (e.g., computing device 102 of FIG. 1). Examples of such computing devices include user computing devices (e.g., laptops, desktops, and mobile computing devices such as tablets, smartphones, wearable computing devices, etc.); embedded computing devices (e.g., devices embedded within a vehicle, camera, image sensor, industrial machine, satellite, gaming console or controller, or home appliance such as a refrigerator, thermostat, energy meter, home energy manager, smart home assistant, etc.); other computing devices; or combinations thereof. Computing device 202 that implements ML models 208 or other aspects of the present disclosure may include a number of hardware components that enable the performance of the techniques described herein.

[0125] ML models 208 described herein may be trained according to one or more of various different training types or techniques. For example, in some implementations, ML models 208 may be trained using supervised learning, in which ML models 208 is trained on a training dataset that includes instances or examples that have labels. The labels may be manually applied by experts, generated through crowdsourcing, or provided by other techniques (e.g., by physics-based or complex mathematical models). In some implementations, if the user has provided consent, the training examples may be provided by the user computing device. In some implementations, this process may be referred to as personalizing the model.

[0126] In some implementations, backward propagation of errors may be used in conjunction with an optimization technique (e.g., gradient-based techniques) to train ML models 208 (e.g., when the machine-learned model is a multi-layer model such as an artificial neural network). For example, an iterative cycle of propagation and model parameter (e.g., weights) update may be performed to train ML models 208. Example backpropagation techniques include truncated backpropagation through time, Levenberg- Marquardt backpropagation, etc.

[0127] In some implementations, ML models 208 described herein may be trained using unsupervised learning techniques. Unsupervised learning may include inferring a function to describe hidden structure from unlabeled data. For example, a classification or categorization may not be included in the data. Unsupervised learning techniques may be used to produce machine-learned models capable of performing clustering, anomaly detection, learning latent variable models, or other tasks.

[0128] ML models 208 may be trained using semi-supervised techniques which combine aspects of supervised learning and unsupervised learning. ML models 208 may be trained orotherwise generated through evolutionary techniques or genetic algorithms. In some implementations, ML models 208 described herein may be trained using reinforcement learning. In reinforcement learning, an agent (e.g., model) may take actions in an environment and learn to maximize rewards and / or minimize penalties that result from such actions. Reinforcement learning may differ from the supervised learning problem in that correct input / output pairs are not presented, nor sub-optimal actions explicitly corrected.

[0129] In some implementations, one or more generalization techniques may be performed during training to improve the generalization of ML models 208. Generalization techniques may help reduce overfitting of ML models 208 to the training data. Example generalization techniques include dropout techniques; weight decay techniques; batch normalization; early stopping; subset selection; stepwise selection; label smoothing; etc.

[0130] In some implementations, ML models 208 described herein may include or otherwise be impacted by a number of hyperparameters, such as, for example, learning rate, number of layers, number of nodes in each layer, number of leaves in a tree, number of clusters; etc. Hyperparameters may affect model performance. Hyperparameters may be hand selected or may be automatically selected through the application of techniques such as, for example, grid search; black-box optimization techniques (e.g., Bayesian optimization, random search, etc.); gradient-based optimization; etc. Example techniques and / or tools for performing automatic hyperparameter optimization include Hyperopt; Auto-WEKA; Spearmint; Metric Optimization Engine (MOE); etc.

[0131] In some implementations, various techniques may be used to optimize and / or adapt the learning rate when the model is trained. Example techniques and / or tools for performing learning rate optimization or adaptation include Adagrad; Adaptive Moment Estimation (ADAM); Adadelta; RMSprop; etc.

[0132] In some implementations, transfer learning techniques may be used to provide an initial model from which to begin training of ML models 208 described herein.

[0133] In some implementations, ML models 208 described herein may be included in different portions of computer-readable code on a computing device. In one example, ML models 208 may be included in a particular application or program and used (e.g., exclusively) by such particular application or program. Thus, in one example, a computing device may include a number of applications, and one or more of such applications may contain its own respective machine learning library and machine-learned model(s).

[0134] In another example, ML models 208 described herein may be included in an operating system of a computing device (e.g., in a central intelligence layer of an operating system) andmay be called or otherwise used by one or more applications that interact with the operating system. In some implementations, each application may communicate with the central intelligence layer (and model(s) stored therein) using an application programming interface (API) (e.g., a common, public API across all applications).

[0135] In some implementations, the central intelligence layer may communicate with a central device data layer. The central device data layer may be a centralized repository of data for the computing device. The central device data layer may communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer may communicate with each device component using an API (e.g., a private API).

[0136] The technology discussed herein refers to servers, databases, software applications, and other computer-based systems, as well as actions taken, and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a great variety of possible configurations, combinations, and divisions of tasks and functionality between and among components. For instance, processes discussed herein may be implemented using a single device or component or multiple devices or components working in combination.

[0137] Databases and applications may be implemented on a single system or distributed across multiple systems. Distributed components may operate sequentially or in parallel.

[0138] In addition, the machine learning techniques described herein are readily interchangeable and combinable. Although certain example techniques have been described, many others exist and may be used in conjunction with aspects of the present disclosure.

[0139] In some implementations, transfer learning (TL) may be used. Transfer learning involves reusing a model and its model parameters obtained while solving one problem and applying it to a different but related problem. Models trained on very large data sets may be retrained or fine-tuned on additional data. Often, all model designs and their parameters on a source model are copied except output layer(s). The output layers(s) are often called the head, and other layers are often called the base. The source parameters may be considered to contain the knowledge learned from the source dataset and this knowledge may also be applicable to a target dataset. Fine-tuning may include updating the head parameters with the body parameters being fixed or updated in a later step.

[0140] Thus, ML models 208 may apply a large language model, in combination with or not in combination with one or more of the machine learning techniques described above, to the input data to generate user interaction data. ML models 208 may provide the user interactiondata to action module 220 as output.

[0141] Action module 220 may determine one or more particular actions to perform based on the user interaction action data. Action module 220 may determine particular actions that are actions for one or more of applications 216, such as updating information in the application of applications 216, causing the application to generate a particular output, and / or other actions for the application to take. Action module 220 may identify a single action or multiple actions to perform. For example, action module 220 may identify five actions to be performed by one or more of applications 216.

[0142] Action module 220 may use ML models 208 to determine the particular action to take by providing the user interaction data to ML models 208. Action module 220 may pre- process or otherwise prepare the user interaction data for ML models 208 (e.g., generate a prompt to accompany the user interaction data, edit the user interaction data to remove unnecessary or irrelevant information, etc.) or may provide unprocessed user interaction data directly to ML models 208. For example, action module 220 may provide user interaction data that includes audio data of a phone call to a language model of ML models 208 and receive an indication of a particular action as an output of the language model. Action module 220 may provide the user interaction data to ML models 208 and receive an indication of one or more actions from ML models 208. In an example, action module 220 obtains audio data of a phone call from phone application 214 in response to user approval and uses ML models 208 to extract user interaction data from the audio data. Action module 220 provides the user interaction data along with a first set of instructions (e.g., a prompt) and information regarding applications 216 to ML models 208 and receives an indication of a particular action from ML models 208. Action module 220 may use ML models 208 to predict the particular action based on the user interaction data. In an example, action module 220 streams the user interaction data to ML models 208 during an ongoing call or while messages are being sent and received. ML models 208 predict the particular action prior to the completion of the call or while messages are still being sent / received.

[0143] In some examples, action module 220 may use contextual information to identify the particular action. Action module 220 may use contextual information that includes information regarding contextual information from computing device 202, such as time and date, location data, identity of a caller, identity of an entity that is exchanging messages with computing device 202, and other information. Action module 220 may store contextual information in app data 222 and / or determine the contextual information in real-time. Action module 220 may use the contextual information to identify the particular action and resolveambiguities in the user interaction data. Action module 220 may use the contextual information to resolve ambiguities, such as when there are multiple applications of the same type (e.g., multiple calendar applications), when there is no application of a particular type (e.g., no calendar application), when a duration is not explicitly indicated (e.g., assume that an appointment will last 30 minutes, action module 220 utilizes heuristics to determine a duration, etc.), relative mentions (e.g., “tomorrow, next week, etc.”), incomplete information (e.g, using a default time when user interaction data does not include the date for an action), and other ambiguities (e.g., “Sunday”, action module 208 may assume an upcoming Sunday). Action module 220 may resolve one or more of the ambiguities and determine the particular action.

[0144] In some examples, action module 220 may determine the particular from one or more categories or types of actions. Action module 220 may maintain a list of known types of actions that includes types of actions such as calendar actions, reminder actions, web search actions, contacts action, email actions, phone actions (calling, texting, etc.), and checklist actions among other actions. Action module 220 may compare the user interaction data to the types of actions as part of determining the particular action. For example, action module 220 may compare the user interaction data to the types of actions and determine the particular action will be an action from the web search category.

[0145] Action module 220 may dynamically identify a particular application from applications 216 and / or other applications executed by computing device 202 based on the particular action. Action module 220 may dynamically identify the particular application using information about applications stored in app data 222. For example, action module 220 may compare information regarding the particular action (e.g., a type of application that could perform the particular identified by ML models 208) to information in app data 222 to determine the particular application.

[0146] Storage components 224 include app data 222. App data 222 may include data storage such as a database or data repository that maintains information regarding applications executed by computing device 202, such as applications 216. Action module 220 may generate and / or obtain data regarding the applications from a variety of sources, such application store descriptions of the applications (e.g., information from application store 238, information such as information regarding the most popular applications available from application store 238, etc.), application intents, pre-training data of ML models 208 (e.g., from a corpus or other set of training data that included information regarding the application), information extracted from APIs of the applications (e.g., inferring what actionsan application may perform based on an associated API), and / or a list of the applications and what actions each application may perform (e.g., a pre-populated list from one or more sources, such as developer of action module 220 and / or OS 218). Action module 220 may determine a set of functionality that a new application may perform in response to the installation of a new application of applications 116. In an example, action module 220 determines that a new application of applications 116 has been installed by computing device 202. Responsive to the installation, action module 220 determines the set of functionality that the new application of applications 116 may perform (e.g., the actions that the new application can perform). Action module 220, based on dynamically identifying the particular application, may cause component generator 210 to generate one or more software components.

[0147] Storage components 224 include component generator 210. Component generator 210 may be a software module, plugin, process, executable, and / or other type of software component. Component generator 210 may dynamically generate software components, such as APIs, executables, nano-applications, plugins, and other types of software components to enable action module 220 to cause a particular application to perform a particular action. Component generator 210 may use one or more ML models, such as a language model of ML models 208, to generate the software component. In an example, component generator 210 generates a prompt or other type of request for a language model of ML models 208 to dynamically generate an API. Component generator 210 provides the prompt to the language model as input and receives code of the API as output.

[0148] Component generator 210 may use one or more of ML models 208 to create instructions for dynamically generating software components. ML models 208 may apply a large language model, in combination with or not in combination with one or more of the machine learning techniques described above, to input data that includes an indication of the particular application, an indication of the particular action for the particular action to perform, and a first set of instructions (e.g., a prompt that includes requirements for generating software components) among other information to generate the second set of instructions, wherein the second set of instructions include instructions for generating the software components. Components generator 210 may use the second set of instructions as a basis to dynamically generate the software components (e.g., use another of ML models 208 to generate the software components, use one or more other components to implement the second set of instructions to dynamically generate the software components, etc.).

[0149] In some examples, component generator 210 may use one or more of ML models 208to dynamically generate the software components themselves. Similarly to the above, ML models 208 may apply an LLM to input data that includes a prompt for generating software instructions and an indication of the particular application in addition to other input data. ML models 208 may output a set of instructions of the software components for execution by computing device 202. In an example, component generator 210 provides input data to an LLM of ML models 208 that includes a prompt for generating software components to cause a running application to update a planned group run. The LLM processes the input data and outputs source code for a nano-application in XML format for computing device 202 to execute the plugin. Action module 220 processes the source code for the nano-application and causes processors 226 to execute the instructions of the nano-application. The nanoapplication executes and causes the running application to update the planned group run.

[0150] In some examples, component generator 210 may generate software components as adhering to one or more requirements. Component generator 210 may generate the software components as adhering to the one or more requirements to ensure consistent functionality when the actions are performed. Component generator 210 may generate the software components as adhering to requirements that include formatting phone numbers into particular formats (e.g., (XXX) XXX-XXXX even when text recognition software provide the numbers as text), validating the number of characters in a phone number (e.g., ensuring that a phone number has the correct number of characters, extracting the first three digits from a user’s area code to complete a number, etc.), formatting email addresses according to predetermined format (e.g., local-part@domain.com), validating an email address domain (e.g., ensuring that a commonly used domain is correctly spelled, correcting “emil” to “email”, adding a domain where missing, etc.), and / or other requirements. In some examples, component generator 210 may include the one or more requirements in a prompt to ML models 208 as part of a prompt to generate the second set of instructions for generating the software components.

[0151] Component generator 210 may generate software components that enable a user of computing device 202 to specify one or more further actions that they would like to take. Component generator 210 may generate software components that enable a user to specify further actions, such as generating a GUI with a phone number highlighted and including visual elements corresponding further actions that the user may select (e.g., save the phone number as contact, to call the number, text the number, etc.), generating a GUI that includes an email address highlighted and as including visual elements corresponding further actions that the user may select (e.g., sending an email to the highlighted address, modifying theemail, etc.), and generating a GUI that includes a reminder with a line describing a task determined by action module 220 and including an area for the user to add more information, among other further actions.

[0152] Component generator 210 may generate software components that provide additional and / or simplified functionality for a user of computing device 202. Component generator 210 may generate software components that interact with applications 216 and enable additional functionality, such as enabling the creation of a group run in a running application that is tied to a calendar event in calendar application 240. In addition, component generator 210 may generate software components that simplify the functionality of applications 216. For example, component generator 210 may generate a software component that causes a health monitoring application of applications 216 to provide an option to autofill one or more fields from a selection of information.

[0153] Action module 220 causes one or more applications, such as the particular application, to perform the particular action using the software component generated by component generator 210. Action module 220 may execute the software component and / or use the software component to interact with the particular application to cause the particular application to perform the particular application. In an example, action module 220 uses an API and an executable generated by component generator to interact with calendar application 240 and cause calendar application to generate an event that includes a name of the event, a list of attendees, and a location.

[0154] FIG. 3 is a flowchart illustrating example operations performed by a computing device that determines actions and dynamically generates software components, in accordance with one or more aspects of the present disclosure. For the purposes of clarity, FIG. 3 is discussed in the context of FIG. 1.

[0155] A computing device, such as computing device 102, extracts user interaction data from one of a phone call or a message (302). Computing device 102 extracts user interaction data in response to obtaining explicit user approval for the data extraction. Computing device 102 may extract user interaction data that may include audio data, text data, and / or other multi-modal data using a software component such as action module 120. Computing device 102 may extract the user interaction data by preprocessing non-textual data (e.g., audio data, image data, etc.) into text and providing the text as input to one or more of ML models 108. Computing device 102 may receive the user interaction data as output from ML models 108. In a particular example, computing device 102 receives user approval to extract user interaction data from a phone call. Computing device 102 obtains audio data of a user ofcomputing device 102 discussing a change in plans for an upcoming group bike ride with another individual. Computing device 102 receives explicit user approval from the user to extract user interaction data. Computing device 102 preprocesses the audio data into text and provides the text along with a prompt to a language model, such as an LLM. The LLM processes the prompt and the text to generate user interaction data that includes identified keywords (e.g., an identification that the user discussed biking, time, and a location) and other information. Computing device 102 receives, as output from the LLM, the user interaction data.

[0156] Computing device 102 determines, using a machine learning model such as one or more of ML models 108, a particular action to perform based on the user interaction data. Computing device 102 may provide the user interaction data to ML models 108 as input and receive an indication of the particular application as output from ML models 108. Computing device 102 may provide user interaction data that includes information relevant to the identification of the particular action (e.g., particular words and phrases, relationships between the words and phrases, etc.) to ML models 108. Continuing the above example, computing device 102 provides the extracted user interaction data and a prompt requesting the identification of a particular action to a model of ML models 108 as input. The model processes the input and generates an output that indicates that an application should generate a reminder for the upcoming group bike ride that includes the location and time of the group ride. Computing device 102 receives the output from computing device 102.

[0157] Computing device 102 dynamically identifies, based on the particular action, a particular application from the one or more applications executed by computing device 102, such as an application of applications 116 (306). Computing device 102 may identify the particular application using information such as a list of actions that includes information about a respective set of functionality of each application. Computing device 102 may use one or more techniques to identify the particular application, such as comparing information in the user interaction data to information regarding one or more applications (e.g., information regarding application intents, information regarding app functionality obtained from an application store, etc.), providing input to ML models 108 that includes the user interaction data and a prompt to identify the particular application, determining application functionality by analyzing APIs associated with applications, and / or other techniques. Further continuing the above example, computing device 102 generates a prompt for ML models 108 that includes a request to identify a particular application. Computing device 102 provides the prompt, an indication of the particular action to update a biking application, and informationregarding application 116 to a model of ML models 108 as input. Computing device 102 receives, as output, an indication of a particular biking application from the ML model as output. In some examples, computing device 102 may identify multiple applications and one or more applications for each application of the multiple applications.

[0158] Computing device 102 dynamically generates a software component that enables computing device 102 to cause the particular application to perform the particular action (308). Computing device 102 may use a component such as component generator 110 to generate the software component. Component generator 110 may use one or more of ML models 108, such as an LLM, to generate the software component. Component generator 110 may generate a prompt for ML models 108 that includes a request to generate a software component, information regarding the particular application and particular action (e.g., an identifier of the particular action, information regarding one or more APIs of the particular application, application intents of the particular application, requirements for generating the software component, etc.). Component generator 110 may retrieve instructions or other information from the particular application to enable ML models 108 to determine what software components are necessary for the particular application to perform the particular action and to enable ML models 108 to generate the necessary software components. Further continuing the above example, component generator 110 retrieves a first set of instructions from the biking application and generates a prompt for a model of ML models 108.Component generator 110 provides the prompt and the first set of instructions to the model. Component generator 110 receives, as output from ML models 108, a second set of instructions that includes code for a nano-application and an API. Component generator 110 packages the code for the nano-application into an executable and the code for the API into an API capable of interfacing with the biking application.

[0159] Computing device 102 causes, using the software component, the particular application to perform the action (310). Computing device 102 may use the software component to interface with the particular application and perform the particular action without requiring further input from a user of computing device 102. Completing the above example, computing device 102 causes one or more processors to execute the instructions of the nano-application. The nano-application interfaces with the biking application through the API. The nano-application causes the biking application to generate a reminder prior to the scheduled group ride for the user that includes a location of the group ride.

[0160] FIG. 4 is a flowchart illustrating example operations performed by a computing device that determines actions and dynamically generating software components, inaccordance with one or more aspects of the present disclosure. For the purposes of clarity, FIG. 4 is described in the context of FIG. 1. For example, a computing device such as computing device 102 may determine actions using an action module, action module 120.

[0161] Computing device 102 receives an indication that a user of computing device 102 has opted-in to the functionality of action module 120 (402). Computing device 102 requires that a user affirmatively opt-in prior to enabling any of the functionality of action module 120 to ensure that the user is aware of the scope of the functionality of action module 120 (e.g., to ensure that the user is aware that action module 120 will collect / extract data from the phone calls and messages of the user prior to performing any such data collection / extraction). Computing device 102 may provide the user with a summary (e.g., a visual element that includes textual / graphical information displayed by a display component, audio that includes a spoken version of the summary output by a speaker of computing device 102, etc.) of the functionality of action module 120 and what types of data action module 120 as part of the opt-in step. In an example, action module 120 causes computing device 102 to output a GUI that includes a summary of the functionality of action module 120 and what types of data action module 120 may collect and interactable visual indicators that enable the user to accept or decline the use of action module 120. Based on receiving user interaction consistent with the user accepting the use of action module 120, computing device 102 enables the use of action module 120.

[0162] Computing device 102 receives a call from or initiates a call with another computing device or system (404). Computing device 102 may receive or initiate a call such as a video call or audio call over one or more networks such as a cellular network or wireless internet network (e.g., WIFI network). Computing device 102 may use calling functionality that is integrated with an operating system of computing device 102 and / or an application, such as phone application 114, to facilitate receiving or initiating the call. In some examples, computing device 102 may output one or more indications of the call to the user such as a vibration or audio tone to alert the user that a call is being received or that an outgoing call is being made.

[0163] Computing device 102 establishes the call (406). In the case of computing device 102 receiving an incoming call, computing device 102 may establish the call in response to receiving an indication of accepting the incoming call from the user. Computing device 102 may output one or more indications that there is an ongoing call such as outputting a GUI for display that includes one or more visual indicators (e.g., GUI 118A as illustrated in FIG. 1). Computing device 102 may establish the call in response to an indication that the user ofcomputing device 102 wishes to establish a call with one or more individuals or systems (e.g., a phone system for a call center). In an example, computing device 102 receives user input via an input component of computing device 102 consistent with a user requesting that computing device 102 call another individual. Responsive to receiving the user input, computing device 102 attempts to connect a call between the user and the individual.

[0164] Computing device 102 presents an action opt-in to the user (408). Computing device 102 presents the action opt-in prior to causing action module 120 to extract user interaction data from the call to request explicit user permission to extract user interaction data.Computing device 102 only extracts user interaction data after having received user approval to do so. Computing device 102 may cause action module 120 to generate a GUI that includes an interactable visual indicator (e.g., extraction approval element 122 as illustrated in FIG. 1) to enable the user to indicate that they wish for action module 120 to generate actions for the ongoing call. Computing device 102 may output the GUI via one or more display components. Computing device 102 waits for user approval prior to extracting user interaction data to ensure user privacy and confidentiality and to only extract data from calls when the user so desires.

[0165] Computing device 102 receives user approval (410). Computing device 102 may receive user approval in the form of the user interacting with a visual element, such as extraction approval element 122 or other type of interaction. For example, computing device 102 may receive spoken input of the user approving the extraction of user interaction data and the generation of actions. Computing device 102 may receive the approval for extraction of the user interaction data in real-time and via a GUI of computing device 102 and / or an audio interface of computing device 102.

[0166] Computing device 102 may provide an indication that the call is being recording and that Al is being used (412). Computing device 102 may provide an indication to both the caller and the callee that the call is being recording that and Al is being used to process the call. Computing device 102 may provide the indication to ensure that computing device 102 complies with regulatory requirements regarding the use of call recording and Al processing software. Computing device 102 may provide the indication in only some scenarios and may refrain from the providing the indication in other scenarios.

[0167] Computing device 102 activates action module 120 (413). Computing device 102 may activate action module 120 in response to receiving approval from the user. Action module 120 may analyze the phone call in response to the user interaction approving the phone call. While active and during the call, action module 120 may obtain audio data from the call viaphone application 114. In some examples, action module 120 may provide or stream the audio data as it is obtained from the call to one or more ML models, such as ML models 108, for ML models 108 to predict actions. Action module 120 may preprocess the audio data of the call into text and provide the text to one or more of ML models 108. Action module 120 may use one or more of ML models 108 to extract user interaction data from the text of the call. In some examples, action module 120 may directly provide the audio data of the call to ML models 108 for user interaction data extraction (e.g., when ML models 108 include a model capable of directly processing audio data).

[0168] Computing device 102 displays an indicator (414) to indicate and remind the user that action module 120 is active and extracting user interaction data from the call. Computing device 102 may display an indicator as a visual indicator within a GUI output by computing device 102 and / or using one an indicator, such as an LED indicator on the exterior of computing device 102. In an example, computing device 102 illuminates an LED on the exterior of computing device 102 with a blinking red light to indicate that action module 120 is obtaining the audio data of the call and extracting user interaction data.

[0169] Computing device 102 determines that the call has concluded (416). Computing device 102 may determine that the call has concluded based on an indication from phone application 114. Action module 120 may terminate the collection of user interaction data upon the conclusion of the call. In some examples, action module 120 may cause computing device 102 to transmit the user interaction data to one or more recipients, such as a cloud computing system for off-device analysis of user interaction data when explicitly approved by the user. Action module 120 deletes the user interaction data and audio data of the call upon completion of executing the generated software component to avoid retention of the user interaction data and the audio data.

[0170] Computing device 102 determines a particular action to perform (418). Computing device 102 may cause action module 120 to determine the action using one or more of ML models 108. Action module 120 may provide the user interaction data, information regarding one or more of applications 116, and / or other information to an ML model of ML models 108 in addition to a prompt requesting the identification of the particular action as input. Action module 120 may receive an indication of the action as an output from the ML model. In some examples, action module 120 may provide or stream at least a portion of the user interaction data to ML models 108 while the call is still ongoing for ML models 108 to predict one or more actions. While illustrated as determining a single action in the example of FIG. 4, action module 120 may determine a plurality of actions.

[0171] Computing device dynamically identifies a particular application to perform the particular action (420). Computing device 102 may cause action module 120 to dynamically identify the particular application using one or more techniques, such as applying heuristics to information regarding applications executed by computing device 102, comparing one or more elements of the user interaction data to information regarding the applications, and / or other techniques. Action module 120 may use one or more of ML models 108 to identify the particular application. For example, action module 120 may provide a prompt, an indication of the particular action, and information regarding applications 116 to ML models 108 for ML models 108 to identify the particular application. Action module 120 may dynamically identify one or more applications, such as one or more of applications 116. In some examples, action module 120 may dynamically identify a particular application that is not currently installed on computing device 102 but that is available from an application store.

[0172] Computing device 102 dynamically generates a software component to enable computing device 102 to cause the particular application to perform the particular application (422). Computing device 102 may use one or more components, such as component generator 110, to generate the software component. Component generator 110 may use one or more of ML models 108, such as an LLM, to generate the software component. In addition, component generator 110 may use information regarding the particular application to generate the software component. As part of generating the software component, component generator 110 may retrieve information regarding the particular application, such as application intents, APIs associated with the application, an app store description of the particular application, and other information. Component generator 110 may provide the retrieved information, an indication of the particular application and particular action, a prompt requesting creation of the software component as a first set of instructions to ML models 108. Component generator 110 may receive a second set of instructions from ML models 108 that includes code of the software component and instructions for executing the code as output from ML models 108.

[0173] In some examples, component generator 110 may generate the software component by providing a first set of instructions to ML models 108 and receiving a second set of instructions as output from ML models 108. Component generator 110 may generate and provide a first set of instructions that includes instructions, such as source code, intents, and information regarding, of the particular application to ML models 108 and receive a second set of instructions. Component generator 110 may receive a second set of instructions that includes code for a software component and other instructions (e.g., instructions forpackaging the code into an executable).

[0174] Computing device 102 causes the particular application to perform the particular action (424). Computing device 102 uses the software component generated by component generator 110 to cause the particular application to perform the particular action. Computing device 102 may require user approval prior to causing the particular application to perform the particular action. Computing device 102 may receive approval to cause the particular application to perform the particular action via a graphical user interface (e.g., GUI 118B) and / or an audio interface. In an example, computing device 102 generates an executable configured to cause a health tracking app to provide reminders for a user to periodically record their heart rate. Computing device 102 executes the executable. The executable interfaces with an API of the health tracking application and causes the health tracking app to begin generating reminders for the user to periodically record their heart rate in the health tracking app. Computing device 102 may use action module 120 to enable end-to-end functionality of extracting the user interaction data through causing applications to perform actions. Computing device 102 may execute the instructions of the software component to enable the software component to cause the particular application to perform the particular action.

[0175] FIG. 5 is a flowchart illustrating example operations performed by a computing device that determines actions and dynamically generates software components, in accordance with one or more aspects of the present disclosure. For the purposes of clarity, FIG. 5 is described in the context of FIG. 1. For example, a computing device, such as computing device 102, may determine actions using an action module, such as action module 120.

[0176] Computing device 102 receives an indication that a user of computing device 102 has opted-in to the functionality of action module 120 (502). Computing device 102 requires that a user affirmatively opt-in prior to enabling any of the functionality of action module 120 to ensure that the user is aware of the scope of the functionality of action module 120 (e.g., to ensure that the user is aware that action module 120 will collect / extract data from the phone calls and messages of the user prior to performing any such data collection / extraction). Computing device 102 may provide the user with a summary of the functionality of action module 120 and what types of data action module 120 as part of the opt-in step. In an example, action module 120 causes computing device 102 to output a GUI that includes a summary of the functionality of action module 120 and what types of data action module 120 may collect and interactable visual indicators that enable the user to accept or decline the useof action module 120. Based on receiving user interaction consistent with the user accepting the use of action module 120, computing device 102 enables action module 120.

[0177] Computing device 102 receives a message (504) or receives a selection of a message thread (508). Computing device 102 may receive a message from another computing device such as other computing device 150. Computing device 102 may receive data regarding the message via one or more communication components of computing device 102 and provide the data regarding the message to one or more messaging applications executed by computing device 102, such as messaging application 112. In an example, computing device 102 receives data regarding a message for messaging application 112. Messaging application 112 generates a GUI that includes the text of the message and causes computing device 102 to output the GUI for display via one or more components, such as UIC 104. Computing device 102 may receive an indication of a selection of a message thread by a user computing device 102. In an example, a UI module 106 generates an indication of a user interacting with UIC 104 and selecting a message thread of messaging application 112. Messaging application 112 generates a GUI displaying the message thread and causes UIC 104 to display the GUI.

[0178] Computing device 102 presents an action opt-in to the user (510). Computing device 102 presents the action opt-in prior to causing action module 120 to extract user interaction data from the received message or selected message thread. Computing device 102 may cause action module 120 to generate a GUI that includes an interactable visual indicator (e.g., extraction approval element 122) to enable the user to indicate that they wish for action module 120 to generate actions for the received message or selected message thread. Computing device 102 may output the GUI via UIC 104. Computing device 102 waits for user approval prior to extracting user interaction data to ensure user privacy and confidentiality and to only extract data from messages when the user so desires. Computing device 102 requires explicit user approval prior to extracting user interaction data.

[0179] Computing device 102 receives user approval (512). Computing device 102 may receive user approval in the form of the user interacting with a visual element such as extraction approval element 122 or other type of interaction. For example, computing device 102 may receive spoken input of the user approving the extraction of user interaction data and the generation of actions. In some examples, computing device 102 may provide an indication to one or more members of the call that the call is being recording and / or that Al is being used to process the call.

[0180] Computing device 102 activates action module 120 (513). Computing device 102 may activate action module 120 in response to receiving approval from the user. While active,action module 120 may extract user interaction data while messages are sent and received by computing device 102. For example, action module 120 may extract user interaction data for a particular message thread until a user of computing device 102 deactivates action module 120 and / or until a predetermined period of time has elapsed since the user first approved use of action module 120. Action module 120 may extract user interaction data by providing the text of the messages along with a prompt requesting the extraction to ML models 108. Action module 120 may receive the user interaction data as output from ML models 108.

[0181] Computing device 102 displays an indicator (514) to indicate and remind the user that action module 120 is active and extracting user interaction data from the messages.Computing device 102 may display an indicator as a visual indicator within a GUI output by computing device 102 and / or using one an indicator, such as an LED indicator on the exterior of computing device 192. In an example, computing device 102 illuminates an LED on the exterior of computing device 102 with a blinking red light to indicate that action module 120 is extracting user interaction data.

[0182] Computing device 102 determines a particular action to perform (516). Computing device 102 may cause action module 120 to determine the action using one or more of ML models 108. Action module 120 may provide the user interaction data, information regarding one or more of applications 116, and / or other information to an ML model of ML models 108 in addition to a prompt requesting the identification of the particular action as input. Action module 120 may receive an indication of the action as an output from the ML model. In some examples, action module 120 may provide or stream at least a portion of the user interaction data to ML models 108 while the call is still ongoing for ML models 108 to predict one or more actions. While illustrated as determining a single action in the example of FIG. 5, action module 120 may determine a plurality of actions.

[0183] Computing device 102 dynamically identifies a particular application to perform the particular action (518). Computing device 102 may cause action module 120 to dynamically identify the particular application using one or more techniques, such as applying heuristics to information regarding applications executed by computing device 102, comparing one or more elements of the user interaction data to information regarding the applications, and / or other techniques. Action module 120 may use one or more of ML models 108 to identify the particular application. For example, action module 120 may provide a prompt, an indication of the particular action, and information regarding applications 116 to ML models 108 for ML models 108 to identify the particular application. Action module 120 may dynamically identify one or more applications, such as one or more of applications 116. In someexamples, action module 120 may dynamically identify a particular application that is not currently installed on computing device 102 but that is available from an application store.

[0184] Computing device 102 dynamically generates a software component to enable computing device 102 to cause the particular application to perform the particular application (520). Computing device 102 may use one or more components, such as component generator 110, to generate the software component. Component generator 110 may use one or more of ML models 108, such as an LLM, to generate the software component. In addition, component generator 110 may use information regarding the particular application to generate the software component. As part of generating the software component, component generator 110 may retrieve information regarding the particular application, such as application intents, APIs associated with the application, an app store description of the particular application, and other information. Component generator 110 may provide the retrieved information, an indication of the particular application and particular action, a prompt requesting creation of the software component as a first set of instructions to ML models 108. Component generator 110 may receive a second set of instructions from ML models 108 that includes code of the software component and instructions for executing the code as output from ML models 108.

[0185] In some examples, component generator 110 may generate a second set of instructions based on a first set of instructions retrieved from one or more of application 116. Component generator 110 may retrieve a first set of instructions associated with the particular application that are associated with functions of the particular application. For example, component generator 110 may retrieve a first set of instructions from the particular application that are associated with functionality related to the particular action. Component generator 110 may generate a prompt for one or more of ML models 108 and provide the prompt along with the first set of instructions to ML models 108. Component generator 110 may receive a second set of instructions from ML models 108 that includes code for the software component. In an example, component generator 110 retrieves a first set of instructions from a health monitoring application that include information regarding an API of the health monitoring application in addition to reminder functionality of the health monitoring application. Component generator 110 provides the first set of instructions to an LLM of ML models 108 as input. Component generator 110 receives a second set of instructions from the LLM as output that includes code that, when executed, enables action module 120 to interface with the API of the health monitoring and the reminder functionality of the health monitoring application to cause the health monitoring application to generateperiodic reminders to record the heart rate of the user.

[0186] Computing device 102 causes the particular application to perform the particular action (522). Computing device 102 uses the software component generated by component generator 110 to cause the particular application to perform the particular action. In an example, computing device 102 generates a nano-application to cause a web browser to promptly execute a search for houses within a given area, an API to enable computing device 102 to add the contact information of a realtor to a list of contacts, and an API and executable to cause a calendar application to add a reminder to meet with the realtor at future place and time. Computing device 102 uses the software components to interact with the applications and cause the applications to execute the actions. Computing device 102 may use sets of instructions generated by ML models 108 to cause the particular application to perform the particular action. In an example, computing device 102 receives a second set of instructions from component generator 110 to cause a health monitoring application to generate periodic reminders to measure their heart rate and to record their heart rate in the health monitoring application, where the second set of instructions include code for an API and a nanoapplication. Computing device 102 processes the instructions to generate the API and to package the code for the nano-application into an executable. Computing device 102 causes processors of computing device 102 to execute the instructions of the nano-application. The nano-application uses the API to interface with the health monitoring application and causes the health monitoring application to generate reminders

[0187] This disclosure includes the following examples.

[0188] Example 1 : A method includes extracting, by a computing device, user interaction data from one of: a phone call or a message; determining, by the computing device and using a machine learning model, a particular action to perform based on the user interaction data; dynamically identifying, by the computing device and based on the particular action, a particular application from one or more applications capable of execution by the computing device; dynamically generating, by the computing device, a software component that enables the computing device to cause the particular application to perform the particular action; and causing, by the computing device and using the software component, the particular application to perform the particular action.

[0189] Example 2: The method of example 1, wherein determining an action to perform includes: predicting, by the computing device and using a machine learning model, the particular action based on the user interaction data.

[0190] Example 3: The method of any of examples 1 and 2, further includes identifying, by the computing device, a plurality of potential actions, where each potential action of the plurality of potential actions is performed by a corresponding application of the one or more applications executed by the computing device, and wherein dynamically identifying the particular action is based on identifying the particular action from the plurality of potential actions.

[0191] Example 4: The method of example 3, wherein dynamically identifying the plurality of potential actions includes: responsive to installation of a new application, determining a set of functionality of the new application.

[0192] Example 5: The method of any of examples 1 through 4, wherein the software component is an application programming interface, and wherein dynamically generating the software component includes generating, using a language model, the application programming interface.

[0193] Example 6: The method of any of examples 1 through 5, wherein dynamically generating the software component includes generating, using a language model, the software component configured to use an application programming interface associated with the particular application to cause the particular application to perform the particular action.

[0194] Example 7: The method of any of examples 1 through 6, wherein the user interaction data include audio data of the phone call, and wherein determining the particular action to perform includes providing the audio data to a language model that outputs an indication of the particular action.

[0195] Example 8: The method of any of examples 1 through 7, wherein dynamically identifying the particular application is based on at least one of: an application store description of the particular application, an application programming interface of the particular application, a list of applications from the application store, a list of actions that includes information about a respective set of functionality of each application, or pretraining data of a language model that includes information regarding the particular application.

[0196] Example 9: The method of any of examples 1 through 8, wherein extracting the user interaction data includes analyzing the phone call in response to user interaction approving analysis of the phone call.

[0197] Example 10: The method of any of examples 1-9, wherein causing the application to perform the particular action is in response to user interaction via at least one of a graphical user interface or an audio interface approving performance of the particular action.

[0198] Example 11 : The method of any of examples 1-10, wherein dynamically generating the software component includes dynamically generating the software component using a machine learning model, and wherein the machine learning model is at least one of a language model or large language model

[0199] Example 12: A computing device includes a memory, and one or more programmable processors in communication with the memory, and configured to: extract user interaction data from one of: a phone call or a message; determine, using a machine learning model, a particular action to perform based on the user interaction data; dynamically identify, based on the particular action, a particular application from one or more applications capable of execution by the computing device; dynamically generate a software component that enables the computing device to cause the particular application to perform the particular action; and cause, using the software component, the particular application to perform the action.

[0200] Example 13: The computing device of example 12, wherein to determine an action to perform, the one or more programmable processors are further configured to: predict, using a machine learning model, the particular action based on the user interaction data.

[0201] Example 14: The computing device of any of examples 12 and 13, wherein the one or more programmable processors are further configured to: identify a plurality of potential actions, where each potential action of the plurality of potential actions is performed by a corresponding application of the one or more applications executed by the computing device, and wherein to dynamically identify the particular action, the one or more programmable processors are configured to dynamically identify the particular action based on identifying the particular action from the plurality of potential actions.

[0202] Example 15: The computing device of example 14, wherein to dynamically identify the plurality of potential actions, the one or more programmable processors are further configured to: responsive to installation of a new application, determine a set of functionality of the new application.

[0203] Example 16: The computing device of any of examples 12 through 15, wherein the software component is an application programming interface, and wherein to dynamically generate the software component, the one or more programmable processors are further configured to generate, using a language model, the application programming interface.

[0204] Example 17: A non-transitory computer-readable storage medium, encoded with instructions that, when executed, cause one or more processors of a computing device to perform any of the methods of examples 1-11.

[0205] Example 16: A computer program product comprising instructions that, when executed, cause one or more processors to perform any of the methods of examples 1-11.

[0206] By way of example, and not limitation, such computer-readable storage media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other storage medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer- readable medium. For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. It should be understood, however, that computer-readable storage mediums and media and data storage media do not include connections, carrier waves, signals, or other transient media, but are instead directed to nontransient, tangible storage media. Disk and disc, as used herein, includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of a computer-readable medium.

[0207] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structures or any other structures suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules. Also, the techniques could be fully implemented in one or more circuits or logic elements.

[0208] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, various units may be combined in a hardware unit or provided by a collection of inter-operative hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.

[0209] Various embodiments have been described. These and other embodiments are within the scope of the following claims.

Claims

WHAT IS CLAIMED IS:

1. A method, comprising: extracting, by a computing device, user interaction data from one of: a phone call or a message; determining, by the computing device and using a machine learning model, a particular action to perform based on the user interaction data; dynamically identifying, by the computing device and based on the particular action, a particular application from one or more applications capable of execution by the computing device; dynamically generating, by the computing device, a software component that enables the computing device to cause the particular application to perform the particular action; and causing, by the computing device and using the software component, the particular application to perform the particular action.

2. The method of claim 1, wherein determining an action to perform includes: predicting, by the computing device and using a machine learning model, the particular action based on the user interaction data.

3. The method of any of claims 1-2, further comprising: identifying, by the computing device, a plurality of potential actions, where each potential action of the plurality of potential actions is performed by a corresponding application of the one or more applications executed by the computing device, and wherein dynamically identifying the particular action is based on identifying the particular action from the plurality of potential actions.

4. The method of claim 3, wherein dynamically identifying the plurality of potential actions includes: responsive to installation of a new application, determining a set of functionality of the new application.

5. The method of any of claims 1-4, wherein the software component is an application programming interface, and wherein dynamically generating the software component includes generating, using a language model, the application programming interface.

6. The method of any of claims 1-5, wherein dynamically generating the software component includes generating, using a language model, the software component configured to use an application programming interface associated with the particular application to cause the particular application to perform the particular action.

7. The method of any of claims 1-6, wherein the user interaction data include audio data of the phone call, and wherein determining the particular action to perform includes providing the audio data to a language model that outputs an indication of the particular action.

8. The method of any of claims 1-7, wherein dynamically identifying the particular application is based on at least one of: an application store description of the particular application, an application programming interface of the particular application, a list of applications from the application store, a list of actions that includes information about a respective set of functionality of each application, or pre-training data of a language model that includes information regarding the particular application.

9. The method of any of claims 1-8, wherein extracting the user interaction data includes analyzing the phone call in response to user interaction approving analysis of the phone call.

10. The method of any of claims 1-9, wherein causing the application to perform the particular action is in response to user interaction via at least one of a graphical user interface or an audio interface approving performance of the particular action.

11. The method of any of claims 1-10, wherein dynamically generating the software component includes dynamically generating the software component using a machine learning model, and wherein the machine learning model is at least one of a language model or large language model.

12. A computing device, comprising: a memory, and one or more programmable processors in communication with the memory, and configured to: extract user interaction data from one of: a phone call or a message; determine, using a machine learning model, a particular action to perform based on the user interaction data; dynamically identify, based on the particular action, a particular application from one or more applications capable of execution by the computing device; dynamically generate a software component that enables the computing device to cause the particular application to perform the particular action; and cause, using the software component, the particular application to perform the action.

13. The computing device of claim 12, further comprising means for performing any combination of the methods of claims 2-11.

14. A non-transitory computer-readable storage medium, encoded with instructions that, when executed, cause one or more processors of a computing device to perform any of the methods of claims 1-11.

15. A computer program product comprising instructions that, when executed, cause one or more processors to perform any of the methods of claims 1-11.

Citation Information

Patent Citations

  • Initializing non-assistant background actions, via an automated assistant, while accessing a non-assistant application

    US20220157317A1

  • User-oriented actions based on audio conversation

    US20220293096A1

  • Language Model Prediction of API Call Invocations and Verbal Responses

    US20230197070A1