Abstraction of computer-based interactions for task automation
By simulating computer interactions at different levels of abstraction to generate synthetic sequences and performing federated learning, the efficiency and privacy issues of individuals performing computer tasks are addressed, achieving cross-individual task automation and privacy protection.
Patent Information
- Application Number
- CN202480038985.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-14
- Filing Date
- 2024-04-09
- Publication Date
- 2026-01-20
AI Technical Summary
Individuals are prone to errors and consume resources when repeatedly performing computer tasks, and it is difficult to automate tasks across individuals without disclosing sensitive information.
By using a private machine learning model to simulate computer interactions at different levels of abstraction, synthetic interaction sequences are generated and federated learning is performed in a global machine learning model to automate tasks while protecting individual semantic privacy.
It enables cross-personal task automation without disclosing sensitive information, improving the efficiency and accuracy of task execution and reducing resource consumption.
Smart Images

Figure CN121368765A_ABST
Abstract
Description
BACKGROUND
[0001] Individuals often operate computing devices to perform tasks that can be replicated to varying degrees by others. For example, an individual can use a first computer application to engage in a sequence of actions to perform a given task, such as setting various application preferences, retrieving / viewing particular data made accessible by the first computer application, performing a series of operations within a particular domain (e.g., 3D modeling, graphics editing, word processing), and so on. Different individuals can later engage in semantically similar sequences of actions, such as when using different computer applications, or when using the same type of computer application but with different purposes, to perform semantically similar tasks in different contexts. Repeatedly performing the actions that comprise these tasks can be cumbersome, error-prone, and can unnecessarily consume computing resources and / or the individual’s attention. SUMMARY
[0002] Embodiments are described herein for protecting individual semantic privacy while facilitating automation of tasks across a group of individuals. More specifically, but not exclusively, embodiments are described herein for enabling an individual (often referred to as a “user”) to adjust a level of abstraction associated with captured (e.g., recorded, observed) sequences of computer-based interactions (e.g., user inputs, presented outputs) before those captured sequences are utilized to automate performance of tasks across a group.
[0003] In some embodiments, a method can be implemented using one or more processors and can include sampling a plurality of interactions between a user and a computer application, wherein the interactions are collectively associated with the user performing a high-level task; encoding the plurality of interactions as one or more task embeddings of a first level of abstraction; processing one or more of the task embeddings using a private machine learning model to simulate, via one or more output devices, performance of the high-level task at the first level of abstraction for the user; training the private machine learning model based on user input rejecting the first level of abstraction, wherein the training generates an updated private machine learning model; and providing parameters of the updated private machine learning model for federated learning of a global machine learning model.
[0004] In various embodiments, the method can include, in response to user input rejecting the first level of abstraction, encoding the plurality of interactions as one or more second task embeddings of a second level of abstraction, the second level of abstraction being different from the first level of abstraction; and prior to the training, processing one or more of the second task embeddings using the private machine learning model to simulate, via one or more of the output devices, performance of the high-level task at the second level of abstraction for the user.
[0005] In various implementations, the simulated execution of the high-level task at the second level of abstraction can exclude one or more interactions of the sampled plurality of interactions. In various implementations, the simulated execution of the high-level task at the second level of abstraction can exclude or obfuscate one or more pieces of information input by the user or output to the user during the sampling. In various implementations, a different softmax layer temperature can be used to encode the first task embedding compared to encoding the second task embedding.
[0006] In various implementations, providing can include providing, to a remote computing system that maintains a global machine learning model, data indicative of the local gradient. In various implementations, the private machine learning model can be a transformer. In various implementations, the private machine learning model can be a large language model (LLM). In various implementations, the token predicted based on the LLM can correspond to the first plurality of interactions.
[0007] In another related aspect, a method can be implemented using one or more processors and includes recording data indicative of an observed set of interactions between a user and a computing device; based on the recorded data, simulating a plurality of different synthetic sets of interactions between the user and the computing device, wherein each synthetic set includes a variation of the observed set of interactions at a different level of abstraction; obtaining user feedback regarding each of the plurality of different sets; based on the user feedback, selecting one of the plurality of different synthetic sets of interactions; and causing a machine learning model to be trained to generate an output indicative of the selected synthetic set of interactions.
[0008] In various implementations, the simulating is performed based on a machine learning model. In various implementations, the machine learning model can be trained to facilitate intelligent process automation.
[0009] In various implementations, the machine learning model can be trained to generate a probability distribution over an action space. In various implementations, the machine learning model can include a private machine learning model, and the method can further include providing parameters of the trained private machine learning model for federated learning of a global machine learning model.
[0010] Additionally, some implementations include one or more processors of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in associated memory, and wherein the instructions are configured to cause performance of any of the aforementioned methods. Some implementations include at least one non-transitory computer- readable storage medium storing computer instructions executable by one or more processors to perform any of the aforementioned methods. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1is a diagram of an example environment in which implementations disclosed herein can be implemented.
[0012] Figure 2 is a diagram illustrating how components of Figure 1 may interact to automate performance of a user's task.
[0013] Figure 3 is a flow diagram of an example method showing selected aspects of practicing the present disclosure in accordance with implementations disclosed herein.
[0014] Figure 4 is a flow diagram of another example method showing selected aspects of practicing the present disclosure in accordance with implementations disclosed herein.
[0015] Figure 5 shows an example architecture of a computing device. DETAILED DESCRIPTION
[0016] Implementations are described herein for protecting individual semantic privacy while facilitating task automation across a group of individuals. More specifically, but not exclusively, implementations are described herein for enabling an individual (often referred to as a "user") to adjust a level of abstraction associated with captured (e.g., recorded, observed) sequences of computer-based interactions (e.g., user inputs, presented outputs) before those captured sequences of computer-based interactions are utilized to automate performance of a task across the group.
[0017] In various implementations, sequences of computer-based interactions between a user and a computing device can be recorded when the user operates the computing device to perform a high-level task (such as writing a letter, reading a scientific article, editing a digital image, etc.). These computer-based interactions can include user inputs (such as keystrokes, pointer device activity, interactions with graphical user interface (GUI) elements, speech inputs, etc.), as well as outputs generated by the computing device (e.g., auditory, visual, haptic, etc.).
[0018] To enable the user to observe and / or control how sensitive and / or private information will be shared to automate performance of the high-level task for others, variations of these computer-based interactions at different levels of abstraction can be simulated for the user. Each simulation can include a different composite set of computer-based interactions involved in the higher-level task previously performed by the user. In other words, each composite set can include a variation of the observed set of interactions at a different level of abstraction. The user can provide feedback (e.g., accept, reject, modify) in response to each simulation, particularly with respect to the respective level of abstraction (e.g., level of detail of factual information) reflected in each simulation.
[0019] As an example, suppose that a sequence of computer-based interactions of a user operating a word processing application to write a letter is recorded. An initial letter composition simulation presented to the user can automatically populate a physical address of a recipient. Suppose that the user does not want to share the recipient's address (e.g., because it is private or otherwise inapplicable), the user can provide feedback (e.g., "No, do not include the recipient's address") that rejects the current / applicable level of abstraction that resulted in the automatic population of the recipient's address. A subsequent letter composition simulation can implement a different level of abstraction by omitting the recipient's address (e.g., instead by including a generic address placeholder). If the user approves the subsequent letter composition simulation, data of the computer-based interactions that indicate the subsequent letter composition simulation can be utilized to automate letter composition tasks generally (e.g., across a population of users).
[0020] High-level tasks can be automated in various ways. In some implementations, a "private" (e.g., local) embedding machine learning model (also referred to herein as an "embedding model") can be trained to generate embeddings at various levels of abstraction. Additionally or alternatively, a "private" (e.g., local) action machine learning model (also referred to herein as an "action model") can be trained to generate outputs that indicate synthetic sequences of computer-based interactions related to performing high-level tasks. These private embedding and / or action models can be stored locally at a client device operated by a user, or can be stored in a "private cloud" that is controlled for access by the user. In either case, the learning parameters of the private embedding and / or action models can be combined, e.g., with learning parameters of other private embedding and / or action models of other users, to train "global" or "public" embedding and / or action models as part of a federated learning framework.
[0021] Sets of computer-based interactions generated based on (private or global) action models as described herein are referred to as "synthetic" because their constituent computer-based interactions are predicted rather than observed. Thus, a synthetic set of computer-based interactions will include at least some predicted computer-based interactions that are not identical to recorded computer-based interactions in real life. More generally, a synthetic set of computer-based interactions can exhibit a different level of abstraction than a recorded set of computer-based interactions. A recorded computer-based interaction that potentially exposes sensitive information (e.g., a social security number) to an untrusted party can be "abstracted" such that the sensitive information is excluded or obfuscated. Similarly, a recorded computer-based interaction that can not necessarily expose sensitive information but is not broadly applicable outside of a narrow context can also be abstracted to be more broadly applicable.
[0022] Private or global action models can take various forms. In some implementations, an action model configured with selected aspects of the present disclosure can take the form of a neural network trained to generate probability distributions over an action space. Based on these probability distributions, a synthetic set of computer-based interactions can be generated. The action space can be populated with, for example, computer-based interactions (particularly user inputs) that can be performed using a computing device. To reduce and / or manage the size of the search space, in some implementations, domain-specific action models can be trained to generate probability distributions over an action space of a particular domain, such as for a particular computer application, a particular context (e.g., various scientific disciplines, various positions in an organization), and / or the like. As used herein, a “domain” can refer to a target area of subject matter in which a computing component intends to operate, e.g., a range of knowledge, influence, and / or activity that the logic of the computing component centers on.
[0023] In other implementations, an action model can take the form of a sequence-to-sequence model that generates a sequence of output tokens corresponding to and / or representing a synthetic set of computer-based interactions. A sequence-to-sequence action model can include, for example, various types of recurrent neural networks (RNNs) such as long short-term memory (LSTM) networks or gated recurrent unit (GRU) networks.
[0024] A sequence-to-sequence action model can alternatively include various types of transformer networks, such as a bidirectional encoder representations from transformers (BERT) transformer or a generative pre-trained transformer (GPT). In some implementations, a large language model (LLM) can be used, which can also take the form of a transformer but with a large number of parameters. In some such implementations, a beam search can be performed at various beam widths (which can be selected by a user, e.g., as part of controlling the level of abstraction) in order to generate synthetic sets of computer-based interactions at different levels of abstraction. In some implementations, a sequence of input tokens can represent observed computer-based interactions that actually occurred between a user and a computing device. In various implementations, output tokens can correspond to and / or represent a synthetic set of computer-based interactions between a user and a computing device.
[0025] The level of abstraction associated with a set of computer-based interactions can be changed in other ways. In some implementations, individual actions can be altered, e.g., based on commands received from a user, to change or exclude data. For example, in the letter writing example described earlier, the user provided feedback— "Don't include the recipient's address." Accordingly, natural language processing and / or pattern recognition can be performed on the letter being composed to identify the recipient's address, which can then be excluded, obfuscated, replaced with a generic placeholder, etc., before being tokenized (e.g., converted to a domain-specific language (DSL) and / or embedded) and applied as input to a cross-action model.
[0026] As another example, individual computer-based interactions can be altered, e.g., symbolically and / or using various meta-heuristics (e.g., simulated annealing, genetic algorithms), to be more generally applicable. Suppose a physics expert is reading a digital paper on biology, and the physics expert receives clarifications (e.g., definitions, acronyms) of various biology terms contained in the digital paper, e.g., automatically from a virtual assistant or by request. These clarifications can be presented to the physics expert in various ways, such as computer-generated speech and / or text of the virtual assistant, annotations in the margin, pop-up windows highlighting the terms being clarified, etc. The interactions between the overall physics expert, the application presenting the digital paper, the biology terms being clarified, the virtual assistant, and / or the entire computing device can be recorded. In some implementations, situational data can also be recorded, such as the fact that an expert in one field (e.g., physics) is consuming media (a digital paper) from another field (e.g., biology) that they are not an expert in.
[0027] Some of these recorded computer-based interactions can be abstracted so that they are applicable outside of the specific context in which they were recorded. For example, clarifying a particular biology term to a physics expert can not be applicable to another situation, e.g., the biology expert is reading a digital paper on computer science. Accordingly, one or more computer-based interactions that collectively resulted in clarifying a biology term to a physics expert can be abstracted to, e.g., a symbolic template that more broadly causes a term or phrase in any field to be clarified to an individual that is not an expert in that field.
[0028] In some implementations, words or phrases can be identified as suitable for clarification based on metrics such as word length, term frequency (e.g., using term frequency-inverse document frequency or “TF-IDF” calculations), and the like. Suppose a physics expert seeks clarification of a biology term that has a TF-IDF score that falls above a particular threshold, is within a range, and the like. The resulting symbolic template can be created to clarify terms or phrases in any domain that have similar TF-IDF scores. In some implementations, the symbolic template can then be tokenized (e.g., converted to a DSL and / or embedded) and applied as input to a cross-action model.
[0029] Another way to change the level of abstraction of the (observed or synthesized) set of computer-based interactions is to reduce the dimensionality of the tokens (e.g., vector / feature embeddings, DSLs, and the like) that encode the set of computer-based interactions. In some implementations, a sequence of records of computer-based interactions can be tokenized / encoded into an x (positive integer)-dimensional embedding, e.g., using one or more of the private embedding models described above (e.g., obtained from an encoder-decoder machine learning model). If a user requests a higher level of abstraction, then dimensionality reduction can be performed to encode the x-dimensional embedding into a y-dimensional new embedding, where y is a positive integer that is less than x. In some such implementations, the dimensionality reduction can be lossy to prevent reconstruction of information that the user’s intent protects. Processing the y-dimensional embedding using an action model configured with selected aspects of the present disclosure can yield a synthesized set of computer-based interactions that is more abstract than the x-dimensional embedding represents. In some implementations, fuzzy semantic privacy can be employed to make it at least probabilistically difficult or impossible to recreate the user’s original scenario. For example, various types of transformers can be applied to data indicative of the user’s recorded actions to create a new sequence of recorded actions. The new synthesized sequence of actions can resemble the user’s original actions but differ in details (e.g., values).
[0030] Another way to change the level of abstraction of the set of computer-based interactions is to adjust one or more hyperparameters of the action model itself. As one example, the temperature of a softmax layer of the action model can be adjusted to alter the probability distribution generated by the action model. A “higher” temperature can yield a “softer,” less confident, and / or more uniform probability distribution, while a “lower” temperature can yield a “harder,” more confident, and / or less uniform probability distribution. Random sampling from the former can yield more variation (and thus a different level of abstraction) than the latter.
[0031] As another example of work, suppose that a user is using an image editing application to edit a digital image. Further suppose that the user performs a particular action on the image, such as removing noise, cropping the image, converting the image to a particular resolution (e.g., a reduced resolution for viewing on a display as opposed to printing), etc. In various implementations, the user can be shown one or more simulations that demonstrate variations in different levels of abstraction of these actions, e.g., performed on the same image, on different images, or in the absence of an image at all. In the last case, the user can be shown the actions as conceptual objects, with adjustable attributes corresponding to parameters of the actions that the user implemented on a real image. The user can specify values for these parameters in order to create a suitable abstraction for generalization. Additionally or alternatively, in some implementations, the user can specify, e.g., using natural language, a quality of the image that the parameters should govern. For example, suppose that the user manually crops an image to focus on a person depicted in the image, while excluding unwanted background. The user can provide a natural language annotation, e.g., as a response to feedback on a subsequent simulation, such as "crop to within 1 cm of the subject's face in all directions."
[0032] Figure 1 An example environment in which selected aspects of the present disclosure can be implemented in accordance with various implementations is schematically depicted. Figure 1 Any of the computing devices depicted in the figures or elsewhere herein can include logic such as one or more microprocessors (e.g., central processing units or "CPUs," graphics processing units or "GPUs," tensor processing units or "TPUs") that run computer-readable instructions stored in memory, or other types of logic such as application specific integrated circuits ("ASICs"), field programmable gate arrays ("FPGAs"), etc. Figure 1 Some of the systems depicted in the figures, such as task automation system 120, can be implemented in whole or in part using one or more server computing devices that form what is sometimes referred to as a "cloud infrastructure."
[0033] Task automation system 120 can be operatively coupled with one or more client computing devices (also referred to herein as "clients"), such as client computing device 110, via one or more computer networks 114. Task automation system 120 can automatically determine a set of computer-based interactions for attempting automation of a higher-level task performed by a user of a client device (e.g., 110).
[0034] A person (which can also be referred to as a "user" in the present context) can operate client device 110 to interact with task automation system 120 to cause task automation system 120 to perform one or more of the following: Figure 1The other components depicted interact. Each client device 110 can be, for example, a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device of a participant's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker (with or without a display), or a wearable device that includes a computing device (such as a head-mounted display ("HMD") that provides an AR or VR immersive computing experience, a "smart" watch), etc. Additional and / or alternative client devices can be provided.
[0035] Examples described herein generally involve a user operating a computing device such as a client device 110 to record a sequence of computer-based interactions that are automated. However, this is not meant to be limiting. The sequence of computer-based interactions can also include interactions with other types of computing devices. For example, when a driver operates a vehicle, the interactions between the driver and the vehicle, which is equipped with various in-vehicle circuitry and / or logic configured with selected aspects of the present disclosure, can be recorded. Later, those interactions can be simulated, e.g., for the driver or for another person (e.g., someone whose task is to train one or more machine learning models), at different levels of abstraction. Based on those simulations, feedback can be provided that allows at least some of the driving interactions to be applied outside the particular context in which the driver initially operated the vehicle that is to be automated. For example, the interactions that the driver makes when parallel parking between two vehicles can be recorded and abstracted to allow parallel parking between any two structures.
[0036] A client device 110 can include one or more applications that interact with the task automation system 120, such as an application 112. For example, the application 112 can be an application via which input of a user can be provided and via which output generated by the task automation system 120 can be presented to the user, such as output that requests user feedback (e.g., output that reflects a final state of a simulation of a candidate set of interactions) and / or output that reflects a set of computer-based interactions determined by the task automation system 120 (e.g., for confirmation of the set of automation). In some implementations, where the task includes control of a computer application running on the client 110, the application 112 (or another application) can be an application that can be controlled using a synthesized set of computer-based interactions determined by the task automation system 120.
[0037] In various implementations, the task automation system 120 includes a global embedding engine 122, a global action engine 124, a global selection engine 126, a global simulation (SIM) engine 128, and / or a global evaluation engine 130. Although the task automation system 120 is depicted as a single system, it can be distributed across multiple computing devices, such as multiple servers in a data center, multiple servers in a cloud computing environment, and / or multiple servers in a distributed computing environment. Figure 1The task automation system 120 is depicted as connected to the client 110 via the network 114, but one or more aspects of the task automation system 120 can be combined and / or can be implemented locally at the client 110 in various implementations. For example, one or more of the engines of the task automation system 120 can additionally or alternatively be implemented at the client 110. For example, the client 110 includes a private embedding engine 122A, a private action engine 124A, a private selection engine 126A, a private SIM engine 128A, and / or a private evaluation engine 130A. Additionally or alternatively, in some implementations, all or part of the task automation system 120 can be implemented in a private cloud (not shown) controlled by the user, such that the virtually limitless resources of the cloud can be leveraged without compromising the user’s personal privacy or security.
[0038] The public embedding engine 122 and / or the private embedding engine 122A can interface with one or more public and / or private embedding ML models 152, 152A in generating the embeddings described herein. Which embedding ML model 152 / 152A the embedding engine 122 / 122A interfaces with, and / or the data it processes in interfacing with one or more of the embedding ML models 152 / 152A can depend on the embedding technique the embedding engine 122 / 122A is utilizing.
[0039] For example, for a given NL input, the embedding engine 122 / 122A can generate an embedding based on a first embedding technique that processes NL input data reflecting the NL input using a domain-specific LLM model of the embedding model 152 / 152A. For various types of computer-based interactions, the embedding engine 122 / 122A can generate an embedding based on a second embedding technique that includes processing the computer-based interaction using some other domain-specific model.
[0040] The embedding technique utilized by the embedding engine 122 / 122A at a given instance can depend on various factors, and in some implementations can be dictated by the selection engine 126 / 126A. For example, the embedding technique utilized can depend on the domain of the task, the computer-based interaction for which the embedding is being generated, and / or the action models 154 / 154A utilized by the action engine 124 / 124A in generating a candidate synthetic set of computer-based interactions. Various embedding ML models 152 / 152A can be provided. For example, the embedding ML models 152 / 152A can include those specific to particular domains, those specific to particular sets of domains, and / or those that are domain-agnostic. As another example, the embedding ML models 152 / 152A can additionally or alternatively include those specific to a first type of data (e.g., natural language data), those specific to a second type of data (e.g., computer-based interactions), etc.
[0041] The global action engine 124 can interface with one or more global (or “public”) action models 154 in facilitating automation of a collection of computer-based interactions. In some implementations, the one or more global action models 154 can be trained to facilitate robotic process automation and / or intelligent process automation. Similarly, the private action engine 124A can interface with one or more private action models 154A in facilitating automation of a collection of computer-based interactions while maintaining user privacy. In some implementations, the action engine 124 / 124A also interfaces with one or more action rules 164 / 164A, which can be used to eliminate some generated candidate synthetic collections of computer-based interactions according to further considerations (e.g., according to further considerations of the evaluation engine 130 / 130A). The action rules 164 / 164A can be specific to a domain and / or specific to a corresponding requesting entity, such as a user or an organization associated with the user. For example, for a particular domain and a particular organization, a given action rule can define that a given action is not allowed at all, or is not allowed if it occurs before or after certain other actions.
[0042] Which action model(s) 154 / 154A the action engine 124 / 124A interfaces with at a given instance can depend on various factors, and in some implementations can be indicated by the selection engine 126 / 126A. For example, the action model 154 / 154A utilized can depend on the domain of the task, the input being generated an embedding for, and / or the embedding generated by the embedding engine 122 / 122A. Further, for example, the action model 154 / 154A utilized for an input in a given instance can additionally or alternatively be based on the action models utilized in prior instances in generating candidate synthetic collections of computer-based interactions for that input and / or in evaluating those candidate synthetic collections.
[0043] Various action models 154 / 154A can be provided. For example, the action models 154 / 154A can include machine learning models and / or heuristic models. As another example, the action models 154 / 154A can include action models specific to a particular domain, action models specific to a particular set of domains, and / or domain-agnostic action models. As another example, the action models 154 / 154A can include those representing reinforcement learning (RL) policies and used to generate candidate synthetic sets of computer-based interactions by iteratively generating corresponding next actions for the candidate synthetic sets based on state data to which updates are applied at each iteration, those used to generate one or more candidate synthetic sets of computer-based interactions in a single iteration, those representing value functions and used to generate a measure reflecting a value of a set of computer-based interactions, a current state pair, and / or other action models. For example, the action models 154 / 154A can include one or more of RL policy machine learning (ML) models, action sequence ML models, constraint satisfaction models, SAT solvers, and / or other models.
[0044] In some implementations, the selection engine 126 / 126A can interact with the embedding engine 122 / 122A in indicating which embedding technique the embedding engine 122 / 122A is to utilize at a given instance, and / or can interact with the action engine 124 / 124A in indicating which action model(s) the global action engine 124 / 124A is to utilize at a given instance. For example, the selection engine 126 / 126A can indicate that the embedding engine 122 / 122A is to initially utilize a first embedding technique. Then, if the evaluation engine 130 / 130A indicates that the corresponding set of computer-based interactions generated based on the first embedding technique is not suitably abstract, e.g., based on feedback from a user, the selection engine 126 / 126A can indicate that the embedding engine 122 / 122A is to utilize a second embedding technique in generating additional embeddings. As another example, the selection engine 126 / 126A can indicate that the embedding engine 122 / 122A is to initially utilize a first embedding technique and a second embedding technique. Then, only if the evaluation engine 130 / 130A indicates that the corresponding candidate synthetic sets of computer-based interactions generated based on the first embedding technique and the second embedding technique are not suitable, the selection engine 126 / 126A can indicate that the embedding engine 122 / 122A is to utilize a third embedding technique and / or a fourth embedding technique in generating additional embeddings.
[0045] In some implementations, the selection engine 126 / 126A can optionally utilize one or more selection models 156 / 156A in determining which embedding technique(s) and / or actions to utilize at a given instance. For example, the selection model 156 / 156A can include a selection ML model that can be used to process the domain of the task and / or the NL input data of the request task (e.g., an embedding of the NL input data) and generate an output indicating a corresponding probability for each of a plurality of embedding techniques and / or action models. The selection engine 126 / 126A can utilize the generated output in selecting which embedding technique(s) and / or action model(s) to utilize. For example, the selection engine 126 / 126A can use the output to select the highest probability embedding technique and / or the highest probability action model for initial utilization. Such a selection ML model can be trained based on supervised training examples that are based on past sets of computer-based interactions that were determined to be appropriate abstractions (and optionally confirmed as suitable after their real-world implementations).
[0046] For each of the set of computer-based interactions generated by the action engine 124, the SIM engine 128 / 128A can be used to simulate implementation of the set of computer-based interactions in a simulated environment, such as a simulated environment that reflects a current state of the domain. Further, the SIM engine 128 / 128A generates simulation data for each of the simulation. In some cases, the set of actions can be generated via the SIM engine 128 during the simulation. For example, some RL policy models can be utilized in the simulation to generate the set of computer-based interactions that will be implemented during the simulation, and their generation will depend on the simulated states encountered during the simulation.
[0047] The evaluation engine 130 / 130A can determine whether a candidate synthesis set of computer-based interactions is an appropriate abstraction and / or determine a most appropriate abstraction of a candidate synthesis set of computer-based interactions from among multiple candidate synthesis sets of computer-based interactions. In various implementations that utilize simulation data in evaluating candidate synthesis sets of computer-based interactions, the evaluation engine 130 / 130A can solicit and / or utilize user feedback based on the simulation data. For example, the evaluation engine 130 / 130A can cause simulation data from a simulation to be presented to a user from which a set of observed computer-based interactions was recorded, and determine suitability based on feedback from the user in response to the presentation. For example, the evaluation engine 130 / 130A can cause a screenshot, final state, video, etc. of a simulated environment from a simulation in its final state to be presented at the client 110. In response, the user can provide user interface input that reflects whether the output is abstracted enough to protect the user’s privacy, e.g., while continuing to be suitable overall for automating higher-level tasks. The evaluation engine 130 / 130A can use instances of negative feedback to eliminate a corresponding candidate synthesis set of computer-based interactions, or negatively impact a suitability metric for the candidate synthesis set. By contrast, the evaluation engine 130 / 130A can use instances of positive feedback to select a candidate synthesis set of computer-based interactions as the most appropriate, or positively impact a suitability metric for the corresponding candidate synthesis set. In various implementations, in evaluating synthesis sets of computer-based interactions, the evaluation engine 130 / 130A utilizes simulation data from simulations of the set of actions by the SIM engine 128.
[0048] In some of those implementations, the evaluation engine 130 / 130A can compare simulation data to one or more state rules 160 / 160A, which can be used to determine that a candidate set of actions is not suitable and / or negatively impact a suitability score for the candidate set of actions that is used to determine suitability of the candidate set of actions. The state rules 160 / 160A can be specific to a domain and / or specific to a corresponding requesting entity, such as a providing user or an organization associated with the user. For example, for a particular domain and a particular organization, a given state rule (e.g., predefined or provided as feedback by a user) can define that a given state or a particular sequence of states is never encountered. If simulation data from a simulation of a candidate set of actions shows that the given state and / or the particular sequence of states is encountered, the evaluation engine 130 / 130A can determine that the candidate set of actions is not suitable. As another example, for a particular domain and a particular organization, a given state rule can define that a given state or a particular sequence of states is undesirable, but not prohibited. If simulation data from a simulation of a candidate set of actions shows that the given state and / or the particular sequence of states is encountered, the evaluation engine 130 / 130A can negatively impact a suitability metric for the candidate set of actions.
[0049] The machine learning models described herein can have various architectures and be trained in various ways. For example, one or more of the models can be a graph-based neural network (e.g., as a graph neural network (GNN), a graph attention neural network (GANN), or a graph convolutional neural network (GCN)), a sequence-to-sequence neural network (such as a transformer, an encoder-decoder, or a recurrent neural network (“RNN,” e.g., long short-term memory or “LSTM,” gated recurrent unit or “GRU,” etc.), BERT (Bidirectional Encoder Representations from Transformers), etc. Moreover, for example, one or more of the machine learning models can be trained with reinforcement learning, supervised learning, and / or imitation learning. Additional descriptions of some implementations of various machine learning models are provided herein.
[0050] SUMMARY Figure 1 The global (or “public”) engines 122, 124, 126, 128, and / or 130 can be used, for example, to select from among a plurality of different candidate synthetic sets of computer-based interactions for real-world implementation in response to user requests to perform various high-level tasks. Additionally, one or more of the global engines 122, 124, 126, 128, and / or 130 can facilitate federated learning of one or more models 152, 154, 156 based on local model parameters / local gradients received from the private (or “local”) engines 122A, 124A, 126A, 128A, and / or 130A. In this regard, the private (or “local”) engines 122A, 124A, 126A, 128A, and / or 130A can facilitate generation and distribution of these local model parameters / local gradients based on provided user feedback regarding simulation of computer-based interactions with syntheses at various levels of abstraction.
[0051] Turning to Figure 2 A description of the following example is provided: engines 122A, 124A, 126A, 128A, and 130A of the task automation system 120 implemented at the client device 110 for the purpose of federated learning. Figure 2 Interactions that can occur between those engines, models 152A, 154A, and 156A, and rules 160A and 164A (which in some cases can be the same as 160 and 164), are also depicted, which the client 110 can utilize to facilitate federated learning.
[0052] In Figure 2In particular embodiments, the private embedding engine 122A processes various data, such as data indicative of observed computer-based interactions (“interactions”) 104, and in some instances, NL input 101. Based on this data, the private embedding engine 122A generates task embeddings 123. Domain-specific knowledge (DSK) 102 can include, for example, various reference documents or other supplemental data, which can be provided as additional input to effectively “condition” or “prime” the model (e.g., 152A, 154A), for example, similar to few-shot learning (except that the model weights can or can not be tuned). Context 103 can include a wide variety of data points, such as signals generated by the client 110 (e.g., location coordinates, time of day, foreground and / or background applications, other sensor signals, etc.), user preferences, user expertise (e.g., physics expert, biology expert, etc.), and the like.
[0053] Data 104 can include any data that is recorded to memorialize a user’s computational interactions with one or more computer applications operating on the client device 110. In some implementations, data 104 can include hardware inputs, such as keystrokes, pointer device movements and / or actions (e.g., click, right-click, scroll, etc.), and the like. In some implementations, data 104 can include application-specific interactions, such as interactions with graphical elements and / or menu items, input command sequences (including NL input 101 provided to interact with one or more computer applications), presented audible and / or visual output, and the like.
[0054] In some implementations, data 104 can include application-specific computer code, such as code that can be generated in some applications (e.g., word processing, spreadsheets, etc.) when a user selects to record a “macro.” In some implementations, data 104 can include information entered by a user, such as information for composing a document (e.g., a letter, an email, a report), information for populating a spreadsheet, information for populating a form (e.g., a portion of a web page or an application), information for creating a graphic design or drawing, information exchanged with a virtual assistant, and the like.
[0055] In either case, the private embedding engine 122A can process data 104 (and other data 101-103, as applicable) to generate one or more task embeddings 123. In some implementations, task embeddings 123 can include separate tokens / embeddings that encode each interaction (e.g., input from a user, output from a computer application). In some implementations, multiple interactions can be combined (e.g., aggregated, concatenated) into a semantically rich task embedding 123 that represents multiple interactions.
[0056] NL input 101 (which is optional and / or can be recorded as part of data 104 or separate from data 104) can be transmitted by the user via a client device (e.g., Figure 1 The task embedding 123 is provided through interaction with a user interface input device of client 110. For example, a user could provide NL input 101, such as, “I am now drafting a letter. Please record my actions so we can automate the process.” In some instances, NL input 101 could be verbal input detected by the user via the microphone of client 110. The proprietary embedding engine 122A can process recognized text generated based on verbal input (e.g., using Automatic Speech Recognition (ASR)) when generating the task embedding 123. As another example, NL input 101 could be typed input provided via a virtual or hardware keyboard of client 110, and the typed text could be processed by the proprietary embedding engine 122A when generating the task embedding 123. For example, an NL ML model of ML model 152A could be used to process the recognized or typed text to generate the NL embedding. The NL ML model could be, for example, an LLM. The task embedding 123 could be an NL embedding, or it could be a function of an NL embedding and other NL embeddings.
[0057] A private action engine 124A uses one or more action ML models to process task embeddings 123 to generate one or more candidate composition sets 125 of computer-based interactions. For example, in a given instance, the private action engine 124A may use a first RL policy model 154B, a second RL policy model 154C, a constraint satisfaction model 154D, an action sequence model 154N, or other models of action model 154 (e.g., by...). Figure 2 The vertical ellipsis in the text indicates one of the other models (the others) to handle task embedding 123. In some of these implementations, the selection engine 126 may indicate which of the action models 154 is used by the private action engine 124A at a given instance. In some implementations, the action sequence model 154N is a sequence-to-sequence model, such as a transformer model (e.g., BERT, GPT, etc.), which can be applied, for example, to iteratively or one-time predict a sequence of labeled sequences representing computer-based interactions.
[0058] For each of the candidate composite sets of computer-based interactions 125, the private SIM engine 128A can be used to simulate implementation of the set of actions in a simulated environment, such as a simulated environment that reflects the current state of the domain. Further, the private SIM engine 128A generates simulation (SIM) data 127 for each of the simulations. In some cases, the candidate composite sets of computer-based interactions for the set 125 can be generated independently of their simulations, and the private SIM engine 128A can be utilized to simulate the set of actions after the candidate composite sets of computer-based interactions are generated. In some other cases, the candidate composite sets of computer-based interactions can be generated by the private action engine 124A during simulation via the private SIM engine 128A. This is reflected by the dashed double-headed arrow line between the private action engine 124A and the private SIM engine 128A.
[0059] The SIM data 127 can be presented (e.g., shown) to the user, for example, by the private SIM engine 128A or the private evaluation engine 130A, such that the user can provide feedback 129. For example, once the SIM data 127 is shown to the user, for example, visually (e.g., as an animation, a snapshot of the current or final state of the domain, a final document, etc.) and / or audibly, the user can be prompted to provide feedback 129 that accepts, rejects, and / or modifies the level of abstraction of one or more of the candidate composite sets of computer-based interactions 125. This can allow the user to control how much personal and / or sensitive information is provided to untrusted and / or public entities.
[0060] The private evaluation engine 130A can determine, based on the user feedback 129, whether the corresponding candidate set of actions of the candidate composite sets of computer-based interactions 125 is appropriately abstracted to protect the user’s privacy, and / or determine the most appropriate abstraction of the candidate composite sets of computer-based interactions from among a plurality of candidate composite sets of computer-based interactions. In some of these implementations, the private evaluation engine 130A can compare the simulation data to one or more state rules 160A, which can be used to determine that the candidate composite sets of computer-based interactions are inappropriately abstracted (e.g., too specific, contain personal data, not widely applicable) and / or negatively impact the suitability score utilized in determining the suitability of the candidate composite sets of computer-based interactions.
[0061] If the private evaluation engine 130A determines, e.g., based on user feedback 129, that none of the candidate synthetic sets of computer-based interactions 125 are suitable, it can output an unsuitable indication 131 to the private selection engine 126A. In response, the private selection engine 126A can alter the embedding technique utilized by the private embedding engine 122A and / or the private action model 154A being utilized by the private action engine 124A. Then, a further candidate synthetic set of computer-based interactions 125 can be generated based on a different task embedding 123 (e.g., generated using an alternative embedding technique) and / or based on a different action model of the private action model 154A.
[0062] For example, the private selection engine 126A can alter the embedding technique being utilized (e.g., by reducing the dimensionality of the embedding), but not the action model 154 being utilized. In response, the private embedding engine 122A can generate a different task embedding 123 using a different altered embedding technique, and the private action engine 124A will utilize the same action model as before to process the different task embedding 123. This can result in a different candidate synthetic set of computer-based interactions 125 being generated due to the different task embedding 123. The private SIM engine 128A can simulate the different action set 125, and the private evaluation engine 130A utilizes the resulting SIM data 127 to allow the user to provide new feedback 129 regarding the different candidate synthetic set of computer-based interactions 125. Such multiple iterations can occur until, e.g., the private evaluation engine 130A and / or the user determine that the candidate synthetic set of computer-based interactions being evaluated is suitable for dissemination to global and / or public entities.
[0063] The private evaluation engine 130A (and / or the private selection engine 126A) can use other techniques to control the level of abstraction used for automating tasks. In some implementations, based on user feedback 129, the private evaluation engine 130A can use various rules and / or heuristics to alter particular pieces of information logged as part of the interaction data 104. One example described previously involved a user indicating (as part of feedback 129) that a particular recipient’s address should not be used when attempting to automate a letter writing task. As another example, during automation of filling out a web form to make a purchase, a user can reject simulation of a candidate synthetic set of computer-based interactions (125) that presents a form field filled out with a particular credit card number. This can result in the user’s credit card number being excluded, obfuscated, replaced with a generic placeholder, etc. prior to being generated by the private embedding engine 122A, e.g.
[0064] In some implementations, the private evaluation engine 130A, the private selection engine 126A, and / or the private action engine 124A can select different private action models 154A and / or alter one or more parameters of a particular action model 154A to control / alter the level of abstraction. As one example, the temperature of a softmax layer of an action model can be adjusted to alter the probability distribution generated by the action model. A“higher” temperature can yield a“softer,” less confident, and / or more uniform probability distribution, while a“lower” temperature can yield a“harder,” more confident, and / or less uniform probability distribution. Random sampling from the former can yield more variation (and thus different levels of abstraction) than the latter.
[0065] Referring back to Figure 2 If the private evaluation engine 130A determines that one of the candidate synthetic sets of computer-based interactions 125 is suitable in a given iteration (e.g., the user feedback 129 indicates that the interaction 125 is a suitable level of abstraction), the private evaluation engine 130A can generate a suitable indication 132. In various implementations, based on the unsuitable indication 131 and / or the suitable indication 132, a local gradient 133 can be computed, e.g., using techniques such as backpropagation, gradient descent, cross-entropy, etc. The local gradient 133 can then be used to update the private embedding ML model 152A and / or the private action model 154A, as indicated by the arrows. The local gradient 133 and / or the updated parameters of the private embedding ML model 152A and / or the private action model 154A can then be provided to the global embedding engine 122 and / or the global action engine 124. The global embedding engine 122 can use the local gradients 133 received from multiple different clients 110 to update the global embedding model 152 as part of a federated learning framework. Similarly, the global action engine 124 can use the local gradients 133 received from multiple different clients 110 to update the global action model 154 as part of a federated learning framework.
[0066] Figure 3 FIG. 3 is a flow diagram illustrating an example method 300 for practicing selected aspects of the disclosure in accordance with implementations disclosed herein. The operations of method 300 are described with reference to a system that performs the operations (e.g., the task automation system 120). The system can include various components of various computing systems, such as one or more components of the task automation system 120. Additionally, while operations of method 300 are shown in a particular order, this is not meant to be limiting. One or more operations can be reordered, omitted, or added.
[0067] At block 302, the system, e.g., through the application 112, operating system, and / or private embedded engine 122A, can record data (e.g., 104) indicative of an observed set of interactions between a user and a computing device. In some implementations, this recording can be triggered by a command from the user, which the user can issue using various types of input (e.g., keyboard input, pointer device input, voice input, etc.). In other implementations, the recording can be triggered automatically, e.g., in response to detecting that the user repeatedly performs a number of actions. For example, if the user repeatedly performs multiple sequences of largely similar actions, an agent (e.g., a virtual assistant or application assistant) configured with selected aspects of the present disclosure can issue a prompt, e.g., “I see that you are again [insert name of task]. Would you like me to generate an automated routine for you to perform these steps in the future?”
[0068] Based on the recorded data, at block 304, the system, e.g., through the private SIM engine 128A, can simulate a plurality of different synthetic sets of interactions between the user and the computing device. Each synthetic set of computer-based interactions can be a variation of the observed set of interactions at a different level of abstraction. In some implementations, the simulation of block 304 can be based on a machine learning model. The machine learning model can be a private action model 154A trained to generate a probability distribution over an action space of a domain (e.g., word processing, graphic design, web browsing, spreadsheet operations, etc.) in which the user is operating.
[0069] For example, the RL policy 154B and / or 154C can generate, at each iteration, a probability distribution over an action space of a domain based on a current state of the domain. Based on the probability distribution, one or more next actions can be selected and simulated, and the process can be repeated. Similarly, the action sequence model 154N can be a sequence-to-sequence model that, e.g., generates a sequence of tokens at a time, each representing one or more actions in the action space of the domain. In some such implementations, each token can include a probability distribution over multiple different actions. In other such implementations, the entire sequence can be assembled based on the probability distributions and then simulated.
[0070] Suppose the repetitive task is writing a letter. One simulation can be presented to the user in which the street address of the recipient is omitted or replaced with a placeholder but the city and / or state of the recipient is preserved, another simulation can be presented to the user in which the recipient address (including city and state) is completely omitted or replaced with a placeholder, and so on. Other simulations can abstract all or part of the body of the letter, the sender address, and so on.
[0071] Suppose the repetitive task is to prepare and / or organize a spreadsheet to perform a number of calculations based on data contained in an input range of cells to populate an output range of cells. One simulation can include specific values and formulas in the input range of cells used to populate the output range of cells. Another simulation can not include specific values in the input range of cells, or can include fuzzy or random values, but still populate the output range of cells with the same formulas used by the user. Yet another simulation can include only the general format employed by the user in the original spreadsheet, without any of the data used or calculated by the user.
[0072] At block 306, the system, e.g., via the private evaluation engine 130A, can obtain user feedback regarding each of the plurality of different synthetic sets of interactions. In some implementations, this feedback can be obtained after each simulation (e.g., the user can be prompted to provide it), in which case no further simulations can be performed once the user accepts the latest simulation. In other implementations, the user can be presented with multiple simulations before obtaining feedback. In the latter case, the user can identify one or more of the simulations that are satisfactory and / or one or more of the simulations that are unsatisfactory (e.g., because it can inadvertently disclose sensitive or private information).
[0073] Based on the user feedback obtained at block 306, at block 308, the system, e.g., via the private evaluation engine 130A, can select one of the plurality of different synthetic sets of interactions. For example, once the user is presented with a satisfactory simulation that the user feels comfortable with, that does not result in the inadvertent disclosure of sensitive (or more generally, not widely applicable) information, the user can accept the simulation.
[0074] At block 310, the system can cause a machine learning model to be trained to generate an output indicative of the selected synthetic set of interactions. For example, at block 310A, the private evaluation engine 130A, the private embedding engine 122A, and / or the private action engine 124A can train the private embedding ML model 152A and / or the private action model 154A. At block 310B, the private evaluation engine 130A, the private embedding engine 122A, and / or the private action engine 124A can provide parameters (e.g., local gradients) of one or more of the private ML models 152A, 154A to the global embedding engine 122 and / or the global action engine 124 for federated learning of the global embedding ML model 152 and / or the global action ML model 154. When combined with other local gradients, the resulting global embedding ML model 152 and / or global action ML model 154 can be distributed to other clients to facilitate automation of the task.
[0075] Figure 4is a flowchart illustrating another example method 400 for practicing selected aspects of the present disclosure in accordance with implementations disclosed herein. The operations of the flowchart are described with reference to a system that performs the operations. The system can include various components of various computer systems, such as one or more components of the task automation system 120. Moreover, while operations of the method 400 are shown in a particular order, this is not meant to be limiting. One or more operations can be reordered, omitted, or added.
[0076] At block 402, the system may, e.g., by the private embedding engine 122A, sample (e.g., record) a plurality of interactions between a user and a computer application. The interactions may, collectively, be associated with the user performing a high-level task, such as writing a letter, reading an academic article (e.g., within or outside the user’s area of expertise), manipulating a spreadsheet or other document, etc.
[0077] At block 404, the system may, e.g., by the private embedding engine 122A, encode the plurality of interactions into a first task embedding of a first level of abstraction. In some such implementations, the private embedding engine 122A may, in performing the encoding of block 404, utilize one or more private embedding ML models 152A. These private embedding ML models 152A may, e.g., by the private selection engine 126A, be selected based on the context 103, the interaction data 104, the NL input 101 (if present), the DSK 102, etc.
[0078] At block 406, the system may, e.g., by the private action engine 124A and / or the private SIM engine 128A (which may, in some cases, be combined into a single unit), process the first task embedding using a private machine learning model (e.g., 154A) to simulate, for the user, performance of the high-level task at the first level of abstraction via one or more output devices.
[0079] At block 408, the system may, e.g., by the private SIM engine 128A and / or the private evaluation engine 130A, receive prompted or unsolicited user feedback. If the user feedback includes a rejection of the first level of abstraction, at block 410, the private embedding ML models 152A and / or the private action ML models 154A may, e.g., by the private evaluation engine 130A and / or the private action engine 124A, be trained, resulting in a first updated private ML model being generated.
[0080] At block 412, the system may, e.g., by the private action engine 124A and / or the private evaluation engine 130A, provide parameters of the first updated private machine learning model to, e.g., the global embedding engine 122 and / or the global action engine 124, for federated learning of the global embedding ML models 152 and / or the global action ML models 154.
[0081] At block 414, the system can determine, e.g., by the private assessment engine 130A, whether the user accepts or rejects the simulated performance of the high-level task at the first level of abstraction. If the answer is yes, the method 400 can end. However, if the answer at block 414 is no, the method 400 can return to block 402 and the process can repeat. For example, in response to user input rejecting the first level of abstraction, the same sampled plurality of interactions or a different sampled plurality of interactions (e.g., in the case of the user rejecting a particular interaction) can be encoded, e.g., by the private embedding engine 122A, into a second task embedding at a second level of abstraction different from the first level of abstraction. The second task embedding can be used to simulate, for the user, performance of the high-level task at the second level of abstraction via one or more of the output devices. The process can repeat for N iterations as described previously, as Figure 4
[0082] Figure 5 is a block diagram of an example computing device 510 that can optionally be utilized to perform one or more aspects of the technology described herein. In some implementations, the client device 110, the task automation system 120, and / or other components can include one or more components of the example computing device 510.
[0083] The computing device 510 typically includes at least one processor 514 which communicates with a number of peripheral devices via bus subsystem 512. These peripheral devices can include a storage subsystem 524, including, for example, a memory subsystem 525 and a file storage subsystem 526, user interface output devices 520, user interface input devices 522, and a network interface subsystem 516. The input and output devices allow for user interaction with the computing device 510. Network interface subsystem 516 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.
[0084] User interface input devices 522 can include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen incorporated into the display, audio input devices such as voice recognition systems, microphones, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways to input information to the computing device 510 or to a communication network.
[0085] User interface output device 520 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide non-visual displays, such as via an audio output device. Generally, the term "output device" is used to encompass all possible types of devices and methods for outputting information from computing device 510 to a user or another machine or computing device.
[0086] Storage subsystem 524 stores the programming and data structures that provide some or all of the functionality of the modules described herein. For example, storage subsystem 524 may include programming and data structures for performing... Figure 3 Method 300 Figure 4 The logic of selecting method 400 and / or other methods described herein.
[0087] These software modules are typically run independently by processor 514 or in combination with other processors. The memory 525 used in storage subsystem 524 may include multiple memories, including main random access memory (RAM) 530 for storing instructions and data during program execution and read-only memory (ROM) 532 for storing fixed instructions. File storage subsystem 526 can provide persistent storage for program and data files and may include hard disk drives, floppy disk drives, and associated removable media, CD-ROM drives, optical drives, or removable media cartridges. Modules implementing the functionality of certain embodiments may be stored by file storage subsystem 526 within storage subsystem 524 or on other machines accessible to processor 514.
[0088] Bus subsystem 512 provides a mechanism for enabling various components and subsystems of computing device 510 to communicate with each other as intended. Although bus subsystem 512 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0089] The computing device 510 can be of various types, including workstations, servers, computing clusters, blade servers, server groups, or any other data processing system or computing device. Due to the constantly evolving nature of computers and networks, Figure 5 The description of the computing device 510 depicted herein is intended only as a specific example for illustrating some implementations. Many other configurations of the computing device 510 may have... Figure 5 The computing device depicted in the text has more or fewer components.
[0090] While several embodiments have been described and illustrated herein, a variety of other embodiments can be constructed and utilized to perform the functions and / or achieve the results and / or advantages of the embodiments described herein, and each of such variations and / or modifications is deemed to be within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the teachings are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that other embodiments may be developed without departing from the scope of the disclosure. Embodiments of the present disclosure relate to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the scope of the present disclosure.
Claims
1. A method implemented using one or more processors and comprising: sampling a plurality of interactions between a user and a computer application, wherein the interactions collectively relate to the user performing a high-level task; encoding the plurality of interactions as one or more first task embeddings of a first level of abstraction; processing one or more of the first task embeddings using a private machine learning model to simulate, via one or more output devices, performance of the high-level task of the first level of abstraction by the user; training the private machine learning model based on rejecting a user input of the first level of abstraction, wherein the training generates an updated private machine learning model; and providing parameters of the updated private machine learning model for federated learning of a global machine learning model.
2. The method of claim 1, further comprising: in response to rejecting the user input of the first level of abstraction, encoding the plurality of interactions as one or more second task embeddings of a second level of abstraction, the second level of abstraction being different than the first level of abstraction; and prior to the training, processing one or more of the second task embeddings using the private machine learning model to simulate, via one or more of the output devices, performance of the high-level task of the second level of abstraction by the user.
3. The method of claim 2, wherein, the simulated performance of the high-level task of the second level of abstraction excludes one or more of the sampled plurality of interactions.
4. The method of claim 2 or 3, wherein, the simulated performance of the high-level task of the second level of abstraction excludes or obfuscates one or more pieces of information input or output to the user by the user during the sampling.
5. The method of any one of claims 2 to 4, wherein, a different softmax layer temperature is used to encode the first task embeddings than is used to encode the second task embeddings.
6. The method according to any one of the preceding claims, wherein, the providing comprises providing data indicative of local gradients to a remote computing system that maintains the global machine learning model.
7. The method of any of the preceding claims, wherein, the private machine learning model comprises a transformer.
8. The method of any of the preceding claims, wherein, the private machine learning model comprises a large language model (LLM).
9. The method of claim 8, wherein, the LLM predicts a token corresponding to the first plurality of interactions.
10. A method implemented using one or more processors and comprising: recording data indicative of an observed set of interactions between a user and a computing device; based on the recorded data, simulating a plurality of different synthetic sets of interactions between the user and the computing device, wherein each synthetic set comprises a variation of the observed set of interactions at a different level of abstraction; obtaining user feedback regarding each of the plurality of different synthetic sets of interactions; based on the user feedback, selecting one of the plurality of different synthetic sets of interactions; and causing a machine learning model to be trained to generate an output indicative of the selected synthetic set of interactions.
11. The method of claim 10, wherein, the simulation is performed based on the machine learning model.
12. The method of claim 11, wherein, the machine learning model is trained to facilitate intelligent processing automation.
13. The method of claim 11 or 12, wherein, the machine learning model is trained to generate a probability distribution over an action space.
14. The method of any one of claims 11-13, wherein, The machine learning model comprises a private machine learning model, and the method further comprises providing parameters of the trained private machine learning model for federated learning of a global machine learning model.
15. A system comprising one or more processors and memory storing instructions that, in response to execution of the instructions by the one or more processors, cause the one or more processors to: record data indicative of an observed set of interactions between a user and a computing device; based on the recorded data, simulate a plurality of different synthetic sets of interactions between the user and the computing device, wherein each synthetic set comprises variations of the observed set of interactions at different levels of abstraction; obtain user feedback regarding each of the plurality of different sets; based on the user feedback, select one of the plurality of different synthetic sets of interactions; and cause a machine learning model to be trained to generate an output indicative of the selected synthetic set of interactions.
16. The system of claim 15, wherein, The machine learning model is used to simulate the plurality of different synthetic sets of interactions between the user and the computing device.
17. The system of any one of claims 15-16, wherein, The machine learning model is trained to facilitate intelligent process automation.
18. The system of any one of claims 15-17, wherein, The machine learning model is trained to generate a probability distribution over an action space.
19. The system of any one of claims 15-18, wherein, The machine learning model comprises a private machine learning model.
20. The system of claim 19, further comprising instructions to provide parameters of the trained private machine learning model for federated learning of a global machine learning model.