Automating semantically relevant computing tasks across contexts
By recording and abstracting user operation sequences and combining natural language processing to generate embeddings, semantic tasks automation in different computer applications and contexts is achieved, solving the problems of high task repetition and error rates in the prior art, and improving the utilization efficiency of computing resources.
Patent Information
- Application Number
- JP2024562104
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-04-21
- Filing Date
- 2023-04-19
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2043-04-19
AI Technical Summary
The prior art is difficult to automate the processing of semantically similar but syntactically different computing tasks in different computer applications and contexts, resulting in repetition of tasks, high error rates and waste of resources.
Cross-domain automation of actions is achieved by recording the user's operation sequence in a computer application and abstracting it into "action embedding" in a general "action embedding space", and combining natural language processing to generate "tasks" or "policy" embeddings.
It realizes semantic tasks automation in different computer applications and contexts, reduces the repetition and error rate of user operations, and improves the efficiency of computing resources utilization.
Smart Images

Figure 2025514092000001_ABST
Abstract
Description
[Technical field]
[0001] Individuals often operate computing devices to perform semantically similar tasks in different contexts. For example, an individual may engage in a series of actions using a first computer application to perform a given task, such as setting various application preferences, retrieving / viewing particular data made accessible by the first computer application, performing a series of actions within a particular domain (e.g., 3D modeling, graphics editing, word processing), etc. The same individual may later engage in a series of semantically similar but syntactically different actions to perform the same or semantically similar task in a different context, such as while using a different computer application. Repeatedly performing actions involving these tasks can be tedious, error-prone, and unnecessarily consuming computing resources and / or the individual's attention.
[0002] Many computer applications provide users with the option to record a sequence of actions so that the sequence can be automated, for example, using a scripting language embedded in the computer application. These recorded sequences are sometimes referred to as "macros." However, these recorded sequences of actions and / or the scripts they generate can suffer from a variety of shortcomings. They tend to be constrained to operate within a particular computer application and are often narrowly tailored to a very specific context. Furthermore, their underlying scripts tend to be too complex to be understood, much less manipulated, by individuals unfamiliar with computer programming. Summary of the Invention
[0003] Described herein are embodiments for automating semantically similar computing tasks across multiple contexts. More particularly, but not exclusively, described herein are embodiments for enabling an individual (often referred to as a "user") to authorize or request a set of actions to be performed to fulfill or accomplish a task that is captured (e.g., recorded) in one context, e.g., in a given computer application, in a given domain, etc., without requiring programming knowledge, and seamlessly extended to other contexts. In various embodiments, the captured set of actions can be abstracted as an "action embedding" in a generalized "action embedding space." This domain-independent action embedding can represent, in abstraction, a "semantic task" that can be transformed into the action space of any number of domains using the respective domain models. In other words, a "semantic task" is a domain-independent higher-order task that finds a representation in a particular domain as a set / several domain-specific actions.
[0004] Along with the captured sequence of actions (captured with the user's permission or at the user's request, as described above), the individual can provide natural language input, e.g., spoken or typed, that provides additional semantic context for these captured sequence of actions. Natural language processing (NLP) can be performed on these natural language inputs to generate "task" or "policy" embeddings, which can then be associated with the simultaneously created action embeddings. It is then possible for the individual to provide natural language input in a different context that can be matched to one or more task / policy embeddings. The matched task / policy embeddings may be used to identify corresponding action embeddings in a generalized action embedding space. These corresponding action embeddings may be processed using a domain model associated with the current domain / context, in which the individual operates to select multiple actions from the action space of the current domain that are syntactically different but may be semantically equivalent to the original sequence of actions captured in the previous domain.
[0005] In some embodiments, a method may be implemented using one or more processors, comprising: obtaining an initial natural language input and a first plurality of actions to be performed using a first computer application; performing natural language processing (NLP) on the initial natural language input to generate a first task embedding representing a first task conveyed by the initial natural language input; processing the first plurality of actions using a first domain model to generate a first action embedding representing the first plurality of actions to be performed using the first computer application, the first domain model being trained to transform between an action space of the first computer application and an action embedding space that includes the first action embedding; and processing the first plurality of actions using a first domain model to generate a first action embedding representing the first task embedding and the first action embedding. The method may include storing in memory an association between the first and second embeddings; performing NLP on the subsequent natural language input to generate a second task embedding representing a second task conveyed by the subsequent natural language input; determining that the second task semantically corresponds to the first task based on a similarity measure between the first task embedding and the second task embedding; in response to the determination, processing the first action embedding using a second domain model to select a second plurality of actions to be performed using the second computer application, the second domain model being trained to transform between an action space and an action embedding space of the second computer application; and performing the second plurality of actions using the second computer application.
[0006] In various implementations, at least one of the first computer application and the second computer application may be an operating system. In various implementations, a first plurality of actions performed using the first computer application may be intercepted from data exchanged between the first computer application and the underlying operating system. In various implementations, the exchanged data may include data indicative of keystrokes and pointing device input.
[0007] In various embodiments, the first plurality of actions performed using the first computer application may be captured from an application programming interface (API) of the first computer program. In various embodiments, the first plurality of actions performed using the first computer application may be captured from a domain-specific programming language associated with the first domain. In various embodiments, the first plurality of actions performed using the first computer application may be captured from a scripting language embedded in the first computer application.
[0008] In various implementations, the first plurality of actions performed using the first computer application may include interactions with a first graphical user interface (GUI) rendered by the first computer application, and in various implementations, the second plurality of actions performed using the second computer application may include interactions with a second GUI rendered by the second computer application.
[0009] In various embodiments, the first computer application may be operable to exchange data with a first database having a first database schema and the second computer application is operable to exchange data with a second database having a second database schema different from the first database schema. In various embodiments, the first plurality of actions may interact with first data from the first database according to the first database schema and the second plurality of actions may interact with second data from the second database according to the second database schema, the second data semantically corresponding to the first data.
[0010] In various implementations, the first computer application may be a first communications application operated to communicate with a first plurality of contacts, and the second computer application may be a second communications application operated to communicate with a second plurality of contacts. In various implementations, the second task may determine past correspondence with one or more contacts in the second plurality of contacts. In various implementations, the second task may also determine past correspondence with one or more contacts in the first plurality of contacts.
[0011] In various embodiments, a first computer application may be operable to exchange data with a first database having a first database schema and a second computer application is operable to exchange data with a second database having a second database schema different from the first database schema. In various embodiments, a first plurality of actions may interact with first data from the first database according to the first database schema and a second plurality of actions may interact with second data from the second database according to the second database schema, and the second data may semantically correspond to the first data.
[0012] In another aspect, a method implemented using one or more processors includes obtaining an initial natural language input and a first plurality of actions to be performed using a first input form configured for a first domain; performing NLP on the initial natural language input to generate a first policy embedding representing a first input policy conveyed by the initial natural language input; processing the first plurality of actions using a first domain model to generate a first action embedding representing the first plurality of actions performed using the first input form, the first domain model being trained to transform between an action space of the first domain and an action embedding space including the first action embedding; and converting between the first policy embedding and the first action embedding. storing the association in a memory, performing NLP on the subsequent natural language input to generate a second policy embedding representing a second policy conveyed by the subsequent natural language input, determining that the second policy semantically corresponds to the first policy based on a similarity measure between the first policy embedding and the second policy embedding, and in response to the determination, processing the first action embedding using a second domain model to select a second plurality of actions to be performed using a second input form configured for the second domain, the second domain model being trained to transform between an action space and an action embedding space of the second domain, and performing the second plurality of actions using the second input form. In various implementations, the first plurality of actions may include inputting a first set of values into the first plurality of form fields, and the second plurality of actions includes inputting at least some of the first set of values into the second plurality of form fields.
[0013] Additionally, some embodiments include one or more processors of one or more computing devices, the one or more processors operable to execute instructions stored in associated memory, the instructions configured to cause any of the aforementioned methods to be performed. Some embodiments include at least one non-transitory computer-readable storage medium storing computer instructions executable by the one or more processors to perform any of the aforementioned methods.
[0014] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are contemplated as being part of the subject matter disclosed herein, for example, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein. [Brief description of the drawings]
[0015] [Figure 1] FIG. 1 is a schematic diagram of an exemplary environment in which embodiments disclosed herein may be implemented. [Diagram 2] 1 illustrates generally one example of how data may be exchanged and / or processed to extend a task performed in one domain to additional domains, according to various embodiments. [Diagram 3] 1 illustrates generally another example of how data may be processed to perform a single task across multiple domains, according to various embodiments. [Figure 4A] 1 illustrates an example of how the techniques described herein can be used to automatically fill in an input form, according to various embodiments. [Figure 4B] 1 illustrates an example of how the techniques described herein can be used to automatically fill in an input form, according to various embodiments. [Figure 5A] 1 illustrates another example of how the techniques described herein can be used to automatically fill in input forms, according to various embodiments. [Figure 5B]1 illustrates another example of how the techniques described herein can be used to automatically fill in input forms, according to various embodiments. [Figure 6] 1 is a flow chart illustrating an exemplary method for implementing selected aspects of the present disclosure, according to embodiments disclosed herein. [Figure 7] 1 is a flow chart illustrating another exemplary method for implementing selected aspects of the present disclosure, according to embodiments disclosed herein. [Figure 8] 1 is a flow chart illustrating another exemplary method for implementing selected aspects of the present disclosure, according to embodiments disclosed herein. [Figure 9] 1 illustrates an exemplary architecture of a computing device. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] Described herein are embodiments for automating semantically similar computing tasks across multiple contexts. More particularly, but not exclusively, described herein are embodiments for enabling an individual (often referred to as a "user") to authorize or request a set of actions to be performed to fulfill or accomplish a task that is captured (e.g., recorded) in one context, e.g., in a given computer application, in a given domain, etc., without requiring programming knowledge, and seamlessly extended to other contexts. In various embodiments, the captured set of actions can be abstracted as an "action embedding" in a generalized "action embedding space." This domain-independent action embedding can represent, in abstraction, a "semantic task" that can be transformed into the action space of any number of domains using the respective domain models. In other words, a "semantic task" is a domain-independent higher-order task that finds a representation in a particular domain as a set / several domain-specific actions.
[0017] Along with the captured sequence of actions (captured with the user's permission or at the user's request, as described above), the individual can provide natural language input, e.g., spoken or typed, that provides additional semantic context for these captured sequence of actions. Natural language processing (NLP) can be performed on these natural language inputs to generate "task" or "policy" embeddings, which can then be associated with the simultaneously created action embeddings. It is then possible for the individual to provide natural language input in a different context that can be matched to one or more task / policy embeddings. The matched task / policy embeddings may be used to identify corresponding action embeddings in a generalized action embedding space. These corresponding action embeddings may be processed using a domain model associated with the current domain / context, in which the individual operates to select multiple actions from the action space of the current domain that are syntactically different but may be semantically equivalent to the original sequence of actions captured in the previous domain.
[0018] As one non-limiting example, a user may allow a local agent computer program (referred to herein as an "automation agent") to capture a series of actions performed by the user using a graphical user interface (GUI) of a first computer application to set various application parameters, such as setting visual parameters to "dark mode," setting application permissions (e.g., location, camera access, etc.), or setting other application preferences (e.g., Celsius vs. Fahrenheit, metric vs. English, preferred font, preferred sort order, etc.). Many of these various application parameters may not be unique to that particular computer application, and other computer applications with similar functionality may have semantically similar application parameters. However, the semantically similar application parameters of other computer applications may be named, organized, and / or accessed differently (e.g., different submenus, command line inputs, etc.).
[0019] Using the techniques described herein, a user may provide natural language input, for example, to describe a sequence of actions performed using the GUI of a first computer application while performing them, or immediately before or after. A first task / policy embedding generated from the NLP of this input may be associated (e.g., mapped, combined) with a first action embedding generated from the sequence of actions captured using the first domain model. As previously mentioned, the first domain model may transform between a general action embedding space and the action space of the first computer application.
[0020] Later, when operating a second computer application having similar functionality to the first computer application, the user may provide a semantically similar natural language input. A second task / policy embedding generated from this subsequent natural language input may be matched to the first task / policy embedding and thus the first action embedding. The first action embedding may then be processed using a second domain model that translates between the general action embedding space and the action space of the second computer application to select actions to be performed in the second computer application. In some implementations, these selected actions may be performed automatically and the user may then be prompted to provide feedback regarding the resulting state of the second computer application. This feedback may be used, for example, to train the second domain model.
[0021] The techniques described herein are not limited to automating semantically similar tasks across separate computer applications. Other types of different contexts and domains are contemplated. For example, a sequence of actions performed by a user to fill in input fields of a first input form, e.g., a web page for ordering takeout, may be associated with an "input policy" that is captured at the user's request and conveyed in the natural language input provided by the user. The task / policy embedding generated from the user's natural language input may provide constraints, rules, and / or other data parameters that the user wants to preserve for extension to other domains. When the user later fills out another input form, e.g., grocery delivery, in a different domain, the user may provide natural language input conveying the same policy, such that at least some input fields of the new input form may be filled in with values from the previous form fill. In this manner, a user may, for example, create multiple different procurement policies or profiles that the user may select from in different contexts (e.g., one for making personal purchases, another for making business purchases, another for making travel purchases, etc.).
[0022] Abstracting both the captured sequence of actions and the accompanying natural language input may provide several technical advantages. If a sequence of actions performed by an individual can be abstracted into a semantically rich action embedding that captures much of the individual's intent, there is no need for the individual to provide a long, detailed natural language input. As a result, an individual can name an automated action using a word or short phrase, and yet the association between that word / phrase and the corresponding action embedding provides sufficient semantic context for cross-domain automation.
[0023] As with many artificial intelligence models, the more training data used to train the domain models, the more accurately they translate between various domains and action embedding spaces. Human-provided feedback, as described above, can provide particularly useful training data for supervised training, but may not be available in abundance due to its cost. Thus, in various implementations, in a process variably referred to herein as "self-supervised training" and "simulation," additional "synthetic" training data may be generated and used to train the domain models. These synthetic training data may include, for example, variations and / or permutations of user-recorded automation that are automatically generated and processed using the domain models. The resulting "synthetic" results may be evaluated, for example, against "ground truth" results of the original user-recorded automation and / or against user-provided natural language input to determine errors. These errors may be used to train the domain models, for example, using techniques such as backpropagation and gradient descent.
[0024] As an example, assume that an individual provides relatively simple and / or undetailed natural language input, such as words or short phrases, to describe a sequence of actions they require that have been recorded in a particular domain. Apart from the individual providing feedback on "ground truth" results that extend those recorded actions to different domains, additional synthetic training data can be generated and used to generate synthetic results that extend those recorded actions to different domains.
[0025] For example, short words / phrases provided by an individual can be used to generate and / or select longer, more detailed, and / or semantically similar synthetic natural language inputs. The process can then be reversed, i.e., the synthetic natural language inputs can be processed using NLP to generate synthetic task / policy embeddings, which can be processed as described herein to select action embeddings and generate synthetic results in one or more domains. These synthetic results may be compared to ground truth results in the same domains, and / or feedback on these synthetic results may be solicited from individuals in order to train domain models for these domains.
[0026] As used herein, a "domain" may refer to a target subject area in which a computing component is intended to operate, e.g., the scope of knowledge, influence, and / or activity around which the computing component's logic revolves. In some implementations, the domain to which a task is extended may be identified by heuristically matching keywords in the user-provided input with domain keywords. In other implementations, the user-provided input may be processed using NLP techniques, such as, for example, word2vec, Bidirectional Encoder Representations from Transformers (BERT) transformers, various types of recurrent neural networks ("RNNs," e.g., Long Short Term Memory or "LSTM," Gated Recurrent Units or "GRUs"), to generate a semantic embedding representing the natural language input. In some implementations, this natural language input semantic embedding, which may also function as a "task" or "policy" embedding as described above, may be used to identify one or more domains, for example, based on the distance in the embedding space between the semantic embedding and other embeddings associated with various domains.
[0027] In various implementations, one or more domain models may have been previously generated for each domain. For example, one or more machine learning models, such as RNNs (e.g., LSTM, GRU), BERT transformers, various types of neural networks, reinforcement learning policies, etc., may be trained based on a corpus of documentation associated with the domain. As a result of this training, one or more of the domain models may be at least bootstrapped so that they can be used to process what is referred to herein as "action embedding" to select multiple candidate computing actions for automation from an action space associated with the target domain.
[0028] FIG. 1 illustrates generally an exemplary environment in which selected aspects of the disclosure may be implemented, according to various embodiments. Any computing device illustrated in FIG. 1 or elsewhere in the figure may include logic such as one or more microprocessors (e.g., central processing units or "CPUs," graphical processing units or "GPUs," tensor processing units ("TPUs")) that execute computer-readable instructions stored in memory, or other types of logic such as application-specific integrated circuits ("ASICs"), field-programmable gate arrays ("FPGAs"). Some of the systems illustrated in FIG. 1, such as the semantic task automation system 102, may be implemented using one or more server computing devices forming what is sometimes referred to as a "cloud infrastructure," although this is not required. In other implementations, aspects of the semantic task automation system 102 may be implemented on the client device 120, for purposes such as, for example, protecting privacy, reducing latency, etc.
[0029] The semantic task automation system 102 may include several different components configured in selected aspects of the disclosure, such as a domain module 104, an interface module 106, and a machine learning ("ML" in FIG. 1) module 108. The semantic task automation system 102 may also include any number of databases for storing machine learning model weights and / or other data used to perform selected aspects of the disclosure. In FIG. 1, for example, the semantic task automation system 102 includes a database 110 that stores a global domain model and another database 112 that stores data indicative of global action embeddings.
[0030] The semantic task automation system 102 may be operatively coupled to any number of client computing devices operated by any number of users via one or more computer networks (114). In FIG. 1, for example, a first user 118-1 operates one or more client devices 120-1. A pth user 118-P operates one or more client devices 120-P. As used herein, a client device 120 may include, for example, one or more of a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device in a user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker (which may in some cases include a visual sensor and / or a touch screen display), a smart appliance such as a smart television (or a standard television with a networked dongle having an automated assistant function), and / or a user's wearable device including a computing device (e.g., a user's watch with a computing device, a user's glasses with a computing device, a virtual or augmented reality computing device). Additional and / or alternative client computing devices may be provided.
[0031] The domain module 104 may be configured to determine a variety of different information regarding the domain in which the user 118 is currently operating, the domain in which the user would like to extend a semantic task, etc., the domain that is relevant to a given user 118 at a given time. To this end, the domain module 104 may collect contextual information regarding, foreground and / or background applications running on the client device 120 operated by the user 118, web pages currently / recently visited by the user 118, domains to which the user 118 has access and / or frequently accessed domains, etc.
[0032] Using this collected context information, in some implementations, the domain module 104 may be configured to identify one or more domains associated with the natural language input provided by the user. For example, a request to record a task performed by the user 118 using a particular computer application and / or a particular input form may be processed by the domain module 104 to identify the domain in which the user 118 performed the task to be recorded, which may be the domain of the particular computer application or input form. If the user 118 later requests that the same task be performed in a different target domain, e.g., using a different computer application or a different input form, the domain module 104 may identify the target domain.
[0033] In some implementations, the domain module 104 may also be configured to retrieve domain knowledge from a variety of different sources related to the identified domain. In some such implementations, this retrieved domain knowledge (and / or embeddings generated therefrom) may be provided to downstream components, for example, in addition to the natural language input or contextual information discussed above. This additional domain knowledge may enable downstream components, particularly machine learning models, to be used to make predictions that are more likely to be satisfactory (e.g., to extend semantic tasks across different domains).
[0034] In some implementations, the domain module 104 can apply the collected context information (e.g., current state) across one or more “domain selection” machine learning models 105 that are distinct from the domain models described herein. These domain selection machine learning models 105 can take various forms, such as various types of neural networks, support vector machines, random forests, BERT transformers, etc. In various implementations, the domain selection machine learning models 105 can be trained to select applicable domains based on attributes (or “context signals”) of the current context or state of the user 118 and / or client device 120. For example, if the user 118 is interacting with an input form of a particular website to procure a product or service, the uniform resource locator (URL) of that website, or attributes of the underlying webpage, such as keywords, tags, document object model (DOM) elements, etc., can be applied as input across the model, either in their native form or in a dimensionality-reduced embedding. Other context signals that may be considered include, but are not limited to, the user's IP address (e.g., work vs. home vs. mobile IP address), time of day, social media status, calendar, email / text messaging content, etc.
[0035] Interface module 106 may provide one or more graphical user interfaces (GUIs) that may be operated by various individuals, such as users 118-1-118-P, to perform various actions made available by the semantic task automation system. In various embodiments, users 118 may not operate a GUI (e.g., a standalone application or a web page) provided by interface module 106 to opt-in or utilize various techniques described herein. For example, users 118-1-118-P may be required to provide explicit permission before any tasks they perform using client devices 120-1-120-P are recorded and automated as described herein.
[0036] The ML module 108 can access data indicative of various global domain / machine learning models / policies in the database 110. These trained global domain / machine learning models / policies can take a variety of forms, including, but not limited to, graph-based networks such as graph neural networks (GNNs), graph attention neural networks (GANNs), or graph convolutional neural networks (GCNs), sequence-to-sequence models such as encoder-decoders, various recurrent neural networks (e.g., LSTMs, GRUs, etc.), BERT transformer networks, reinforcement learning policies, and any other type of machine learning model that may be applied to facilitate selected aspects of the present disclosure. The ML module 108 may process various data based on these machine learning models at the request or command of other components, such as the domain module 104 and / or the interface module 106.
[0037] Each client device 120 may operate at least a portion of what is referred to herein as an "automation agent" 122. The automation agent 122 may be a computer application operable by a user 118 to perform selected aspects of the present disclosure to facilitate the extension of semantic tasks across different domains. For example, the automation agent 122 may receive a request and / or permission from a user 118 to record a series of actions performed by the user 118 using the client device 120 to complete some task. Without such explicit permission, the automation agent 122 may not be able to monitor the user's activities.
[0038] In some implementations, the automation agent 122 may take the form of what are often referred to as "virtual assistants" or "automation assistants" configured to engage in human-computer natural language interactions with the user 118. For example, the automation agent 122 may be configured to semantically process natural language input provided by the user 118 to identify one or more intents. Based on these intents, the automation agent 122 may perform various tasks, such as operating a smart device, retrieving information, performing a task, etc. In some implementations, the interactions between the user 118 and the automation agent 122 (or a separate automation assistant accessible to / by the automation agent 122) may constitute a set of tasks that may be captured and abstracted into domain-independent embeddings, as described herein, and then extended to other domains.
[0039] For example, a human-computer interaction between a user 118 and an automated agent 122 (or a separate automated assistant, or even between an automated assistant and a third-party application) to order a pizza from a third-party agent (hence, a first domain) at a first restaurant may be captured and used to generate an "order pizza" action embedding. This action embedding may later be extended to ordering pizza from a different restaurant, for example, via an automated assistant or via a separate interface.
[0040] In FIG. 1, each of the client devices 120-1 may include an automation agent 122-1 that provides services to the first user 118-1. The first user 118-1 and its automation agent 122-1 may have access to and / or be associated with a “profile” that includes various data relevant to performing selected aspects of the present disclosure on behalf of the first user 118-1. For example, the automation agent 122 may access one or more edge databases or data stores associated with the first user 118-1, including an edge database 124-1 that stores local domain models and action embeddings, and / or another edge database 126-1 that stores recorded actions. The other users 118 may have similar configurations. Any of the data stored in the edge databases 124-1 and 126-1 may be stored partially or entirely on the client device 120-1, for example, to protect the privacy of the first user 118-1. For example, the recorded actions 126-1 may include confidential and / or personal user information of the first user 118-1, such as payment information, address, phone number, etc., and may be stored locally in its raw form on the client device 120-1.
[0041] The local domain models stored in edge database 124-1 may include, for example, local versions of the global models stored in global domain model database 110. For example, in some implementations, global models may be propagated to edges for purposes of bootstrapping automated agents 122 to extend tasks to new domains associated with those propagated models, after which local models at the edges may or may not be trained locally based on activity and / or feedback of users 118. In some such implementations, the local models (alternatively referred to as "local gradients" in edge database 124) may be used periodically to train the global models (in database 110), for example as part of a federated learning framework. Because the global models are trained based on the local models, the global models may, in some cases, be propagated back to other edge databases (124), thereby keeping the local models up to date.
[0042] However, it is not a requirement in all embodiments that federated learning be employed. In some implementations, the automation agent 122 can provide scrubbed data to the semantic task automation system 102, and the ML module 108 can remotely apply the model to the scraped data. In some implementations, the "scrubbed" data can be data from which sensitive and / or personal information has been removed and / or obfuscated. In some implementations, personal information can be scrubbed at the edge by the automation agent 122, for example, based on various rules. In other implementations, the scrubbed data provided by the automation agent 122 to the semantic task automation system 102 can be in the form of dimensionality-reduced embeddings generated from raw data at the client device 120.
[0043] As previously mentioned, the edge database 126-1 can store actions recorded by the automation agent 122-1. The automation agent 122-1 can record actions in a variety of different ways depending on the access level of the automation agent 122-1 to computer applications executing on the client device 120-1 and the permissions granted by the user 118. For example, most smart phones include an operating system (OS) interface for providing or revoking permissions (e.g., location, camera access, etc.) to various computer applications. In various implementations, such an OS interface can be operable to provide / revoke access to the automation agent 122 and / or to select the particular level of access the automation agent 122 has to particular computer applications.
[0044] The automation agent 122-1 may have various levels of access to the work of the computer application, depending on the permissions granted by the user 118, as well as cooperation from the software developer providing the computer application. Some computer applications may provide the automation agent 122 with “veiled” access, for example, with the permission of the user 118, to the application's API or to scripts written using a programming language (e.g., macros) embedded in the computer application. Other computer applications may not provide as much access. In such cases, the automation agent 122 may record actions in other ways, such as by capturing screenshots, performing optical character recognition (OCR) on those screenshots to identify menu items, and / or monitoring user input (e.g., interrupts captured by the OS) to determine which graphical elements were manipulated by the user 118 and in what order. In some implementations, the automation agent 122 may intercept actions performed using the computer application from data exchanged between the computer application and the underlying OS (e.g., via system calls). In some implementations, the automation agent 122 may intercept and / or access data exchanged between or used by the window manager and / or window system.
[0045] 2 shows an example of how data may be processed by and / or using various components across domains. Starting at the top left, a user 118 operates a client device 120 to provide typed or spoken natural language input NLP-1. In the latter case, the spoken utterance may first be processed using a speech-to-text (STT) engine (not shown) to generate a speech recognition output. In either case, NLP-1 may be provided to an automated agent 122.
[0046] In addition, the user 118 operates the client device 120 to request and / or allow recording of actions performed by the user 118 using the client device 120. In various embodiments, the automation agent 122 cannot record actions without receiving this permission. In some implementations, this permission may be granted on a per-application basis, much like an application is granted permission to access GPS coordinates, local files, the use of an on-board camera, etc. In other implementations, this permission may only be granted until the user 118 states otherwise, for example, by pressing a "stop recording" button, similar to recording a macro, or by providing a voice input such as "stop recording" or "done."
[0047] Once the request / permission is received, in some implementations, the automation agent 122 may approve the request / permission. A series of actions {A1, A2, ...} performed by the user 118 in domain A using the client device 120 may then be captured and stored in the edge database 126. These actions {A1, A2, ...} may take various forms or combinations of forms, such as command line inputs, and interactions with one or more graphical elements of the GUI using various types of inputs, such as pointer device (e.g., mouse) inputs, keyboard inputs, voice inputs, gaze inputs, and any other type of input capable of interacting with the graphical elements of the GUI.
[0048] In various implementations, the domain (A) in which these actions are performed may be identified by the domain module 104 using, for example, any combination of NLP-1, computer applications operated by the user 118, remote services accessed by the user (e.g., email, text messaging, social media), projects the user is working on, etc. In some implementations, the domain may be identified at least in part by areas of a simulated digital world, sometimes referred to as the "metaverse," that the user 118 virtually operates or visits. For example, the user 118 may record an action that causes his score and a brief video replay of his performance in a first metaverse game (i.e., first domain) to be posted to his social media. The user 118 may later wish to perform a semantically similar task for a completely different metaverse game (i.e., second domain), and the techniques described herein may enable the user 118 to seamlessly extend actions previously recorded in the first domain to semantically corresponding or semantically equivalent actions in the second domain.
[0049] 2, based at least in part on the natural language input, the automation agent 122 may generate a task / policy embedding T'. For example, the automation agent 122 may perform (or cause to be performed) an STT process on the speech input provided by the user 118. The resulting speech recognition output may then be processed using various natural language processing techniques, including but not limited to techniques such as word2vec, BERT transformers, etc., to generate a task / policy embedding T' that represents the semantics of what the user 118 said.
[0050] Based on the captured domain-specific actions {A1, A2, ...}, the automation agent 122 may generate an action embedding A' that semantically represents the semantic task represented by the domain-specific actions {A1, A2, ...}. The automation agent 122 may associate this action embedding A' with the task / policy embedding T' in various ways. In some implementations, these embeddings A', T' may be combined, for example, via concatenation or by being processed together, to generate a joint embedding in a joint embedding space that captures the semantics of both the natural language input from the user 118 and the actions {A1, A2, ...}. In other implementations, these embeddings A', T' may be in separate embedding spaces, i.e., a generalized action embedding space for the action embedding A' and a task / policy embedding space for the task / policy embedding T'. A mapping (e.g., a lookup table) may be stored between these two embeddings A', T' in these two embedding spaces.
[0051] At some later point in time, the user 118 may issue another natural language input NLP-2 at the client device 120 or at another computing device associated with the user 118, such as another computing device in a collaborative ecosystem of computing devices registered in the user 118's online profile. NLP-2 may be identical to, or at least semantically equivalent to, NLP-1. However, the user 118 may be operating in a different domain, domain B. Natural language processing may be performed on NLP-2 to generate another task / policy embedding T2'. The automation agent 122 may match T2' with a previous task / policy embedding T' generated from NLP-1 ("Match Embedding" in FIG. 2). This match may be based, for example, on the distance or similarity in the embedding space between T' and T2'. These distance and / or similarity measures may be calculated using various techniques, such as Euclidean distance, cosine similarity, dot product, etc.
[0052] After the automation agent 122 matches the task / policy embedding T2' generated from NLP-2 to the task / policy embedding T' generated from NLP-1, the automation agent 122 can process the action embedding A' using domain model B (or provide it to another component for processing) based on the previously created association between A' and T'. Domain model B can be trained to transform between a generic action embedding space and an action space associated with domain B. Thus, the processing action of embedding A' using domain model B can generate a probability distribution over the action space of domain B. This probability distribution can be used, for example, by the automation agent 122 to select one or more domain-specific actions {B1, B2, ...} from the action space of domain B.
[0053] Actions such as {B1, B2, ...} can be selected in a variety of ways from the action space of the domain. In some implementations, the actions may be selected in a random order or in the order of their probabilities. In some implementations, various sequences or permutations of the selected actions may be performed, for example, as part of a real-time simulation, and the outcome (e.g., success or failure) may determine which permutation is actually performed for the user 118. Once the domain model is well trained, it may be better at predicting the order in which the actions will be performed.
[0054] In either case, the selected actions {B1, B2, ...} may be provided to client device 120 by automation agent 122 so that client device 120 can execute them. In some cases, this may automatically operate interactive elements of the GUI displayed on client device 120, with the actions being rendered as they are performed. In other implementations, these GUI actions may be executed without re-rendering.
[0055] In various embodiments, simulations can be performed, for example, by the automated agent 122 and / or components of the semantic task automation system 102, to further train the domain model. More specifically, various permutations of actions can be simulated to determine synthesis results. These synthesis results can be compared, for example, to the natural language input associated with the set of actions from which the simulated permutations were selected. The success or failure of these synthesis results can be used as positive and / or negative training examples for the domain model. In this way, it is possible to train a domain model based on much more than user-recorded actions and accompanying natural language input.
[0056] An example of a simulation to generate a composite result is shown below the bottom dashed horizontal line in FIG. 2. At some point, e.g., overnight when the server computer is under a relatively low computational load and is not specifically called upon to do so, the automation agent 122 can run one or more simulations. For example, the automation agent 122 can process an action embedding A' based on a domain model C. As described above, this can generate a probability distribution over the action space of domain C. Based on this probability distribution, the automation agent 122 can select the actions {C1, C2, ...} to be performed to generate a composite result. In some implementations, the automation agent 122 can simulate multiple different permutations of the actions {C1, C2, ...} to generate multiple different composite results. Each of these composite results can then be evaluated for success or failure (or various intermediate measures thereof). The domain model C can be trained based on this evaluation.
[0057] Although the simulation is shown as being performed in FIG. 2 as part of a domain (C) that the user 118 has not yet operated on, this is not meant to be limiting. In various embodiments, the simulation may be performed in other domains, such as domain B of FIG. 2. For example, different permutations of the actions {B1, B2, ...} may be performed to determine how resilient the domain-independent semantic task is to reordering of these actions. Furthermore, it is not necessary that all selected actions in any given domain be performed to generate all composite results.
[0058] FIG. 3, from a different perspective than FIG. 2, illustrates another example of how the techniques described herein can be used to perform semantic tasks across multiple different domains. Starting at the bottom left, a user 118 operates a client device 120 (in this example, a standalone interactive speaker) and speaks a natural language command: "Find the latest messages from Redmond." An STT module 330 can perform STT processing to generate a speech recognition output. The speech recognition output may be processed by a natural language processing (NLP) module 332 to generate a text / policy embedding 334. An action embedding finder ("AEF" in FIG. 3) module 336 can match the text / policy embedding 334 to an action embedding (white star) in an action embedding space 338. In various implementations, the STT module 330, the NLP module 332, and / or the AEF module 336 can be implemented as part of an automation agent 122, as part of a semantic task automation system 102, or any combination thereof.
[0059] The action embedding space 338 may include multiple action embeddings, each action embedding represented by a block dot in FIG. 3. These domain-independent action embeddings may be abstractions of domain-specific actions recorded in the action spaces of various domains. These recorded actions, when processed based on the respective domain models, may be abstracted into embeddings shown as part of the embedding space 338. The embedding space 338 is shown as two-dimensional for purposes of illustration and understanding only. It should be understood that the embedding space 338 may in fact have as many dimensions as the individual embeddings, and may be hundreds or thousands of dimensions.
[0060] The white star represents a coordinate in the action embedding space 338 associated with the task / policy embedding 334. As can be seen in FIG. 3, this white star is actually between two action embeddings enclosed by oval 340. In some implementations, multiple action embeddings may match a single natural language input, for example, because the action embeddings are semantically similar to each other and / or were executed in response to semantically similar natural language inputs. In some implementations, multiple matching action embeddings, such as the two in oval 340, may be combined into a unified representation, for example, via concatenation or averaging, and the unified action embeddings may be processed by downstream components.
[0061] The automation agent 122 can then process or have processed the action embeddings using multiple domain models A-C, each associated with a different domain in which the user 118 communicates with others. Domain A can represent, for example, an email domain served by one or more email servers 342A. Domain B can represent, for example, a simple messaging service (SMS) or multimedia messaging service (MMS) domain served by one or more SMS / MMS servers 342B. Domain C can represent, for example, a social media domain served by one or more social media servers 342C. Any of the servers 342A-C may or may not be part of a cloud infrastructure and thus may not necessarily be tied to a particular server instance.
[0062] Processing the action embeddings selected based on domain model A may generate actions {A1, A2, ...}, as described above. Similarly, processing the action embeddings selected based on domain models B and C may generate actions {B1, B2, ...} and {C1, C2, ...}, respectively. These actions may be executed by servers 342A-C in their respective domains. As a result, email server 342A may retrieve and return to client device 120 (e.g., via automation agent 122) the most recent email in which someone named "Redmond" was the sender or recipient. SMS / MMS server 342B may retrieve and return to client device 120 (e.g., via automation agent 122) the most recent text message in which someone named "Redmond" was the sender or recipient. The social media server 342C can then retrieve the most recent social media posts or messages (e.g., a “direct message” by, from, or to someone named “Redmond” who is a friend of the user 118) and return them, for example, to the client device 120 (e.g., by the automated agent 122). In some implementations, all of these returned messages can be matched and presented to the user 118. In other implementations, these returned messages may be compared to identify the most recent one, and only that message may be presented to the user 118. For example, if the client device 120 is a standalone interactive speaker without a display capability, as in FIG. 3 (or, for example, as in a vehicle), it may be advantageous to minimize the amount of output to avoid overwhelm or distract the user 118, in which case the most recent message of any of the domains may be read aloud.
[0063] 4A and 4B show an example of an input form that may be rendered, for example, by a web browser based on HTML (hypertext markup language) or XML (extensible markup language) when a user (not shown) visits the website "www.hyptotheticalair.com" to purchase a plane ticket. A fairly standard input form is shown that requests the user's personal information and associated payment information. It can be assumed that the user interacting with this input form has created several payment profiles or policies, each corresponding to different payment information. For example, the user may use the techniques described herein to record filling out one form (i.e., performing a set of actions) to create a "work profile" in which a "work" credit card was used, and filling out another form (i.e., performing a different set of actions) to create a "marketing profile" in which a different "marketing" credit card was used.
[0064] 4A, the input fields of the input form are pre-populated as indicated by the natural language output provided by the automation agent 122 (not shown), which states, "I used your main work profile to fill out this form. Is this correct?" This may be because the domain module 104 has applied various contextual signals, such as the URL "hypotheticalair.com," or attributes of the underlying web page (e.g., tags, keywords, DOM objects, topics, etc.), as inputs across the domain selection model 105 to select the domain with which the user's email work profile is associated.
[0065] The user responds, "No, I want to use my marketing profile." As a result, in FIG. 4B, the values used to pre-populate the input form have been changed to the user's marketing profile. The automation agent 122 responds, "Okay, I updated the form with your marketing profile. Is this correct?" The user responds, "Yes." In some implementations, the user's affirmative response may cause the input form to be submitted, for example to complete or advance a different stage of a purchase. Additionally or alternatively, in some implementations, the automation agent 122 may capture various features of the input form, such as the website URL, the layout of the input fields, or other contextual features of the input form. Using these extracted contextual features, the automation agent 122 may train one or more domain selection models 105, such that future visits to the same website or semantically similar websites in a similar context will cause the user's marketing profile information to be used to auto-populate the input form instead of the user's main work profile.
[0066] The domain model (as opposed to the domain selection model 105) may also be trained based on user feedback: for example, if a user identifies a particular field that was entered incorrectly (e.g., an incorrect expiration date for a credit card used), a domain-specific model may be trained based on that error, for example, using gradient descent and / or backpropagation.
[0067] 5A and 5B show further examples of input forms that may be rendered, for example, by a web browser based on HTML or XML, when a user (not shown) visits the website "www.hypotheticalpizza.com" to order pizza. In this example, it can be assumed that the user has not yet set up a "personal" profile for procuring / purchasing personal items, in particular food. Thus, in FIG. 4A, the user has manually filled in the input fields to include, among other things, the address and credit card information that the user would like to use to purchase pizza in this example. In addition, the user has provided a natural language input to the automated agent 122 (not shown) of "I would like to use this personal profile when purchasing food."
[0068] The automation agent 122 responds, "Okay, I'll default this profile when I confirm that you are ordering food." Using techniques described herein, the automation agent 122 can then capture the actions performed by the user to fill out these fields. The automation agent 122 can perform techniques described herein to associate action embeddings that abstract these actions with all or a portion of the user's natural language input, such as a "personal profile." This domain-independent action embedding can later be extended to other domains, such as other websites operable to order other types of food, such as grocery stores, different restaurants (e.g., other pizza restaurants or other types of restaurants), as described herein.
[0069] In particular, the automation agent 122 in this example can retroactively record actions previously performed by the user instead of recording actions that occur following the user's natural language input. In some implementations, the automation agent 122 or another component can, with explicit permission or opt-in by the user, maintain a stack or buffer of actions performed by the user, for example, when filling out an input form. If the user, after performing these actions, decides that he or she wants to record them for extension across different domains, the user can make a declaration such as the one shown in FIG. 5A, and this stack or buffer can be used to retroactively retrieve actions already performed for automation of semantic tasks.
[0070] FIG. 5B shows another input form, this time for a website with the URL "hypotheticalgrocery.com." Similar to "hypotheticalpizza.com," "hypotheticalgrocery.com" is about food, albeit in different entities (hypotheticalpizza vs. hypotheticalgrocery). As shown in FIG. 5A, the automation agent 122 has auto-populated the input fields with the same payment information used in FIG. 5A when the user created his or her "personal profile." More specifically, the domain module 104 has applied various contextual signals to the domain selection model 105 to identify the domain of the input form. The automation agent 122 then applies the action embeddings generated with respect to FIG. 5A as inputs across a domain model associated with the identified domain (e.g., a general "food" or "grocery" domain). The domain model may be trained to transform between the action embedding space and the identified domain. As a result, the application of the action embeddings across the domain model generates a probability distribution across actions in the identified domain. These actions include filling in the input fields of FIG. 5B, as shown.
[0071] In particular, despite the fact that two of the name input fields in FIG. 5B differ from their corresponding fields in FIG. 5A ("Given name" instead of "First" and "Surname" instead of "Last"), the abstraction of actions performed on the input form in FIG. 5A allows these details to be abstracted from the domain-independent action embedding resulting from FIG. 5A. As a result, when this domain-independent action embedding is transformed into concrete actions in the "hypotheticalgrocery" domain, the domain model associated with the "hypotheticalgrocery" domain can translate between these semantically equivalent terms accordingly.
[0072] The automation agent 122 states, "I used your personal profile to fill out this form. Is this correct?" The user responds affirmatively. In some implementations, this may be used as a positive training example to further train the domain model used to auto-populate the input form of FIG. 5B. Additionally or alternatively, in some implementations, the automation agent 122 and / or the domain model 104 may train one or more domain selection machine learning models 105 to which the user's "personal" payment profile is applied based on the context of FIG. 5B. The "context" of FIG. 5B may include the URL "hypotheticalgrocery.com", tags or other attributes of the underlying web page, information about the user, information about the computing device operated by the user (e.g., home IP address vs. work IP address), etc.
[0073] 6 is a flow chart illustrating an example method 600 for implementing selected aspects of the disclosure, according to embodiments disclosed herein. For convenience, the operations of the flow chart are described with reference to a system that performs the operations. The system may include various components of various computer systems, such as one or more components of the semantic task automation system 102. Additionally, although the operations of the method 600 are shown in a particular order, this is not meant to be limiting. One or more operations may be rearranged, omitted, or added.
[0074] At block 602, the system may record a first plurality of actions performed by the user 118 in a first computer application in response to a request by the user 118. At block 604, the system may receive an initial natural language input, for example by the automated agent 122, conveying information about a task performed or to be performed by the user 118. This natural language input may be received as typed text or spoken speech. In some implementations where the client device 120 being used also includes a camera, the user 118 may provide gestures) or other visual cues (e.g., sign language) as additional input.
[0075] At block 606, the system may perform NLP, e.g., by the automation agent 122 or the ML module 108, on the initial natural language input to generate a first task (or policy) embedding that represents a first task (or policy) conveyed by the initial natural language input. For example, the automation agent 122 or the ML module 108 may process the natural language input using an NLP machine learning model and / or techniques such as word2vec, BERT transformers, etc., to generate the first task (or policy) embedding.
[0076] At block 608, the system, e.g., by the automation agent 122 and / or the ML module 108, may process the first plurality of actions using the first domain model (e.g., selected by the domain module 104 using one or more domain selection machine learning models 105) to generate a first action embedding. The first action embedding may represent, in a dimensionality-reduced form, the first plurality of domain-specific actions performed using the first computer application. To this end, the first domain model may be trained to transform between the action space of the first computer application and an action embedding space that includes the first action embedding.
[0077] At block 610, the system may store in memory, for example by the automation agent 122, an association between the first task embedding and the first action embedding. For example, the automation agent 122 may store a single embedding (e.g., as an average or concatenation of the two) in a joint task / action embedding space that includes both the task (or policy) embedding generated at block 606 and the action embedding generated at block 608. Additionally or alternatively, in some implementations, the automation agent 122 may store a mapping between the two embeddings (e.g., as part of a lookup table).
[0078] The operations forming the method 600 of Figure 6 may be performed when a user 118 initially desires to record a semantic task. The automated semantic task may then be extended to other domains. As indicated in Figure 6 by block 700, exemplary operations for extending such semantic tasks to other domains are shown as part of the method 700 of Figure 7.
[0079] 7, in block 702, the system may perform NLP on the subsequent natural language input to generate a second task embedding representing a second task conveyed by the subsequent natural language input, for example, by the automated agent 122. As with block 606, other types of contextual signals and / or visual cues (e.g., gestures) may be used as inputs as part of the processing of block 702.
[0080] At block 704, the system may determine, for example by the automation agent 122, that the second task semantically corresponds to the first task based on a similarity measure between the first task embedding and the second task embedding. Such a similarity measure may be determined in a variety of ways, such as Euclidean distance, cosine similarity, dot product, etc.
[0081] In response to the determination of block 704, the system may identify, e.g., by the automation agent 122 or the domain module 104, one or more applicable domains in which the user desires to perform a semantic task. In many cases, this may be a single domain, e.g., the domain of a computer application currently being operated by the user in which the user desires to perform a semantic task (e.g., applying a dark theme, setting preferences, etc.). However, as shown in FIG. 3, for example, there may be multiple applicable domains in which the user desires to have a semantic task extended / performed at once.
[0082] Thus, in block 706, the system may determine whether there are more applicable domains in which the semantic task may be expanded / executed, for example, by the automation agent 122 or the domain module 104. If the answer is no, the method 700 ends. However, if the answer in block 706 is yes, the method 700 proceeds to block 708, where the next applicable domain is selected as the current domain.
[0083] At block 710, the system may process the first action embedding using a domain model associated with the current domain to select a current set of domain-specific actions to be performed in the current domain, e.g., by the automation agent 122 or the ML module 108. Similar to the first domain model described with respect to Figure 6, the current domain model may be trained to transform between the action space of the current domain (e.g., a second computer application, a second input form, etc.) and the action embedding space.
[0084] At block 712, the system may cause a second plurality of actions to be performed in the current domain, for example, by automation agent 122. If the current domain is a computer application having a GUI, the plurality of actions may be performed automatically on the GUI, for example, by automation agent 122. In some implementations, the GUI may be visually updated at each step so that the user can see the actions being performed. In other implementations, the actions may be performed without updating the GUI, so that the user sees only the end result of the plurality of actions.
[0085] In optional block 714, the system may receive feedback from the user regarding the performance of the semantic task in the current domain, for example by the automation agent 122. This feedback may be solicited by the automation agent 122 or may be provided unsolicited by the user. In block 716, the system may train a current domain model based on the feedback, for example by the automation agent 122 or the ML module 108. If the current domain model is local to the client device (e.g., stored in the edge database 124 in the federated learning framework shown in FIG. 1), the local model may be trained and may or may not be propagated (as local gradients) to the global model database 110 to train a corresponding global model.
[0086] Method 700 may return from block 716 (or 712 if blocks 714-716 are omitted) to block 706, where the system may again determine whether any applicable domains exist. In some implementations, the applicable domains may be configured by the user. For example, the user may register multiple domains, such as email, SMS / MMS, and social media, as shown in FIG. 3. In other implementations, with the user's permission or opt-in, the system may automatically determine the applicable domains, such as by domain module 104. For example, domain module 104 may reference the user's digital contact list to identify communication modalities associated with the individual's contacts (e.g., email addresses, phone numbers, social media or instant messaging handles, etc.). Assume that a user requests a recent written correspondence from Delia Sue. If Delia Sue's digital contact card includes an email address, phone number, and a gaming pseudonym, three domains may be applicable: email, SMS / MMS, and the domain where Delia Sue communicates with other gamers using her gaming pseudonym.
[0087] FIG. 8 is a flow chart illustrating another exemplary method 800 for implementing selected aspects of the present disclosure, according to embodiments disclosed herein. Method 800 represents a variation of methods 600-700 applicable to the scenarios illustrated in FIGS. 4A-4B and 5A-5B, for example, where an action performed using one input form (first domain) is transformed into a semantically equivalent action performed using another input form (second domain). For convenience, the operations of the flow chart are described with reference to a system that performs the operations. The system may include various components of various computer systems, such as one or more components of the semantic task automation system 102. Additionally, although the operations of method 800 are illustrated in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.
[0088] At block 802, the system may obtain, e.g., by the automation agent 122, an initial natural language input and a first plurality of actions to be performed using a first input form configured for a first domain. For example, in FIG. 5A, a user first provided an action and then provided the natural language input "I would like to use this personal profile when purchasing food." At block 804, the system may perform NLP, e.g., by the automation agent 122, on the initial natural language input to generate a first policy embedding representing a first input policy conveyed by the initial natural language input. In FIG. 5A, for example, a user defines a policy to be used when payment information associated with the "personal profile" purchases food.
[0089] At block 806, the system, e.g., by the automation agent 122, may process the first plurality of actions using the first domain model to generate a first action embedding representing the first plurality of actions to be performed using the first input form. As with other domain models described herein, the first domain model may be trained to translate between an action space of the first domain (e.g., the specific URL of FIG. 5A or the more general domain of "food") and an action embedding space that includes the first action embedding. At block 808, the system, e.g., by the automation agent 122, may store in memory an association between the first policy embedding and the first action embedding.
[0090] At block 810, the system may perform NLP on the subsequent natural language input to generate, e.g., by the automation agent 122, a second policy embedding representing a second policy conveyed by the subsequent natural language input. At block 812, the system may determine, e.g., by the automation agent 122, that the second policy semantically corresponds to the first policy based on a similarity measure between the first and second policy embeddings. For example, in FIG. 4A, the user declines to use the payment information associated with his / her "main work" profile and utters "I want to use my marketing profile." The speech recognition output generated from this utterance may be processed and matched (e.g., by the AEF module 336 of FIG. 3 using Euclidean distance or cosine similarity) to the policy embeddings and therefore action embeddings generated from actions performed by the user in the previous domain.
[0091] In response to the determination of block 812, in block 814, the system may process the first action embeddings using the second domain model to select, e.g., by the automation agent 122 or the ML module 106, a second plurality of actions to be performed using a second input form configured for the second domain. The second domain model may be trained to transform between the action space and the action embedding space of the second domain. In block 816, the system may cause, e.g., by the automation agent 122, to perform the second plurality of actions using the second input form. For example, once an action embedding associated with the user's marketing profile has been identified (e.g., by the AEF module 336 of FIG. 3), it has been applied as an input across the domain model associated with the user's marketing profile to auto-populate the input fields of the input form of FIG. 4B.
[0092] 9 is a block diagram of an example computing device 910 that may optionally be utilized to perform one or more aspects of the techniques described herein. In some implementations, one or more of client computing devices 120-1 through 120-P, semantic task automation system 102, and / or other components may comprise one or more components of example computing device 910.
[0093] The computing device 910 typically includes at least one processor 914 that communicates with a number of peripheral devices via a bus subsystem 912. These peripheral devices may include, for example, a storage subsystem 924 including a memory subsystem 925 and a file storage subsystem 926, user interface output devices 920, user interface input devices 922, and a network interface subsystem 916. The input and output devices enable user interaction with the computing device 910. The network interface subsystem 916 provides an interface to external networks and is coupled to corresponding interface devices in other computing devices.
[0094] The user interface input devices 922 may include pointing devices such as a keyboard, a mouse, a trackball, a touchpad, or a graphics tablet, a scanner, a touch screen integrated into a display, a voice input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 910 or a communications network.
[0095] The user interface output devices 920 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 910 to a user or to another machine or computing device.
[0096] Storage subsystem 924 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 924 may include logic for performing selected aspects of methods 600, 700, and 800 of FIGS.
[0097] These software modules are generally executed by the processor 914 alone or in combination with other processors. The memory 925 used within the storage subsystem 924 may include several memories, including a main random access memory (RAM) 930 for storing instructions and data during program execution, and a read only memory (ROM) 932 in which fixed instructions are stored. The file storage subsystem 926 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of a particular embodiment may be stored by the file storage subsystem 926, within the storage subsystem 924, or other machines accessible by the processor 914.
[0098] The bus subsystem 912 provides a mechanism for allowing the various components and subsystems of the computing device 910 to communicate with each other as intended. Although the bus subsystem 912 is illustrated generally as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0099] Computing device 910 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computing device 910 shown in Figure 9 is intended only as a specific example to illustrate some implementations. Many other configurations of computing device 910 are possible, having more or fewer components than the computing device shown in Figure 9.
[0100] Although several embodiments have been described and illustrated herein, various other means and / or structures for performing the functions and / or obtaining one or more of the results and / or advantages described herein can be utilized, and each such variation and / or modification is deemed to be within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on the particular application or applications in which the teachings are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. It is therefore to be understood that the above-described embodiments are presented by way of example only, and that, within the scope of the appended claims and equivalents thereof, the embodiments may be practiced otherwise than as specifically described and claimed. The embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.
Claims
1. 1. A method implemented using one or more processors, comprising: obtaining an initial natural language input and a first plurality of actions to be performed using a first computer application; performing natural language processing (NLP) on the initial natural language input to generate a first task embedding representing a first task conveyed by the initial natural language input; processing the first plurality of actions using a first domain model of a first domain to generate first action embeddings representing the first plurality of actions performed using the first computer application, the first domain model being trained to transform between an action space of the first computer application and an action embedding space that includes the first action embeddings; storing in a memory an association between the first task embedding and the first action embedding; performing NLP on the subsequent natural language input to generate a second task embedding representing a second task conveyed by the subsequent natural language input; determining that the second task semantically corresponds to the first task based on a similarity measure between the first task embedding and the second task embedding; processing the first action embeddings using a second domain model to select a second plurality of actions to be performed using a second computer application in response to the determination, the second domain model being trained to transform between an action space of the second computer application and the action embedding space; and causing the second plurality of actions to be performed using the second computer application.
2. The method of claim 1 , wherein at least one of the first computer application and the second computer application comprises an operating system.
3. 3. The method of claim 1 or 2, wherein the first plurality of actions performed using the first computer application are intercepted from data exchanged between the first computer application and an underlying operating system.
4. The method of claim 2 or 3, wherein the data exchanged includes data indicative of keystrokes and pointing device input.
5. The method of any one of claims 1 to 4, wherein the first plurality of actions performed using the first computer application are captured from an Application Programming Interface (API) of the first computer application.
6. 6. The method of claim 1, wherein the first plurality of actions performed using the first computer application are captured from a domain-specific programming language associated with the first domain.
7. The method of any one of claims 1 to 6, wherein the first plurality of actions performed using the first computer application are captured from a scripting language embedded in the first computer application.
8. 8. The method of claim 1, wherein the first plurality of actions performed using the first computer application include interactions with a first graphical user interface (GUI) rendered by the first computer application.
9. The method of any one of claims 1 to 8, wherein the second plurality of actions performed using the second computer application includes interactions with a second GUI rendered by the second computer application.
10. 10. The method of claim 1, wherein the first computer application is operable to exchange data with a first database having a first database schema, and the second computer application is operable to exchange data with a second database having a second database schema different from the first database schema.
11. 11. The method of claim 10, wherein the first plurality of actions interact with first data from the first database according to the first database schema and the second plurality of actions interact with second data from the second database according to the second database schema, the second data semantically corresponding to the first data.
12. 12. The method of claim 1, wherein the first computer application comprises a first communications application operated to communicate with a first plurality of contacts, and the second computer application comprises a second communications application operated to communicate with a second plurality of contacts.
13. The method of claim 12 , wherein the second task determines past correspondence with one or more contacts in the second plurality of contacts.
14. The method of claim 13 , wherein the second task also determines past correspondence with one or more contacts included in the first plurality of contacts.
15. 15. A method according to any preceding claim, wherein the first computer application is operable to exchange data with a first database having a first database schema and the second computer application is operable to exchange data with a second database having a second database schema different from the first database schema.
16. 16. The method of claim 15, wherein the first plurality of actions interact with first data from the first database according to the first database schema and the second plurality of actions interact with second data from the second database according to the second database schema, the second data semantically corresponding to the first data.
17. 1. A method implemented using one or more processors, comprising: obtaining an initial natural language input and a first plurality of actions to be performed using a first input form configured for a first domain; performing natural language processing (NLP) on the initial natural language input to generate a first policy embedding representing a first input policy conveyed by the initial natural language input; processing the first plurality of actions using a first domain model to generate first action embeddings representing the first plurality of actions performed using the first input form, the first domain model being trained to transform between an action space of the first domain and an action embedding space that includes the first action embeddings; storing in a memory an association between the first policy embedding and the first action embedding; performing NLP on the subsequent natural language input to generate a second policy embedding representing a second policy conveyed by the subsequent natural language input; determining that the second policy semantically corresponds to the first policy based on a similarity measure between the first policy embedding and the second policy embedding; processing the first action embeddings using a second domain model to select a second plurality of actions to be performed using a second input form configured for a second domain, in response to the determination, the second domain model being trained to transform between an action space of the second domain and the action embedding space; and performing the second plurality of actions using the second input form.
18. 20. The method of claim 17, wherein the first plurality of actions includes inputting a first set of values into a first plurality of form fields, and the second plurality of actions includes inputting at least some of the first set of values into a second plurality of form fields.
19. 1. A system comprising one or more processors and a memory storing instructions, the instructions causing the one or more processors to: obtaining an initial natural language input and a first plurality of actions to be performed using a first computer application; performing natural language processing (NLP) on the initial natural language input to generate a first task embedding representing a first task conveyed by the initial natural language input; processing the first plurality of actions using a first domain model to generate first action embeddings representing the first plurality of actions performed using the first computer application, the first domain model being trained to transform between an action space of the first computer application and an action embedding space that includes the first action embeddings; storing an association between the first task embedding and the first action embedding in a memory; performing NLP on the subsequent natural language input to generate a second task embedding representing a second task conveyed by the subsequent natural language input; determining that the second task semantically corresponds to the first task based on a similarity measure between the first task embedding and the second task embedding; processing the first action embeddings using a second domain model to select a second plurality of actions to be performed using a second computer application in response to the determination, the second domain model being trained to transform between an action space of the second computer application and the action embedding space; The system causes the second plurality of actions to be performed using the second computer application.
20. 20. The system of claim 19, wherein at least one of the first computer application and the second computer application comprises an operating system.
Citation Information
Patent Citations
Intelligent automated assistant
JP2020173835A
Providing command bundle suggestions for automated assistants
JP2020530581A