Generating cross-domain guidance for navigating HCI
By capturing and translating domain-specific actions into domain-independent embeddings, the method facilitates efficient task navigation across different computer applications, addressing user interface familiarity issues and improving task performance through continuous learning.
Patent Information
- Application Number
- JP2026081133
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-15
- Filing Date
- 2026-05-13
- Publication Date
- 2026-08-25
AI Technical Summary
Individuals face challenges in performing semantically similar tasks across different computer applications due to unfamiliarity with the application's interface, leading to inefficiencies and difficulties in navigating human-computer interactions.
A method that captures actions in one context and abstracts them into domain-independent action embeddings, which can be translated into other domains using domain models, providing guidance through visual and audible annotations or natural language outputs to facilitate task navigation across different computer applications.
Enables seamless transition of tasks across various applications by leveraging learned interactions, enhancing user familiarity and efficiency in performing similar tasks, with continuous model training for improved guidance based on user feedback.
Smart Images

Figure 2026136170000001_ABST
Abstract
Description
Technical Field
[0001] Individuals often operate computing devices to perform semantically similar tasks in different contexts. For example, an individual may be involved in a series of actions to perform a given semantic task, such as setting preferences for various applications, retrieving / viewing specific data accessible by a first computer application, performing a series of operations within a specific domain (e.g., 3D modeling, graphics editing, word processing), etc., using the first computer application. The same individual may later be involved in a series of semantically equivalent but syntactically different actions to perform a semantically similar task (e.g., the same semantic task) in a different context, such as while using a second computer application. However, the individual may not be very familiar with the second computer application, and as a result, may not be able to perform the semantic task.
Summary of the Invention
[0002] This specification describes embodiments for automatically generating and providing guidance for navigating a human-computer interface (HCI) to perform semantically equivalent and / or semantically similar computing tasks across different computer applications. More specifically, this specification describes, but is not limited to, embodiments that enable an individual (often referred to as a “user”) to leverage actions taken in one context (e.g., when performing a semantic task) to generate guidance for performing semantically identical or similar tasks in other contexts. In various embodiments, captured actions can be abstracted as “action embeddings” in a generalized “action embedding space.” This domain-independent action embedding can, in its abstraction, represent “semantic tasks” that can be translated into action spaces for any number of domains using their respective domain models. In other words, a “semantic task” is a domain-independent higher-order task that finds representations within a particular domain as a set / group of domain-specific actions.
[0003] In some embodiments, the method may be implemented using one or more processors and may include: identifying a first domain of a first computer application operable using a first human-computer interface (HCI); selecting a domain model to translate between the action space of the first computer application and another space based on the identified domain; processing action embeddings based on the selected domain model to generate one or more probability distributions over actions in the action space of the first computer program, wherein the action embeddings generate one or more probability distributions representing a plurality of actions previously performed using a second HCI of a second computer application to perform a semantic task; identifying a plurality of second actions that can be performed using the first computer application based on the one or more probability distributions; and presenting outputs in one or more output devices. In various embodiments, the outputs may include guidance for navigating the first HCI to perform a semantic task using the first computer application, the guidance being based on the identified plurality of second actions that can be performed using the first computer application.
[0004] In various embodiments, the domain model may be trained to translate between the action space of a first computer program and the domain-independent action embedding space. In various embodiments, the domain model may be trained to translate directly between the action space of a first computer program and the action space of a second computer program.
[0005] In various embodiments, the first HCI may take the form of a graphical user interface. In various embodiments, guidance for navigating the first HCI may include one or more visual annotations that overlay on the GUI. In various embodiments, one or more of the visual annotations may be rendered to draw attention to one or more graphical elements of the GUI.
[0006] In various embodiments, the guidance for navigating the first HCI may include one or more natural language outputs. In various embodiments, the method may further include obtaining user input that conveys a semantic task and identifying action embeddings based on the semantic task. In various embodiments, the user input may be natural language input, and the method may further include performing natural language processing (NLP) on the natural language input to generate a first task embedding that represents a semantic task and determining a similarity scale between the first task embedding and action embeddings, the action embeddings being processed based on the similarity scale.
[0007] In addition, some embodiments include one or more processors in one or more computing devices, the one or more processors being operable to execute instructions stored in associated memory, the instructions being configured to perform any of the methods described above. Some embodiments include at least one non-temporary computer-readable storage medium storing computer instructions that can be executed by one or more processors to perform any of the methods described above.
[0008] It should be understood that all combinations of the aforementioned and additional concepts described in detail herein are intended to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are intended to be part of the subject matter disclosed herein. [Brief explanation of the drawing]
[0009] [Figure 1] This is a schematic diagram of an exemplary environment in which the embodiments disclosed herein may be implemented. [Figure 2] This outlines an example of how data can be exchanged and / or processed to extend tasks performed in one domain to additional domains, through various embodiments. [Figure 3A] Examples of how the techniques described herein may be used are provided to give guidance on interacting with computer applications in various embodiments. [Figure 3B] Examples of how the techniques described herein may be used are provided to give guidance on interacting with computer applications in various embodiments. [Figure 3C] Examples of how the techniques described herein may be used are provided to give guidance on interacting with computer applications in various embodiments. [Figure 3D] Examples of how the techniques described herein may be used are provided to give guidance on interacting with computer applications in various embodiments. [Figure 4] This flowchart shows an exemplary method for practicing a selected aspect of the disclosure, according to embodiments disclosed herein. [Figure 5] This is an exemplary architecture for computing devices. [Modes for carrying out the invention]
[0010] This specification describes embodiments for automatically generating and providing guidance for navigating a human-computer interface (HCI) to perform semantically equivalent and / or semantically similar computing tasks across different computer applications. More specifically, this specification describes, but is not limited to, embodiments that enable an individual (often referred to as a “user”) to leverage actions taken in one context (e.g., when performing a semantic task) to generate guidance for performing semantically identical or similar tasks in other contexts. In various embodiments, captured actions can be abstracted as “action embeddings” in a generalized “action embedding space.” This domain-independent action embedding can, in its abstraction, represent “semantic tasks” that can be translated into action spaces for any number of domains using their respective domain models. In other words, a “semantic task” is a domain-independent higher-order task that finds representations within a particular domain as a set / group of domain-specific actions.
[0011] In one non-limiting embodiment, a user may allow a local agent computer program (also referred to herein as “Agent” or “Assistant”) to monitor the user’s interactions with one or more local computer applications. This monitoring may include, for example, capturing the user’s interactions with the HCI of a first computer application, such as a graphical user interface (GUI). The user can interact with the HCI to perform a variety of semantic tasks that can be performed using the first computer application. Each semantic task may include multiple individual or atomic interactions with the HCI of the first computer application.
[0012] For example, if the first computer application is a three-dimensional (3D) design application, the semantic task may include designing a 3D structure, and atomic interactions may include, for example, navigating to specific menus, selecting specific tools from those menus, selecting specific settings for those tools, and operating those tools on a canvas. If the first computer application is a spreadsheet application, the semantic task may include, for example, creating a chart based on underlying data. Atomic interactions may include, for example, sorting the data, adding columns (e.g., with formulas that utilize existing column values for operands), selecting a range to navigate through menus, and selecting specific items from those menus to create a desired chart.
[0013] Returning to the embodiment, domain-specific actions captured in relation to a semantic task performed using the first computer application can be abstracted into action embeddings using a domain model associated with the domain of the first computer application. These action embeddings can then be translated into any number of other domain action spaces, such as the action space of the second computer program. For example, a probability distribution may be generated over actions in the action space of the second computer program. Multiple domain-specific actions that can be executed using the second computer program may be selected from the action space based, for example, on their probabilities (e.g., generated using a softmax layer of the domain model). The selected domain-specific actions of the second computer program can then be used to generate guidance for navigating the HCI provided by the second computer program.
[0014] Guidance for navigating the HCI can be generated and / or presented in a variety of ways. In some embodiments where the HCI is a GUI, for example, visual annotations may be presented that overlay all or part of the GUI. In some embodiments, these visual annotations may draw attention to graphical elements that can be manipulated (e.g., using a pointer device or finger if the display is a touchscreen) to perform a set of domain-specific actions identified from the action space of a second computer program. Visual annotations may include, for example, arrows, animations, natural language text, shapes, etc. Additionally or alternatively, audible guidance may be presented to audibly guide the user to specific graphical elements corresponding to identified domain-specific actions from the action space of a second computer program. Audible guidance may include, for example, natural language output, sounds accompanied by visual annotations (e.g., animations), etc.
[0015] Accordingly, the techniques described herein allow a user to permit an agent configured in a selected aspect of this disclosure to monitor these types of interactions with various HCIs in various domains (e.g., by "opting in"). The knowledge gained by the agent can be captured (e.g., in various domain machine learning models) and utilized to generate guidance for performing semantically similar actions in other domains. In some embodiments, the extent to which other users follow or deviate from such guidance may then be used to train domain models, for example, so that they can subsequently select "better" actions. In some cases, if a sufficient number of users follow the same (or substantially similar) guidance in a given computer application, a tool may be created that can be automatically invoked using that guidance, so that subsequent users do not need to repeat the same atomic actions provided in the guidance.
[0016] In some embodiments, a user may provide natural language input, for example, while performing, immediately before, or immediately after performing, a set of actions performed using the HCI of a first domain. For example, while running a first spreadsheet application, a user might say, "I am creating a bar chart showing net losses for the past 90 days." A first task / policy embedding generated from natural language processing (NLP) of this input can be associated (e.g., mapped, combined) with a first action embedding generated from a set of captured actions, using a first domain model associated with the first spreadsheet application. As previously mentioned, the first domain model can be translated between the action space of the first spreadsheet and, for example, a general action embedding space and / or one or more other domain-specific action embedding spaces.
[0017] Later, when running a second spreadsheet application with similar functionality to the first spreadsheet application, the user can provide semantically similar natural language input to learn how to perform semantically equivalent (or at least semantically similar) tasks using the second spreadsheet application. For example, the user might utter, "How do I create a bar chart showing net losses for the past 120 days?" The second task / policy embedding generated from this subsequent natural language input can be matched to the first task / policy embedding, and therefore the first action embedding. The first action embedding can then be processed using a second domain model that translates between the general action embedding space and the action space of the second spreadsheet application to identify the actions(s) that can be performed in the second spreadsheet application to perform the semantic task. These identified actions(s) can then be used to generate guidance for performing the semantic task in the second spreadsheet application.
[0018] For example, to perform the task of creating a bar chart showing net losses over the past 120 days, visual annotations and / or audible guidance may be provided to guide the user through various menus, sheets, cells, etc., in a second spreadsheet application. In particular, the fact that the first chart shows losses over 90 days while the second chart shows losses over 120 days can be handled by the agent, for example, by storing the number of days as a parameter associated with an action in the action space of the second spreadsheet application. In addition to capturing the meaning of the HCI itself, the domain model may also be trained to identify where semantically equivalent data resides.
[0019] For example, when operating a first spreadsheet application to edit a first spreadsheet file, the data needed to determine net loss may be on a specific tab that also contains various other data. In contrast, a second spreadsheet editable using a second spreadsheet application may contain semantically similar data—that is, the data needed to determine net loss—on different tabs, with or without other data. However, given sufficient training examples provided over time, a domain model used by an agent configured using a selected aspect of this disclosure may be able to pinpoint the appropriate location of data to determine net loss. For example, different columns in different spreadsheets containing data related to net loss may contain semantically similar column headers. Furthermore, the actual data itself may share semantic characteristics, i.e., be similarly formatted, have roughly similar values (e.g., within the same order of magnitude, millions versus hundreds of millions), or exhibit similar temporal patterns (e.g., higher sales in certain seasons).
[0020] In addition to, or instead of, the guidance, in some embodiments, the HCI itself can be configured to suit the behavior or abilities of a particular user using the techniques described herein. For example, the visual settings of a GUI can be configured through a variety of different actions to make the GUI easier to use for visually impaired users. This may include, for example, increasing font size, increasing contrast, reducing the number of menu items presented (e.g., based on frequency of use across a user population), increasing the size of operable graphical elements such as sliders and buttons, and activating user accessibility settings as voice prompts. These actions can be captured in a given computer application and abstracted into action embeddings, along with task / policy embeddings created from natural language input, such as "enforce visual impairment settings."
[0021] Later, in a different context (for example, when interacting with a different computer application), the user may provide natural language input such as, "I am visually impaired, please make this interface easier to use." The previously generated action embedding may be processed using a domain model associated with the new context to automatically perform at least some of the aforementioned adjustments and / or show the user how to do so. If any of the adjustments are not available or applicable, the user may be notified of this and / or offered a different recommendation that may satisfy a similar need.
[0022] The techniques described herein are not limited to generating guidance for performing semantic tasks across similar domains (e.g., from one spreadsheet application to another spreadsheet application). In various embodiments, guidance for performing semantic tasks may also be generated across semantically different domains / contexts. For example, although semantically similar, domain-independent application parameters of various computer applications (plural) may be named, organized, and / or accessed differently (e.g., different sub-menus, command-line inputs, etc.). Such application parameters may include, for example, visual parameters that can be set to various modes such as "dark mode", application permissions (e.g., access to location, camera, files, other applications, etc.), or settings of other applications (e.g., setting of Celsius or Fahrenheit, setting of metric or English unit system, preferred font, preferred sort order, etc.). Many of these various application parameters may not be specific to a particular computer application or domain. In fact, some application parameters such as the "skin" applied to the GUI may also be applicable to the operating system (OS).
[0023] The above example of the spreadsheet included two different spreadsheet applications. However, this was not intended to be limiting. The techniques described herein may be implemented to generate guidance for performing semantic tasks across multiple different use cases within a single domain. Assume that a user operates a first "pocket" spreadsheet to organize pocket and schedule data in a particular way and then generates a pocket report in a particular format. The actions that the user performs to create this report can be captured and abstracted into action embeddings using the domain model of the spreadsheet application that the user is operating, regardless of what that spreadsheet application is, as described above.
[0024] Subsequently, the user may receive, for example, a second "document" spreadsheet created by a different docket creation system or for a different entity. This second document spreadsheet contains data that is semantically similar to the first document spreadsheet, but may be differently organized and / or have a different schema. The columns may have different names and / or be in a different order. The data may be represented using different syntaxes (e.g., "MM / DD / YY" and "DD / MM / YYYY"). Nevertheless, previously created action embeddings can be processed, for example, in conjunction with the second document spreadsheet (e.g., as additional context input data) to generate guidance for performing the same semantic tasks using the second document spreadsheet. For example, the same domain model that can also process context input data, or a separate domain model (e.g., trained in reverse order), may be applied to identify actions that are feasible to perform semantic tasks using the second document spreadsheet.
[0025] In various embodiments, the domain model can be continuously trained based on how the guidance generated using the techniques described herein interacts with the user. This can, in turn, affect how various guidance is provided or whether it is provided at all. For a particular domain-independent semantic action such as "set to dark mode", assume that the particular proposed action is rarely or never performed in a particular domain, for example, using a particular computer application of that domain. Perhaps the native settings of that computer application already address the underlying issues that would require that proposed action in other domains or render such issues practically meaningless.
[0026] In such scenarios, a domain model associated with a particular domain may be further trained so that, when the same (or similar) action embedding is processed, the resulting probability distribution across the domain's action space assigns a lower probability to the proposed action. Conversely, in other domains where the underlying problem still exists, the proposed action may be given a higher probability. The assigned probability may indicate how the proposed action is presented to the user (e.g., how it is highlighted, whether it is an animation or a small visual annotation, whether it is auditory or visual, etc.), when the proposed action is presented to the user (e.g., for other proposals), whether the proposed action is presented to the user at all, or whether the action should be performed automatically without providing user guidance.
[0027] The continuous training of a domain model is not limited to monitoring user feedback / responses to HCI guidance provided in new domains. In some embodiments, a domain model may be trained without the user leaving the original domain in which the semantic task is performed. For example, when a user performs a series of actions using a computer application to complete a given semantic task, a domain model associated with the domain of the computer application can be used to process the actions (or data indicating them, such as embeddings) to generate a domain-independent embedding that semantically represents the given task. This domain-independent embedding can then be processed using a machine learning model (e.g., a sequence decoder) trained to produce natural language output intended to describe the given semantic task performed by the user. For example, in response to a change in a particular visual setting in an application or operating system, the user may be presented with natural language output such as, "It looks like you've changed the graphical interface to 'dark mode'."
[0028] This natural language output may be presented to the user audibly or visually, along with a request for user feedback ("Is that what you did?" or "Did I accurately describe your action?"). The user's positive feedback ("Yes, that's correct") or negative feedback ("No, that's not what I did") may be used to train the domain model using techniques such as backpropagation and gradient descent. The user may be able to adjust or influence how often (or even whether) requests for feedback are presented to them. In some cases, the user may be offered an incentive to be asked for and / or to provide such feedback. These incentives may be offered in various forms, such as monetary rewards or credits related to the computer application (e.g., special items in a game). Additionally or alternatively, the agent itself may self-adjust how often it requests such feedback from the user based on signals such as the user's response (e.g., abandonment or cooperation) or a measure of accuracy associated with the domain model in question (more accurate models may not need to be trained as frequently as less accurate models).
[0029] As used herein, “domain” can mean a target area in which a computing component is intended to operate, for example, the scope of knowledge, influences, and / or activities around which the logic of the computing component revolves. In some embodiments, a domain may be identified by heuristically matching keywords in user-provided input with domain keywords. In other embodiments, user-provided input may be processed using NLP techniques such as word2vec, Bidirectional Encoder Representations from Transformers (BERT) converters, and various types of recurrent neural networks ("recurrent neural network, RNN," e.g., Long Short-Term Memory, i.e., LSTM, Gated Recurrent Unit, i.e., GRU) to generate semantic embeddings representing the user input.
[0030] In various embodiments, one or more domain models may be previously generated for each domain. For example, one or more machine learning models such as RNNs (e.g., LSTM, GRU), BERT transformers, various types of neural networks, reinforcement learning policies, etc., may be trained on a corpus of documentation associated with the domain. As a result of this training, one or more of the domain models may be at least bootstrapped and thus available for processing what is referred to herein as “action embeddings” to select a number of candidate computing actions from the action space associated with the target domain that can be used to provide the guidance described herein.
[0031] Figure 1 schematically illustrates exemplary environments in which selected aspects of the present disclosure may be implemented in various embodiments. Any computing devices shown in Figure 1 or elsewhere in the Figure may include logic such as one or more microprocessors that execute computer-readable instructions stored in memory (e.g., a central processing unit or "CPU", a graphical processing unit or "GPU", a tensor processing unit ("tensor processing unit, TPU")), or other types of logic such as an application-specific integrated circuit ("ASIC"), a field-programmable gate array ("FPGA"). Some of the systems shown in Figure 1, such as the semantic task guidance system 102, may be implemented using one or more server computing devices that form what may be referred to as a "cloud infrastructure," but this is not required. In other embodiments, aspects of the semantic task guidance system 102 may be implemented on a client device 120 for purposes such as protecting privacy or reducing latency.
[0032] The semantic task guidance system 102 may include several different components configured in a selected aspect of the disclosure, such as a domain module 104, an interface module 106, a machine learning (ML in Figure 1) module 108, and / or a task identification (ID in Figure 1) module 110. The semantic task guidance system 102 may also include any number of databases for storing machine learning model weights and / or other data used to perform a selected aspect of the disclosure. In Figure 1, for example, the semantic task guidance system 102 includes a database 111 for storing global domain models and another database 112 for storing data indicating global action embeddings.
[0033] The semantic task guidance system 102 can be operably coupled with any number of client computing devices operated by any number of users via one or more computer networks (114). In Figure 1, for example, the first user 118-1 operates one or more client devices 120-1. The p-th user 118-P operates one or more client devices 120-P. As used herein, the client device(s) 120 may include, for example, a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device in the user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker (which may optionally include a visual sensor and / or a touchscreen display), smart devices such as a smart TV (or a standard TV with a networked dongle having automation assistant functionality), and / or a user's wearable device including a computing device (e.g., a user's watch with a computing device, a user's glasses with a computing device, a virtual or augmented reality computing device). Additional and / or alternative client computing devices may be provided.
[0034] The domain module 104 may be configured to determine various different pieces of information about domains relevant to a given user 118 at a given time, such as the domain the user 118 is currently interacting with, the domain(s) the user has previously interacted with, the domain(s) the user wants to extend semantic tasks to, or the domain(s) from which the user wants to receive guidance on how to perform semantic tasks. For this purpose, the domain module 104 may collect contextual information such as foreground and / or background applications running on the client device(s) 120 that the user 118 is interacting with, web pages the user 118 has currently / recently visited, and the domain(s) the user 118 has access to and / or frequently accesses.
[0035] Using this collected contextual information, in some embodiments, the domain module 104 may be configured to identify one or more domains currently associated with the user. For example, a request to record or observe a task performed by user 118 using a specific computer application and / or a specific input form may be processed by the domain module 104 to identify the domain in which user 118 performs the task to be recorded, and the domain may be the domain of the specific computer application or input form. If user 118 later requests guidance on performing the same task in a different target domain, for example, using a different computer application or a different input form, the domain module 104 may identify the target domain. The user does not necessarily request guidance in a different target domain. In some embodiments, the techniques described herein may be implemented to provide the user with unsolicited guidance on how to perform a similar semantic task to one previously performed by the user in a different domain, simply by operating a different computing application or input form.
[0036] In some embodiments, the domain module 104 may also be configured to extract domain knowledge from various different sources related to the identified domain. In some such embodiments, this extracted domain knowledge (and / or embedded(s) generated therefrom) may be provided to downstream components(s), for example, in addition to the aforementioned natural language input or contextual information. This additional domain knowledge may be used by downstream components(s), particularly machine learning models, to make predictions that are more likely to be satisfactory (e.g., to generate guidance for performing semantic tasks across different domains).
[0037] In some embodiments, the domain module 104 can apply the collected contextual information (e.g., current state) across one or more “domain selection” machine learning models 105 that differ from the domain models described herein. These domain selection machine learning models 105 can take various forms, such as various types of neural networks, support vector machines, random forests, and BERT transformers. In various embodiments, the domain selection machine learning models 105 can be trained to select applicable domains based on attributes (or “context signals”) of the current context or state of the user 118 and / or client device 120. For example, if user 118 is interacting with an input form on a particular website to procure goods or services, the website’s uniform resource locator (URL), or attributes of the underlying webpage (or DOM) such as keywords, tags, and document object model (DOM) elements, can be applied as input across the model, either in their native form or as dimensionality-reducing embeddings. Other contextual signals that may be considered include, but are not limited to, the user's IP address (e.g., work vs. home vs. mobile IP address), time, social media status, calendar, and email / text messaging content.
[0038] Interface module 106 may provide one or more graphical user interfaces (GUIs) that can be operated by various individuals, such as users 118-1 to 118-P, to perform various actions made available by the semantic task guidance system 102. In various embodiments, user 118 may operate a GUI provided by interface module 106 (e.g., a standalone application or a web page) without opting in to or utilizing the various technologies described herein. For example, users 118-1 to 118-P may be required to provide explicit permission before any task they perform using client device(s) 120-1 to 120-P is observed and used to generate guidance as described herein.
[0039] Additionally, interface module 106 may be configured to provide or cause to provide the guidance described above for performing a semantic task in a different domain, in practice of a selected aspect of the present disclosure. For example, interface module 106 may receive from ML module 108 one or more sampled actions from the action space of a particular domain. Interface module 106 may then cause the user to be presented with graphical and / or audio data representing these actions.
[0040] Assume a designer is operating a new computer-aided design (CAD) computer application, and that the designer previously operated an older CAD computer application, for example, as part of their employment. The actions of a given task frequently performed by the designer using the older CAD computer application can be processed by ML module 108, for example, using a domain model associated with the older CAD computing application, to generate domain-independent action embeddings. These action embeddings can then be translated, for example, by ML module 108 into the domain of the new CAD computer application to generate one or more actions (e.g., sampling) that can be performed using the new CAD computer application. These actions(s) can be used by interface module 106 to generate audio and / or visual guidance explaining to the user how to perform a given task using the new CAD computer application.
[0041] ML module 108 can access data representing various global domains / machine learning models / policies in database 111. These trained global domains / machine learning models / policies can take various forms, including, but are not limited to, graph-based networks such as graph neural networks (GNNs), graph attention neural networks (GANNs), or graph convolutional neural networks (GCNs), sequence-to-sequence models such as encoder-decoders, various recurrent neural networks (e.g., LSTM, GRU, etc.), BERT transformer networks, reinforcement learning policies, and any other types of machine learning models that may be applied to facilitate selected aspects of this disclosure. ML module 108 may process various data based on these machine learning models in response to requests or commands from other components such as domain module 104 and / or interface module 106.
[0042] The Task ID module 110 may be configured to analyze interactions between individuals and computer applications(s) collected by the semantic coordination agent 122 (described in more detail below). Based on these observations, the Task ID module 110 can determine which self-contained semantic tasks performed by an individual in one domain are likely to be performed by the same individual or other individuals in other domains. In other words, the Task ID module 110 may selectively trigger the creation of domain-independent action embeddings (e.g., by the ML module 108), which can then be used by the ML module 108 to sample actions in different domains for the purpose of providing semantic task guidance across these domains.
[0043] In some embodiments, the task ID module 110 may selectively trigger the creation of domain-independent action embeddings on an individual basis. If a particular individual appears to repeatedly perform the same semantic task in one domain, guidance for performing that semantic task in other domains may be specifically provided to that individual. Additionally or alternatively, the task ID module 110 may selectively trigger the creation of domain-independent action embeddings applicable across a group of individuals. When different individuals are observed, for example, several threshold numbers, threshold frequencies, etc., of performing the same semantic task across one or more domains may trigger the task ID module 110 to generate domain-independent embeddings. The interface module 106 and / or ML module 108 can then use these domain-independent embeddings to provide guidance for performing the semantic task to any number of different individuals.
[0044] In various embodiments, the task ID module 110 and / or the semantic coordination agent 122 can observe an individual's interactions only if the individual has given permission. For example, when installing (or updating) a computer application on a specific client device 120, the semantic coordination agent 122 may request permission from the individual to observe their interactions with the new computer application.
[0045] Each client device 120 may operate at least a portion of the semantic coordination agent 122 described above. The semantic coordination agent 122 may be a computer application that a user 118 can operate to implement selected aspects of the present disclosure in order to facilitate the extension of semantic tasks across different domains. For example, the semantic coordination agent 122 may receive a request and / or permission from the user 118 to observe / record a series of actions performed by the user 118 using the client device 120 in order to complete some task. Without such an explicit request or permission, the semantic coordination agent 122 may not be able to observe the user's interaction.
[0046] In some embodiments, the semantic adjustment agent 122 may take the form of what is often referred to as a “virtual assistant” or “automation assistant” configured to engage in human-computer natural language interaction with the user 118. For example, the semantic adjustment agent 122 may be configured to semantically process natural language input(s) provided by the user 118 to identify one or more intentions(s). Based on these intentions(s), the semantic adjustment agent 122 can perform a variety of tasks, such as operating smart devices, retrieving information, or performing tasks. In some embodiments, the interaction between the user 118 and the semantic adjustment agent 122 (or a separate automation assistant accessible by / through the semantic adjustment agent 122) may constitute a set of tasks that can be captured, abstracted into domain-independent embeddings, and then extended into other domains, as described herein.
[0047] For example, a human-computer interaction between user 118 and semantic coordination agent 122 (or a separate automated assistant, or even between an automated assistant and a third-party application) for ordering a pizza from a third-party agent of a first restaurant (and thus a first domain) may be captured and used to generate a “order pizza” action embedding. This action embedding may later be extended to ordering pizzas from different restaurants, for example, via an automated assistant or through a separate interface.
[0048] In Figure 1, each of the client devices 120-1 may include a semantic coordination agent 122-1 that provides services to a first user 118-1. The first user 118-1 and its semantic coordination agent 122-1 can access and / or associate with a “profile” containing various data related to carrying out selected aspects of the Disclosure on behalf of the first user 118-1. For example, the semantic coordination agent 122 can access one or more edge databases or datastores associated with the first user 118-1, including an edge database 124-1 that stores local domain models and action embeddings, and / or another edge database 126-1 that stores recorded actions. Other users 118 may have a similar configuration. Any data stored in edge databases 124-1 and 126-1 may be stored partially or entirely on the client device 120-1, for example, to protect the privacy of the first user 118-1. For example, the recorded action 126-1 may include confidential and / or personal information of the first user 118-1, such as payment information, address, and telephone number, and may be stored locally in its raw form on the client device 120-1.
[0049] Local domain models (or multiple models) stored in edge database 124-1 may include, for example, local versions of global models (or multiple models) stored in global domain model database 111. For example, in some embodiments, global models can be propagated to edges for the purpose of bootstrapping semantic adjustment agents 122 to extend tasks to new domains associated with those propagated models, and local models at the edges may or may not be trained locally based on user 118 activity and / or feedback. In some such embodiments, local models (in edge database 124, substituted by “local gradients”) may be used periodically to train global models (in database 111), for example, as part of a federated learning framework. Since global models are trained based on local models, global models can, in some cases, be propagated back to other edge databases (124) to keep local models up-to-date.
[0050] However, the adoption of associative learning is not a requirement in all embodiments. In some embodiments, the semantic adjustment agent 122 can provide scrubbed data to the semantic task guidance system 102, and the ML module 108 can remotely apply a model to the scrubbed data. In some embodiments, the “scrubbed” data may be data from which sensitive and / or personal information has been removed and / or obscured. In some embodiments, personal information may be scrubbed at the edge by the semantic adjustment automation agent 122 based on various rules, for example. In other embodiments, the scrubbed data provided to the semantic task guidance system 102 by the semantic adjustment agent 122 may be in the form of dimensionality reduction embeddings generated from raw data in the client device 120.
[0051] As described above, the edge database 126-1 can store actions recorded by the semantic adjustment agent 122-1. The semantic adjustment agent 122-1 can observe and / or record actions in various different ways depending on the level of access the semantic adjustment agent 122-1 has to the computer application running on the client device 120-1 and the permissions granted by the user 118-1. For example, most smartphones include an operating system (OS) interface for granting or revoking permissions to various computer applications (e.g., location, camera access, etc.). In various embodiments, such an OS interface may be operable to grant / revoke access to the semantic adjustment agent 122 and / or to select a particular level of access that the semantic adjustment agent 122 has to a particular computer application.
[0052] The semantic adjustment agent 122-1 may have varying levels of access to the computer application's operations, depending on the permissions granted by user 118 and the cooperation of the software developers providing the computer application. Some computer applications may, for example, grant semantic adjustment agent 122 "covered" access to the application's API or to scripts written using a programming language (e.g., macros) embedded in the computer application, with the permission of user 118. Other computer applications may not provide as much access. In such cases, semantic adjustment agent 122 may record actions in other ways, such as by capturing screenshots, performing optical character recognition (OCR) on those screenshots to identify menu items, and / or by monitoring user input (e.g., interrupts captured by the OS) to determine which graphical elements were manipulated by user 118 and in what order. In some embodiments, semantic adjustment agent 122 may intercept actions performed using the computer application from data exchanged between the computer application and the underlying OS (e.g., via system calls). In some embodiments, the semantic adjustment agent 122 can intercept and / or access data exchanged between the window manager and / or the window system, or data used by the window manager and / or the window system.
[0053] Figure 2 schematically illustrates an example of how data may be processed and / or used by various components across multiple domains. Starting from the top left, user 118 operates client device 120 to request or provide permission for semantic adjustment agent 122, which operates at least partially on client device 120, to observe user 118's interaction with a first computer application, app A. In various embodiments, semantic adjustment agent 122 cannot record actions without receiving this permission. In some embodiments, this permission may be granted on an application-by-application basis, much like how an application may be granted permission to access GPS coordinates, local files, use of an onboard camera, etc. In other embodiments, this permission may only be granted until user 118 states otherwise, for example, by pressing a “stop recording” button similar to recording a macro, or by providing voice input such as “stop recording” or “end.”
[0054] Once a request / permission is received, in some embodiments, the semantic coordination agent 122 may, but is not required, acknowledge (ACK) the request / permission. The user 118 can then launch application A and use client device 120 to perform a series of actions {A1, A2, ...} within domain A, which may be captured and stored in edge database 126. These actions {A1, A2, ...} may take various forms or combinations of forms, such as command-line input, as well as interaction with one or more graphical elements of a GUI using various types of input, including pointer device (e.g., mouse) input, keyboard input, voice input, eye-tracking input, and any other type of input that can interact with the graphical elements of a GUI.
[0055] In some embodiments, a group of actions performed together logically, for example, within a specific time interval, without interruption, may be grouped together as a semantic task by, for example, the task ID module 110. For example, a user may perform actions A1-A6 during one session, stop interacting with app A for a period of time, and then perform actions A7-A15. In various embodiments, actions A1-A6 may be grouped together as one semantic task, and actions A7-A15 may be grouped together as another semantic task.
[0056] In various embodiments, the domain (A) in which these actions are performed may be identified by, for example, the domain module 104 using any number of signals, such as that user 118 has launched app A, and other signals if available. These other signals may include, for example, natural language input (NLI) provided by the user, the user's calendar, the user's electronic communications, and the user's social media posts.
[0057] In various embodiments, the semantic adjustment agent 122 may observe / record actions {A1, A2, ...} and pass them (or data indicating them, such as dimensionality reduction embeddings) to another component, such as the ML module 108 (not shown in Figure 2). The ML module 108 can then process these actions using all or part of the domain model A to generate a domain-independent action embedding A' (also referred to as an "intermediate representation"). In various embodiments, the domain model A may include, for example, an encoder portion of a larger encoder-decoder architecture that can be used to process a sequence of tokens (e.g., actions {A1, A2, ...}) to generate an intermediate representation, e.g., action embedding A'.
[0058] In some embodiments, as indicated by the dashed lines, user 118 may optionally provide an NLI-1 to describe what user 118 is doing when performing an action {A1, A2, ...}. This NLI-1 may be captured by a semantic adjustment agent 122, which may pass the NLI-1 to the ML module 108 for natural language processing to generate a task embedding T'. The task embedding T' may be used to provide additional context for the actions {A1, A2, ...}. This additional context may be used in various ways, such as additional input for a domain model, or as an anchor to allow semantically similar actions to be requested in the future (by user 118 or someone else). As indicated by the dashed lines, in some embodiments, the task embedding T' and action embedding A' may be related to each other, for example, in a database, via a shared / joined embedding space, etc.
[0059] In some embodiments, if user 118 does not provide natural language input describing actions {A1, A2, ...}, the semantic adjustment agent 122 can formulate (or be made to formulate) predicted descriptions of the action(s) and then request feedback from user 118 regarding the accuracy or quality of the descriptions. In Figure 2, for example, additional dashed arrows indicate how the semantic adjustment agent 122 produced natural language output (NLO) using action embedding A' by processing (or being made to process) embedding A' using, for example, a semantic decoder machine learning mode trained to translate between a domain-independent action embedding space and a natural language vocabulary. The semantic adjustment agent 122 can then present this natural language output to user 118 as part of a request for feedback ("Request for Feedback" in Figure 2). User 118 may provide feedback (e.g., "Yes, that's correct," or "No, you're wrong"). Based on that feedback, the semantic adjustment agent 122 can train or allow the domain model A to be trained.
[0060] As indicated by the horizontal dashed line, after some time, user 118 launches application B, which causes semantic coordination agent 122 to identify domain B as the active domain. In various embodiments, semantic coordination agent 122 may use domain model B so that, for example, action embedding A' is processed by ML module 108. For example, action embedding A' may collectively form domain B or be processed using the decoder portion of an encoder-decoder network associated with domain B. This decoding may, for example, generate probability distributions across the action space of domain B. Based on these probability distributions, various actions {B1, B2, ...} selected from the action space of domain B may be generated and provided to semantic coordination agent 122. Semantic coordination agent 122 may then work with interface module 106 (not shown in Figure 2) to generate HCI guidance for user 118.
[0061] In various embodiments, components such as the semantic adjustment agent 122 and the ML module 108 may continuously train the domain model, for example, to improve the quality of the HCI guidance provided and to allow the HCI guidance to be more narrowly tailored to a particular context. When this HCI guidance is presented, it is assumed that user 118 will follow some parts of the guidance but not others, and will follow them in a different order. For example, suppose the HCI guidance is to perform actions {B1, B2, B3, B4, B5, B6, B7} in order, but in Figure 2, it is assumed that user 118 will perform fewer than all of the actions in a different order {B1, B5, B4, B3}. In some embodiments, the difference between the recommended actions {B1, B2, B3, B4, B5, B6, B7} and the actions {B1, B5, B4, B3} ultimately performed by user 118 can be used as an error, which can then be used, for example, by ML module 108 to train domain model B using techniques such as gradient descent and backpropagation.
[0062] Figures 3A–3D illustrate examples of how components such as the semantic coordination agent 122 may, or may be made to, implement selected embodiments of the present disclosure to provide guidance for performing semantic tasks in a new domain. In this example, we assume that a user (not shown) is operating HCI360 by a “virtual CAD” computer application, or in the form of a rendered GUI. Furthermore, we should assume that the user has previously worked with a different CAD computer application called “FakeCAD” (e.g., as part of the user’s employment) to perform any number of repetitive tasks. Finally, we should assume that the user has recently migrated from using FakeCAD to using virtual CAD (e.g., as a result of taking a new job).
[0063] In Figure 3A, the HCI is in the home state, with the "Home" menu item active. As a result, context-specific menu items such as "New," "Open," "Save," and "Print" are active. These are merely examples of what may be presented and are not intended to be limiting. On the left are several design tools (1-8) which may include tools commonly found in typical CAD programs, such as tools for drawing lines, ellipses, and circles, shape filling tools, and different brushes.
[0064] In this example, when operating FakeCAD, a previous CAD software, to perform various semantic tasks, domain-specific actions previously performed by the user are processed to generate domain-independent action embeddings. Specifically, a domain model trained to convert to and from the action space associated with FakeCAD was used to process these domain-specific actions into domain-independent action embeddings. Then, for example, ML module 108 processed one or more of these domain-independent action embeddings using a domain model configured to convert to and from the action space associated with the new software, Virtual CAD.
[0065] The output of processing using the virtual CAD domain model may include a probability distribution(s) across actions within the virtual CAD action space. Based on these probability distributions, the ML module 108 or the semantic adjustment agent 122 can select one or more actions within the virtual CAD action space. These selected actions(s) can then be used, for example, by the interface module 106 to generate HCI guidance that helps the user navigate HCI360 to perform tasks previously performed using FakeCAD.
[0066] For example, in Figure 3A, the HCI guidance is presented in the form of an overlaid annotation pointing to the "View" menu and explaining, "These are the items that have been consistently adjusted." If the user selects the "View" menu as suggested, HCI360 may transition to the state shown in Figure 3B.
[0067] In Figure 3B, the "View" menu is currently active, which brings up several context-specific menu items. "Zoom," "Ruler," "Mode," "Wireframe," and "Connect" are merely illustrative examples of what may be presented and are not intended to be limiting. Here, new HCI guidance is provided, prompting the user to interact with the "Ruler" and "Wireframe" menu items. This may be because, for example, the user had a habit of frequently adjusting similar parameters when using FakeCAD in the past.
[0068] In Figure 3C, the "Home" menu is active again. Here, the user is presented with additional HCI guidance related to "Tool 2". Specifically, the user is notified via an overlaid annotation that "This is the tool corresponding to tool x, which you frequently used in FakeCAD." For example, the HCI guidance previously presented in Figures 3A and 3B corresponded to actions the user had typically performed previously (e.g., setting common parameters), so the user may be presented with this HCI guidance at various points, such as after the HCI guidance presented in Figures 3A and 3B. In contrast, the HCI guidance presented in Figure 3C may correspond to actions the user would traditionally perform later, for example, after the user has set common parameters to their liking and is ready to work.
[0069] Figure 3D illustrates, for example, the next iteration of the HCI guidance that may be presented after the user selects Tool 2 (as indicated by the shading in Figure 3D). Here again, the HCI guidance is an overlaid visual annotation that points to the "Format" menu and informs the user that "When using this tool, you typically used 8pt brush strokes with anti-aliasing. Here are their settings." If the user selects the "Format" menu, one or more graphical elements will become available for the user to select the brush stroke size.
[0070] Figure 4 is a flowchart illustrating an exemplary method 400 for practicing a selected aspect of the disclosure according to embodiments disclosed herein. For convenience, the operations in the flowchart are described with reference to a system that performs the operations. This system may include various components of various computer systems, such as one or more components of the semantic task guidance system 102. Furthermore, the operations of method 400 are shown in a particular order, but this is not intended to be limiting. One or more operations can be rearranged, omitted, or added.
[0071] In block 402, the system may, for example, use domain module 104 to identify a first domain of a first computer application that can be operated using the first HCI. For example, in Figure 3A, the domain module 104 can identify the virtual CAD domain as active because the user has launched a virtual CAD. As previously mentioned, the domain module 104 can use any number of signals in addition to, or instead of, computer applications to identify the current domain. For example, if the user is operating the HMD to explore a virtual reality world sometimes referred to as the "metaverse," the domain can be identified using the region of the metaverse that the user is currently exploring.
[0072] Based on the domain identified in block 402, in block 404, the system may, for example, use the semantic coordination agent 122 or the ML module 108 to select a first domain model to convert between the action space of the first computer application and another space. For example, in Figures 3A to 3D, the ML module 108 selected a virtual CAD domain model.
[0073] Based on the selected first domain model, in block 406, the system may process domain-independent action embeddings, for example by ML module 108, to generate one or more probability distributions over actions in the action space of the first computer program. The action embeddings may represent a series of actions previously performed using the second HCI of the second computer application to perform a semantic task. This is shown in Figures 3A to 3D, where the domain-independent action embeddings were previously generated based on user interaction with the HCI of the previous software, FakeCAD, and these domain-independent action embeddings were processed using the domain model of the virtual CAD to generate a probability distribution(s) over the action space of the virtual CAD.
[0074] Based on one or more probability distributions generated in block 406, in block 408, the system may identify a second set of actions that can be performed using a first computer application, for example, by the ML module 108 or the semantic adjustment agent 122. In block 410, the system may have the output presented on one or more output devices, for example, by the interface module 106. The output may include guidance for navigating the first HCI to perform a semantic task using the first computer application. The guidance can be generated, for example, by the interface module 106, based on the identified second set of actions that can be performed using the first computer application. The HCI guidance may be provided in various forms. Visually, it may be presented as overlaid annotations (e.g., as shown in Figures 3A to 3D), animations, videos, written natural language output (e.g., text balloons), 3D renderings (e.g., in a virtual reality or metaverse setting), etc. Auditory, the HCI guidance may be presented as natural language output, various sounds, or sound effects, etc.
[0075] The examples described herein primarily focus on providing HCI guidance across semantically similar computer applications, such as between FakeCAD and VirtualCAD in a CAD computer application, or between different spreadsheet applications. However, this is not intended to be limiting. As long as a particular semantic task is relatively independent of a particular domain, that task can be used to generate actions in any number of domains that are otherwise dissimilar. For example, setting a particular computer application to "dark mode" may be relatively universal and can therefore be leveraged to provide HCI guidance across various domains, such as other computer applications or even operating systems.
[0076] As another example, an automated assistant (sometimes referred to as a “virtual” assistant or “virtual agent”) can interface with any number of third-party agents, enabling the automated assistant to function as a liaison for performing tasks such as ordering or booking goods or services, or booking rideshares. Different companies offering similar services (e.g., rideshares) may require users to interact with their respective third-party agents using natural language dialogue. However, the natural language dialogue available for interacting with a first rideshare agent may differ from the dialogue used for interacting with a second rideshare agent. Nevertheless, the final parameters or “slot values” to be satisfied to complete a rideshare request may be semantically similar, even if they are differently named and requested at different points in the conversation. Thus, the techniques described herein can be used to help a user familiar with a first rideshare agent interact more efficiently with a second rideshare agent by providing the user with HCI guidance, for example, as auditory or visual natural language, graphical elements on a display, etc. For example, domain-independent action embeddings can be processed to generate scripts for automated agents that function as liaisons. This script can request necessary parameters or slot values from the user in a domain-independent manner. The automated agent can then use these requested values to engage with any rideshare agent, without requiring the user to learn the nuances of each agent.
[0077] Figure 5 is a block diagram of an exemplary computing device 510 that can be optionally used to implement one or more embodiments of the technology described herein. In some embodiments, one or more of the client computing devices 120-1 to 120-P, the semantic task guidance system 102, and / or other components may include one or more components of the exemplary computing device 510.
[0078] The computing device 510 typically includes at least one processor 514 that communicates with several peripheral devices via a bus subsystem 512. These peripheral devices may include, for example, a storage subsystem 524 including a memory subsystem 525 and a file storage subsystem 526, a user interface output device 520, a user interface input device 522, and a network interface subsystem 516. The input and output devices enable user interaction with the computing device 510. The network interface subsystem 516 provides an interface to an external network and is coupled to a corresponding interface device in another computing device.
[0079] The user interface input device 522 may include pointing devices such as keyboards, mice, trackballs, touchpads, or graphics tablets, scanners, touchscreens integrated into displays, voice recognition systems, microphones, and / or other types of input devices. Generally, the use of the term “input device” is intended to include all possible types of devices and methods for inputting information into the computing device 510 or a communication network.
[0080] The user interface output device 520 may include a non-visual display such as a display subsystem, printer, fax machine, or audio output device. The display subsystem may include a flat panel device such as a cathode ray tube (CRT), liquid crystal display (LCD), projection device, or any other mechanism for creating a visible image. The display subsystem may also provide a non-visual display via an audio output device, etc. In general, the use of the term “output device” is intended to include all possible types of devices and methods for outputting information from the computing device 510 to a user or another machine or computing device.
[0081] The storage subsystem 524 stores programming and data structures that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 524 may include logic for carrying out a selected embodiment of method 400 shown in Figure 4.
[0082] These software modules are generally executed by processor 514 alone or in combination with other processors. The memory 525 used within the storage subsystem 524 may include several memories, including main random access memory (RAM) 530 for storing instructions and data during program execution, and read-only memory (ROM) 532 for storing fixed instructions. The file storage subsystem 526 can provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules performing the functions of a particular embodiment may be stored by the file storage subsystem 526 within the storage subsystem 524 or on other machines accessible by processor(s) 514.
[0083] The bus subsystem 512 provides a mechanism for various components and subsystems of the computing device 510 to communicate with each other as intended. Although the bus subsystem 512 is schematically shown as a single bus, alternative embodiments of the bus subsystem may use multiple buses.
[0084] The computing device 510 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing systems or computing devices. Due to the constantly changing nature of computers and networks, the description of the computing device 510 shown in Figure 5 is intended only as a specific example to illustrate several embodiments. Many other configurations of the computing device 510 are possible, having more or fewer components than the computing device shown in Figure 5.
[0085] While several embodiments have been described and illustrated herein, various other means and / or structures can be utilized to perform the function and / or to obtain one or more of the results and / or benefits described herein, and each of such variations and / or modifications is considered to be within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials and configurations described herein are intended to be illustrative, and actual parameters, dimensions, materials and / or configurations depend on one or more specific uses in which this teaching is used. Those skilled in the art will be able to recognize or investigate many equivalents to the specific embodiments described herein simply by using routine experiments. Thus, it will be understood that the embodiments described herein are presented only as examples, and that within the scope of the appended claims and their equivalents, embodiments may be practiced in ways other than those specifically described and claimed. Embodiments of this disclosure cover each individual feature, system, article, material, kit and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of this disclosure, provided that such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.
Claims
1. A method that is carried out using one or more processors, Identifying a domain of a first computer application that can be operated using a first human-computer interface (HCI), Based on the identified domain, select a domain model to translate between the action space of the first computer application and another space, Based on the selected domain model, the process of action embedding generates one or more probability distributions over actions in the action space of the first computer application, wherein the action embedding generates one or more probability distributions representing a plurality of actions previously performed using the second HCI of the second computer application to perform a semantic task. Identifying a second set of actions that can be performed using the first computer application based on one or more probability distributions, A method comprising causing an output to be presented on one or more output devices, wherein the output includes guidance for navigating the first HCI to perform the semantic task using the first computer application, and the guidance causes an output to be presented on one or more output devices based on the identified second plurality of actions that can be performed using the first computer application.
2. The method according to claim 1, wherein the domain model is trained to translate between the action space and the domain-independent action embedding space of the first computer application.
3. The method according to claim 1 or 2, wherein the domain model is trained to directly translate between the action space of the first computer application and the action space of the second computer application.
4. The method according to any one of claims 1 to 3, wherein the first HCI includes a graphical user interface.
5. The method according to claim 4, wherein the guidance for navigating the first HCI includes one or more visual annotations that are overlaid on the GUI.
6. The method according to claim 5, wherein one or more of the visual annotations are rendered to draw attention to one or more graphical elements of the GUI.
7. The method according to any one of claims 1 to 6, wherein the guidance for navigating the first HCI includes one or more natural language outputs.
8. Obtaining user input that conveys the aforementioned semantic task, Identifying the action embedding based on the aforementioned semantic task, The method according to any one of claims 1 to 7, further comprising:
9. The user input includes natural language input, and the method is Performing natural language processing (NLP) on the aforementioned natural language input to generate a first task embedding that represents the aforementioned semantic task, The method further includes determining a similarity measure between the first task embedding and the action embedding, The method according to claim 8, wherein the action embedding is processed based on the similarity scale.
10. A system comprising one or more processors and memory for storing instructions, wherein the instructions, in response to the execution of the instructions, are provided to the one or more processors. Identifying a domain of a first computer application that can be operated using a first human-computer interface (HCI), Based on the identified domain, select a domain model to translate between the action space of the first computer application and another space, Based on the selected domain model, the process of action embedding generates one or more probability distributions over actions in the action space of the first computer application, wherein the action embedding generates one or more probability distributions representing a plurality of actions previously performed using the second HCI of the second computer application to perform a semantic task. Identifying a second set of actions that can be performed using the first computer application based on one or more probability distributions, To cause an output to be presented on one or more output devices, wherein the output includes guidance for navigating the first HCI to perform the semantic task using the first computer application, and the guidance causes an output to be presented on one or more output devices based on the identified second plurality of actions that can be performed using the first computer application. A system that executes an action.
11. The system according to claim 10, wherein the domain model is trained to translate between the action space and the domain-independent action embedding space of the first computer application.
12. The system according to claim 10 or 11, wherein the domain model is trained to directly translate between the action space of the first computer application and the action space of the second computer application.
13. The system according to any one of claims 10 to 12, wherein the first HCI includes a graphical user interface.
14. The system according to claim 13, wherein the guidance for navigating the first HCI includes one or more visual annotations that are overlaid on the GUI.
15. The system according to claim 14, wherein one or more of the visual annotations are rendered to draw attention to one or more graphical elements of the GUI.
16. The system according to any one of claims 10 to 15, wherein the guidance for navigating the first HCI includes one or more natural language outputs.
17. Obtain user input that conveys the aforementioned semantic task, Based on the aforementioned semantic task, identify the action embedding. The system according to any one of claims 10 to 16, further comprising instructions for
18. The user input includes natural language input, and the system, Natural language processing (NLP) is performed on the aforementioned natural language input to generate a first task embedding that represents the semantic task. The instruction further includes instructions for determining a similarity measure between the first task embedding and the action embedding, The system according to claim 17, wherein the action embedding is processed based on the similarity scale.
19. A non-temporary computer-readable medium containing instructions, wherein the instructions, in response to the execution of the instructions by the processor, Identifying a domain of a first computer application that can be operated using a first human-computer interface (HCI), Based on the identified domain, select a domain model to translate between the action space of the first computer application and another space, Based on the selected domain model, the process of action embedding generates one or more probability distributions over actions in the action space of the first computer application, wherein the action embedding generates one or more probability distributions representing a plurality of actions previously performed using the second HCI of the second computer application to perform a semantic task. Identifying a second set of actions that can be performed using the first computer application based on one or more probability distributions, To cause an output to be presented on one or more output devices, wherein the output includes guidance for navigating the first HCI to perform the semantic task using the first computer application, and the guidance causes an output to be presented on one or more output devices based on the identified second plurality of actions that can be performed using the first computer application. A non-temporary computer-readable medium that enables execution.
20. The computer-readable medium according to claim 19, wherein the domain model is trained to translate between the action space and the domain-independent action embedding space of the first computer application.