Generating cross-domain guidance for navigating HCI
The method uses action embeddings and domain models to generate guidance for navigating human-computer interfaces across different applications, addressing the challenge of task familiarity and improving user efficiency.
Patent Information
- Application Number
- JP2024571301
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-15
- Filing Date
- 2023-06-12
- Publication Date
- 2025-07-31
- Estimated Expiration
- 2043-06-12
AI Technical Summary
Individuals often struggle to perform semantically similar computing tasks across different computer applications due to unfamiliarity with the application's user interface, leading to inefficiencies and difficulties in navigating and completing tasks.
A method utilizing action embeddings in a generalized action embedding space, converted through domain models, to generate guidance for navigating human-computer interfaces across different applications, including visual and audible annotations, and natural language outputs.
Enables users to perform semantically equivalent or similar tasks across different computer applications by providing tailored guidance, enhancing user efficiency and adaptability.
Smart Images

Figure 2025524734000001_ABST
Abstract
Description
[Technical Field]
[0001] Individuals often operate computing devices to perform semantically similar tasks in different contexts. For example, an individual may engage in a series of actions using a first computer application to perform a given semantic task, such as setting preferences for various applications, retrieving / viewing particular data made accessible by the first computer application, or performing a series of operations within a particular domain (e.g., 3D modeling, graphics editing, word processing). The same individual may later engage in a semantically similar but syntactically different series of actions to perform a semantically equivalent task (e.g., the same semantic task) in a different context, such as while using a second computer application. However, the individual may be less familiar with the second computer application and, as a result, may be unable to perform the semantic task. Summary of the Invention
[0002] Embodiments are described for automatically generating and providing guidance for navigating a human-computer interface (HCI) to perform semantically equivalent and / or semantically similar computing tasks across different computer applications. More specifically, but not by way of limitation, embodiments are described that enable an individual (often referred to as a “user”) to utilize actions taken in one context (e.g., when performing semantic tasks) to generate guidance for performing semantically identical or similar tasks in other contexts. In various embodiments, the captured actions can be abstracted as “action embeddings” in a generalized “action embedding space”. This domain-independent action embedding can represent “semantic tasks” that can be transformed into the action spaces of any number of domains using their respective domain models during abstraction. Put another way, a “semantic task” is a domain-independent higher-order task that finds an expression within a particular domain as a set of / multiple domain-specific actions.
[0003] In some embodiments, the method may be implemented using one or more processors and may include identifying a first domain of a first computer application that is operable using a first human-computer interface (HCI), selecting a domain model that converts between an action space of the first computer application and another space based on the identified domain, processing action embeddings based on the selected domain model to generate one or more probability distributions over actions within the action space of the first computer program, where the action embeddings represent multiple actions previously performed using a second HCI of a second computer application to perform semantic tasks, generating the one or more probability distributions, identifying a second plurality of actions that are executable using the first computer application based on the one or more probability distributions, and presenting an output on one or more output devices. In various embodiments, the output may include guidance for navigating the first HCI to perform a semantic task using the first computer application, the guidance being based on the identified second plurality of actions executable using the first computer application.
[0004] In various embodiments, the domain model may be trained to convert between the action space of the first computer program and a domain-independent action embedding space. In various embodiments, the domain model may be trained to directly convert between the action space of the first computer program and the action space of a second computer program.
[0005] In various embodiments, the first HCI may take the form of a graphical user interface. In various embodiments, guidance for navigating the first HCI may include one or more visual annotations overlaying the GUI. In various embodiments, one or more of the visual annotations may be rendered to draw attention to one or more graphical elements of the GUI.
[0006] In various embodiments, the guidance for navigating the first HCI may include one or more natural language outputs. In various embodiments, the method may further include obtaining user input communicating a semantic task and identifying action embeddings based on the semantic task. In various embodiments, the user input may be natural language input, and the method may further include performing natural language processing (NLP) on the natural language input to generate a first task embedding representing the semantic task and determining a similarity measure between the first task embedding and the action embedding, where the action embedding is processed based on the similarity measure.
[0007] Additionally, some embodiments include one or more processors of one or more computing devices, the one or more processors operable to execute instructions stored in associated memory, the instructions configured to cause any of the aforementioned methods to be performed. Some embodiments include at least one non-transitory computer-readable storage medium storing computer instructions executable by the one or more processors to perform any of the aforementioned methods.
[0008] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are contemplated as being part of the subject matter disclosed herein, for example, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a schematic diagram of an exemplary environment in which embodiments disclosed herein may be implemented. [Figure 2] 1 illustrates generally one example of how data may be exchanged and / or processed to extend a task performed in one domain to additional domains, according to various embodiments. [Figure 3A] 1 illustrates an example of how the techniques described herein can be used to provide guidance for interacting with a computer application, according to various embodiments. [Figure 3B] 1 illustrates an example of how the techniques described herein can be used to provide guidance for interacting with a computer application, according to various embodiments. [Figure 3C] 1 illustrates an example of how the techniques described herein can be used to provide guidance for interacting with a computer application, according to various embodiments. [Figure 3D] 1 illustrates an example of how the techniques described herein can be used to provide guidance for interacting with a computer application, according to various embodiments. [Figure 4] 1 is a flowchart illustrating an exemplary method for practicing selected aspects of the present disclosure, according to implementations disclosed herein. [Figure 5] 1 is an exemplary architecture of a computing device. DETAILED DESCRIPTION OF THE INVENTION
[0010] Described herein are embodiments for automatically generating and providing guidance for navigating a human-computer interface (HCI) to perform semantically equivalent and / or semantically similar computing tasks across different computer applications. More specifically, but not by way of limitation, described herein are embodiments that enable an individual (often referred to as a “user”) to leverage actions taken in one context (e.g., while performing a semantic task) to generate guidance for performing semantically identical or similar tasks in other contexts. In various embodiments, the captured actions can be abstracted as “action embeddings” in a generalized “action embedding space.” This domain-independent action embedding can represent, in abstraction, a “semantic task” that can be transformed into the action spaces of any number of domains using the respective domain models. In other words, a “semantic task” is a domain-independent, higher-level task that finds a representation in a specific domain as a sequence / several domain-specific actions.
[0011] As one non-limiting example, a user may allow a local agent computer program (also referred to herein as an "agent" or "assistant") to monitor the user's interactions with one or more local computer applications. This monitoring may include, for example, capturing the user's interactions with an HCI, such as a graphical user interface (GUI), of a first computer application. The user may interact with the HCI to perform various semantic tasks that may be performed using the first computer application. Each semantic task may include multiple individual or atomic interactions with the HCI of the first computer application.
[0012] For example, if the first computer application is a three-dimensional (3D) design application, the semantic task may include designing a 3D structure, and the atomic interactions may include, for example, navigating to a specific menu, selecting a specific tool from those menus, selecting specific settings for those tools, operating those tools on a canvas, and so on. If the first computer application is a spreadsheet application, the semantic task may include, for example, creating a chart based on underlying data. The atomic interactions may include, for example, sorting data, adding columns (e.g., with formulas that utilize existing column values of operands), selecting a range to navigate a menu, selecting specific items from those menus to create a desired chart, and so on.
[0013] Returning to the example, domain-specific actions captured in relation to the semantic task executed using the first computer application can be abstracted into action embeddings using a domain model associated with the domain of the first computer application. This action embedding can then be converted into any number of other domain action spaces, such as the action space of a second computer program. For example, a probability distribution may be generated over the actions within the action space of the second computer program. A plurality of domain-specific actions executable using the second computer program may be selected from the action space, for example, based on their probabilities (e.g., generated using the softmax layer of the domain model). Then, the selected domain-specific actions of the second computer program can be used to generate guidance for navigating the HCI provided by the second computer program.
[0014] Guidance for navigating HCI can be generated and / or presented in various ways. In some embodiments where the HCI is a GUI, for example, visual annotations that overlay all or part of the GUI can be presented. In some embodiments, these visual annotations can draw attention to graphical elements that can be manipulated (e.g., using a pointer device or finger if the display is a touch screen) to perform a plurality of domain-specific actions identified from the action space of a second computer program. The visual annotations can include, for example, arrows, animations, natural language text, shapes, etc. Additionally or alternatively, audible guidance may be presented to aurally guide the user to specific graphical elements corresponding to domain-specific actions identified from the action space of a second computer program. The audible guidance can include, for example, natural language output, sounds with visual annotations (e.g., animations), etc.
[0015] Accordingly, with the techniques described herein, a user can permit an agent configured in a selected aspect of the present disclosure to monitor these types of interactions with various HCIs within various domains (e.g., by “opting in”). The knowledge obtained by the agent can be captured (e.g., in various domain machine learning models) and utilized to generate guidance for performing semantically similar actions in other domains. In some embodiments, the extent to which other users follow or deviate from such guidance may then be used to train the domain model, such that, as a result, they can select “better” actions in the future. In some cases, if a sufficient number of users follow the same (or substantially similar) guidance in a given computer application, a tool can be created that can be automatically invoked using that guidance, eliminating the need for subsequent users to repeat the same atomic actions provided in the guidance.
[0016] In some implementations, a user can provide natural language input to describe a series of actions to be performed using the HCI of the first domain, e.g., while performing them, or immediately before or after. For example, while operating a first spreadsheet application, a user can state, "I am creating a bar graph showing net loss for the past 90 days." A first task / policy embedding generated from natural language processing (NLP) of this input can be associated (e.g., mapped, combined) with a first action embedding generated from the captured series of actions using a first domain model associated with the first spreadsheet application. As previously described, the first domain model can transform between the action space of the first spreadsheet and, for example, a general action embedding space and / or one or more other domain-specific action embedding spaces.
[0017] Later, when operating a second spreadsheet application having similar functionality to the first spreadsheet application, the user can provide semantically similar natural language input to learn how to perform a semantically equivalent (or at least semantically similar) task using the second spreadsheet application. For example, the user may utter, "How do I create a bar graph showing net loss for the last 120 days?" A second task / policy embedding generated from this subsequent natural language input can be matched to the first task / policy embedding, and thus to the first action embedding. The first action embedding can then be processed using a second domain model that translates between the generic action embedding space and the action space of the second spreadsheet application to identify action(s) that are performable in the second spreadsheet application to perform the semantic task. These identified action(s) can be used to generate guidance for performing the semantic task in the second spreadsheet application.
[0018] For example, visual annotations and / or audible guidance may be provided to guide a user through various menus, sheets, cells, etc. of the second spreadsheet application to perform the creation of a bar chart showing net losses for the past 120 days. In particular, the fact that the first chart showed a 90-day loss while the second chart shows a 120-day loss may be handled by the agent, for example, by storing the number of days as a parameter associated with an action in the action space of the second spreadsheet application. In addition to capturing the semantics of the HCI itself, the domain model may also be trained to identify where semantically equivalent data exists.
[0019] For example, when operating a first spreadsheet application to edit a first spreadsheet file, the data needed to determine net loss may be on a particular tab that also contains various other data. In contrast, a second spreadsheet editable using a second spreadsheet application may contain semantically similar data, i.e., the data needed to determine net loss, on a different tab, with or without other data. However, if sufficient training examples are provided over time, a domain model used by an agent configured using selected aspects of the present disclosure may be able to locate the appropriate data for determining net loss. For example, different columns in different spreadsheets containing data related to net loss may contain semantically similar column headings. Additionally or alternatively, the actual data itself may share semantic characteristics, i.e., be formatted similarly, have generally similar values (e.g., within the same order of magnitude, millions versus hundreds of millions, etc.), or exhibit similar temporal patterns (e.g., higher sales during certain seasons).
[0020] In addition to, or instead of, guidance, in some embodiments, the techniques described herein can be used to configure the HCI itself to adapt to the behavior or capabilities of a particular user. For example, the visual settings of a GUI can be configured via a variety of different actions to make the GUI easier to operate for visually impaired users. This can include, for example, increasing the font size, increasing the contrast, decreasing the number of presented menu items (e.g., based on usage frequency across the user population), increasing the size of actionable graphical elements such as sliders, buttons, etc., activating user accessibility settings as voice prompts, and the like. These actions can be captured in a given computer application and abstracted into action embeddings, along with task / policy embeddings created from natural language inputs such as "impose visual impairment settings".
[0021] Later, in a different context (e.g., when operating a different computer application), the user can provide a natural language input such as "I am visually impaired, please make this interface easier to operate". The previously generated action embeddings can be processed using a domain model associated with the new context to automatically perform at least some of the aforementioned adjustments and / or show the user how to perform them. If any of the adjustments are not available or applicable, the user may be so notified and / or provided with different recommendations that meet a similar need.
[0022] The techniques described herein are not limited to generating guidance for performing semantic tasks across similar domains (e.g., from one spreadsheet application to another). In various implementations, guidance for performing semantic tasks may also be generated across semantically distinct domains / contexts. For example, semantically similar but domain-independent application parameters of various computer application(s) may be named, organized, and / or accessed differently (e.g., different submenus, command line inputs, etc.). Such application parameters may include, for example, visual parameters that may be set for various modes such as "dark mode," application permissions (e.g., access to location, camera, files, other applications, etc.), or other application settings (e.g., Celsius vs. Fahrenheit settings, metric vs. Imperial unit settings, preferred fonts, preferred sort orders, etc.). Many of these various application parameters may not be specific to a particular computer application or domain. Indeed, some application parameters, such as "skins" applied to a GUI, may also be applicable to an operating system (OS).
[0023] The spreadsheet example provided above involved two different spreadsheet applications. However, this is not intended to be limiting. The techniques described herein may be implemented to generate guidance for performing semantic tasks across multiple different use cases within a single domain. Assume a user manipulates a first "docket" spreadsheet to organize docket and schedule data in a particular way and then generates a docket report in a particular format. The actions the user performs to create this report may be captured and abstracted into action embeddings, as described above, using, for example, the domain model of whatever spreadsheet application the user is manipulating.
[0024] At a later time, the user may receive a second “docket” spreadsheet, created, for example, by a different docketing system or for a different entity. This second docket spreadsheet contains semantically similar data as the first docket spreadsheet, but may be organized differently and / or have a different schema. Columns may have different names and / or be in a different order. Data may be expressed using a different syntax (e.g., “MM / DD / YY” vs. “DD / MM / YYYY”). Nevertheless, the previously created action embeddings may be processed in conjunction with the second docket spreadsheet (e.g., as additional contextual input data) to generate guidance for performing the same semantic task using the second docket spreadsheet. For example, the same domain model, or a separate domain model (e.g., trained in reverse order), that can also process contextual input data, may be applied to identify actions that are executable to perform the semantic task using the second docket spreadsheet.
[0025] In various implementations, the domain model may be continuously trained based on how users interact with guidance generated using the techniques described herein. This, in turn, may affect how or whether various guidance is provided. For a particular domain-independent semantic action, such as "set dark mode," assume that the particular suggested action is rarely or never performed in a particular domain, e.g., using a particular computer application in that domain. Perhaps the native settings of that computer application already address the underlying problem that necessitated the suggested action in other domains, or make such problem practically irrelevant.
[0026] In such scenarios, the domain model associated with a particular domain may be further trained so that when the same (or similar) action embeddings are processed, the probability distribution resulting across the action space of the domain assigns a lower probability to the proposed action. In contrast, in other domains where the underlying problem still exists, a higher probability may be given to the proposed action. The assigned probability may indicate how the proposed action is presented to the user (e.g., how prominent, whether animated or a small visual annotation, whether auditory or visual, etc.), when the proposed action is presented to the user (e.g., relative to other proposals), whether the proposed action is presented to the user at all, or whether the action should be automatically performed without user guidance.
[0027] The continuous training of the domain model is not limited to monitoring user feedback / response to the HCI guidance provided in the new domain. In some embodiments, the domain model can be trained without the user leaving the original domain in which the semantic task is being performed. For example, when a user performs a series of actions using a computer application to complete a given semantic task, the domain model associated with the domain of the computer application can be used to process the actions (or data indicating them, such as embeddings) to generate a domain-independent embedding that semantically represents the given task. The domain-independent embedding can then be processed using a machine learning model (e.g., a sequence decoder) trained to generate a natural language output intended to describe the given semantic task performed by the user. For example, in response to a change in a particular visual setting in an application or operating system, a natural language output such as "It seems like the graphical interface has been changed to 'dark mode'" may be presented to the user.
[0028] This natural language output may be presented to the user audibly or visually, along with a request for user feedback (“Is that what you did?” or “Did I describe your action accurately?”). The user's positive feedback (“Yes, that's right”) or negative feedback (“No, that's not what I did”) may be used to train the domain model, for example, using techniques such as backpropagation and gradient descent. The user may be able to adjust or influence how often (or even whether) the request for feedback is presented to the user. In some cases, the user may be prompted for such feedback and / or offered incentives to provide feedback. These incentives may come in various forms, such as monetary rewards, credits associated with the computer application (e.g., special items in a game), etc. Additionally or alternatively, the agent itself may self-regulate the frequency with which it prompts the user for such feedback based on signals such as the user's response (e.g., abandonment or cooperation), a measure of accuracy associated with the domain model in question (more accurate models may be trained less frequently than less accurate models), etc.
[0029] As used herein, "domain" can refer to the target area in which a computing component is intended to operate, e.g., the scope of knowledge, influence, and / or activities around which the logic of the computing component revolves. In some embodiments, the domain can be identified by heuristically matching keywords in the user-provided input with domain keywords. In other embodiments, the user-provided input can be processed using NLP techniques such as, for example, word2vec, a Bidirectional Encoder Representations from Transformers (BERT) transformer, various types of recurrent neural networks ("recurrent neural network", "RNN", e.g., long short-term memory, i.e., "LSTM", gated recurrent unit, i.e., "GRU") to generate a meaning embedding representing the user's input.
[0030] In various embodiments, one or more domain models may have been previously generated for each domain. For example, one or more machine learning models such as RNN (e.g., LSTM, GRU), BERT transformers, various types of neural networks, reinforcement learning policies, etc. can be trained based on a corpus of documentation associated with the domain. As a result of this training, one or more of the domain model(s) can be at least bootstrapped, and as a result, process what is referred to herein as an "action embedding" to select a plurality of candidate computing actions from the action space associated with the target domain that can be used to provide the guidance described herein.
[0031] FIG. 1 schematically illustrates an exemplary environment in which selected aspects of the present disclosure may be implemented, according to various embodiments. Any computing device illustrated in FIG. 1 or elsewhere in the figure may include logic such as one or more microprocessors (e.g., central processing units or “CPUs,” graphical processing units or “GPUs,” tensor processing units (“TPUs”)) that execute computer-readable instructions stored in memory, or other types of logic such as application-specific integrated circuits (“ASICs”), field-programmable gate arrays (“FPGAs”), etc. Some of the systems illustrated in FIG. 1 , such as the semantic task guidance system 102, may be implemented using one or more server computing devices forming what is sometimes referred to as a “cloud infrastructure,” although this is not required. In other embodiments, aspects of the semantic task guidance system 102 may be implemented on the client device 120, e.g., for purposes of protecting privacy, reducing latency, etc.
[0032] The semantic task guidance system 102 may include several different components configured in accordance with selected aspects of the present disclosure, such as a domain module 104, an interface module 106, a machine learning ("ML" in FIG. 1 ) module 108, and / or a task identification ("ID" in FIG. 1 ) module 110. The semantic task guidance system 102 may also include any number of databases for storing machine learning model weights and / or other data used to perform selected aspects of the present disclosure. In FIG. 1 , for example, the semantic task guidance system 102 includes a database 111 that stores a global domain model and another database 112 that stores data indicative of global action embeddings.
[0033] The semantic task guidance system 102 may be operatively coupled to any number of client computing devices operated by any number of users via one or more computer networks (114). In FIG. 1 , for example, a first user 118-1 operates one or more client devices 120-1. A pth user 118-P operates one or more client device(s) 120-P. As used herein, client device(s) 120 may include, for example, one or more of a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device in a user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker (which may in some cases include a visual sensor and / or a touchscreen display), a smart appliance such as a smart television (or a standard television with a networked dongle having automated assistant functionality), and / or a user wearable device including a computing device (e.g., a user's watch with a computing device, a user's glasses with a computing device, a virtual or augmented reality computing device). Additional and / or alternative client computing devices may be provided.
[0034] Domain module 104 can be configured to determine various different information related to a given domain related to a given user 118 at a given time, such as the domain that user 118 is currently operating on, the domain(s) that the user has previously operated on, the domain(s) that the user wants to expand semantic tasks for, or the domain(s) for which the user wants to receive guidance on how to perform semantic tasks. For this purpose, the domain module 104 can collect context information, for example, on foreground applications and / or background applications running on client device(s) 120 that user 118 operates on, web pages that user 118 has currently / recently visited, domain(s) that user 118 has access to and / or frequently accesses.
[0035] Using this collected context information, in some embodiments, the domain module 104 may be configured to identify one or more domains related to the current user. For example, a request to record or observe the tasks performed by user 118 using a specific computer application and / or a specific input form is processed by the domain module 104 to identify the domain in which user 118 performs the recorded tasks, and the domain may be the domain of a specific computer application or input form. If user 118 later requests guidance on performing the same task in a different target domain, for example, using a different computer application or a different input form, the domain module 104 can identify the target domain. The user may not need to request guidance in different target domains. In some embodiments, the techniques described herein can be implemented by simply operating different computing applications or input forms to provide the user with unsolicited guidance on how to perform semantic tasks similar to those previously performed by the user in another domain.
[0036] In some embodiments, the domain module 104 may also be configured to extract domain knowledge from various different sources related to the specified domain. In some such embodiments, this extracted domain knowledge (and / or the embedding(s) generated therefrom) may be provided to downstream component(s), for example, in addition to the aforementioned natural language input or context information. This additional domain knowledge may enable the downstream component(s), particularly the machine learning model, to make more likely satisfactory predictions (e.g., generate guidance for performing semantic tasks across different domains).
[0037] In some implementations, the domain module 104 can apply the collected context information (e.g., the current state) across one or more “domain selection” machine learning model(s) 105 that differ from the domain models described herein. These domain selection machine learning model(s) 105 can take various forms, such as various types of neural networks, support vector machines, random forests, BERT transformers, etc. In various implementations, the domain selection machine learning model(s) 105 can be trained to select applicable domains based on attributes (or “context signals”) of the current context or state of the user 118 and / or client device 120. For example, if the user 118 is interacting with an input form on a particular website to purchase a product or service, the website's uniform resource locator (URL), or attributes of the underlying webpage(s), such as keywords, tags, document object model (DOM) elements, etc., can be applied as input across the model, either in their native form or as a dimensionality-reduced embedding. Other context signals that may be considered include, but are not limited to, the user's IP address (e.g., work vs. home vs. mobile IP address), time of day, social media status, calendar, email / text messaging content, etc.
[0038] Interface module 106 may provide one or more graphical user interfaces (GUIs) that various individuals, such as users 118-1 through 118-P, can operate to perform various actions made available by semantic task guidance system 102. In various implementations, users 118 may operate a GUI (e.g., a standalone application or a web page) provided by interface module 106 to opt-in to or utilize various techniques described herein. For example, users 118-1 through 118-P may be required to provide explicit permission before any tasks they perform using client device(s) 120-1 through 120-P are observed and used to generate guidance as described herein.
[0039] Additionally, interface module 106 may be configured to practice selected aspects of the present disclosure to present or cause the presentation of the above-described guidance for performing semantic tasks in different domains. For example, interface module 106 may receive one or more sampled actions from the action space of a particular domain from ML module 108. Interface module 106 may then cause graphical and / or audio data representing these actions to be presented to the user.
[0040] A designer is operating a new computer-aided design (CAD) computer application. Assume that the designer previously operated an old CAD computer application, for example, as part of their employment. The actions of a given task that the designer frequently performed using the old CAD computer application can be processed by the ML module 108 using, for example, a domain model associated with the old CAD computing application to generate domain-independent action embeddings. This action embedding can then be transformed, for example, by the ML module 108, into the domain of the new CAD computer application to generate (e.g., sample) one or more actions that can be performed using the new CAD computer application. These actions can be used by the interface module 106 to generate audio and / or visual guidance that explains to the user how to perform a given task using the new CAD computer application.
[0041] The ML module 108 can access data representing various global domain / machine learning models / policies in the database 111. These trained global domain / machine learning models / policies can take a variety of forms, including, but not limited to, graph-based networks such as graph neural networks (GNNs), graph attention neural networks (GANNs), or graph convolutional neural networks (GCNs), sequence-to-sequence models such as encoder-decoders, various recurrent neural networks (e.g., LSTMs, GRUs, etc.), BERT transformer networks, reinforcement learning policies, and any other type of machine learning model that can be applied to facilitate selected aspects of the present disclosure. The ML module 108 may process various data based on these machine learning models at the request or command of other components, such as the domain module 104 and / or the interface module 106.
[0042] Task ID module 110 may be configured to analyze interactions between individuals and computer application(s) collected by semantic coordination agent 122 (described in more detail below). Based on these observations, task ID module 110 can determine which self-contained semantic tasks an individual performs in one domain are likely to be performed by the same individual or other individuals in other domains. In other words, task ID module 110 may selectively trigger (e.g., by ML module 108) the creation of domain-independent action embeddings, which can then be used by ML module 108 to sample actions in different domains for the purpose of providing semantic task guidance across these domains.
[0043] In some embodiments, the task ID module 110 may selectively trigger the creation of person - based domain - independent action embeddings. If a particular person appears to repeatedly perform the same semantic task in a certain domain, guidance for performing that semantic task in other domains may be specifically provided to that person. Additionally or alternatively, the task ID module 110 may selectively trigger the creation of domain - independent action embeddings applicable across a group of persons. When different persons are observed, for example, some threshold number of times, threshold frequency, etc., of performing the same semantic task across one or more domains, may trigger the task ID module 110 to generate the domain - independent embedding(s). Then, the interface module 106 and / or the ML module 108 can use these domain - independent embedding(s) to provide guidance for performing semantic tasks to any number of different persons.
[0044] In various embodiments, the task ID module 110 and / or the meaning adjustment agent 122 can observe a person's interaction only with the person's permission. For example, when installing (or updating) a computer application on a particular client device 120, the meaning adjustment agent 122 can request the person's permission to observe the person's interaction with the new computer application.
[0045] Each client device 120 may operate at least a portion of the aforementioned meaning adjustment agent 122. The meaning adjustment agent 122 may be a computer application operable by the user 118 to implement selected aspects of the present disclosure to facilitate the extension of meaning tasks across different domains. For example, the meaning adjustment agent 122 may receive from the user 118 a request and / or permission to observe / record a series of actions performed by the user 118 using the client device 120 in order to complete some task. Without such an explicit request or permission, the meaning adjustment agent 122 may not be able to observe the user's interaction.
[0046] In some embodiments, the meaning adjustment agent 122 may take the form of what is often referred to as a "virtual assistant" or "automation assistant" configured to participate in a human-computer natural language dialogue with the user 118. For example, the meaning adjustment agent 122 may be configured to semantically process natural language input(s) provided by the user 118 to identify one or more intents. Based on these intents, the meaning adjustment agent 122 may perform various tasks such as operating smart devices, retrieving information, and performing tasks. In some embodiments, the dialogue between the user 118 and the meaning adjustment agent 122 (or a separate automation assistant accessible by / through the meaning adjustment agent 122) may constitute a series of tasks that can be captured, abstracted into domain-independent embeddings, and then extended to other domains as described herein.
[0047] For example, a human-computer interaction between a user 118 and a semantic coordination agent 122 (or a separate automated assistant, or even between an automated assistant and a third-party application) to order a pizza from a third-party agent at a first restaurant (and thus a first domain) can be captured and used to generate an "order pizza" action embedding. This action embedding can later be extended to ordering pizza from a different restaurant, for example, via an automated assistant or via a separate interface.
[0048] In FIG. 1 , each of client device(s) 120-1 may include a semantic coordination agent 122-1 that provides services to first user 118-1. First user 118-1 and its semantic coordination agent 122-1 may have access to and / or associate with a “profile” that includes various data relevant to implementing selected aspects of the present disclosure on behalf of first user 118-1. For example, semantic coordination agent 122 may access one or more edge databases or data stores associated with first user 118-1, including edge database 124-1 that stores local domain model(s) and action embeddings, and / or another edge database 126-1 that stores recorded actions. Other users 118 may have a similar configuration. Any data stored in edge databases 124-1 and 126-1 may be stored partially or entirely on client device 120-1, for example, to protect the privacy of first user 118-1. For example, the recorded actions 126-1 may include confidential and / or personal user information of the first user 118-1, such as payment information, address, phone number, etc., and may be stored locally in its raw form on the client device 120-1.
[0049] The local domain model(s) stored in the edge database 124-1 may include, for example, a local version of the global model(s) stored in the global domain model(s) database 111. For example, in some embodiments, the global model is propagated to the edge for the purpose of bootstrapping the meaning adjustment agent 122 and can extend tasks to new domains associated with those propagated models. Thereafter, the local model at the edge may or may not be locally trained based on the activities and / or feedback of the user 118. In some such embodiments, the local model (alternatively referred to as "local gradient" within the edge database 124) may be periodically used, for example, as part of a federated learning framework, to train the global model (within the database 111). Since the global model is trained based on the local model, the global model may, in some cases, be propagated back to other edge databases (124), thereby keeping the local model up to date.
[0050] However, the adoption of federated learning is not a requirement in all embodiments. In some embodiments, the meaning adjustment agent 122 can provide scrubbed data to the meaning task guidance system 102, and the ML module 108 can remotely apply the model to the scrubbed data. In some embodiments, the "scrubbed" data can be data from which confidential information and / or personal information has been removed and / or obfuscated. In some embodiments, personal information can be scrubbed at the edge by the meaning adjustment automation agent 122, for example, based on various rules. In other embodiments, the scrubbed data provided by the meaning adjustment agent 122 to the meaning task guidance system 102 may be in the form of dimensionality reduction embeddings generated from raw data at the client device 120.
[0051] As described above, the edge database 126-1 can store actions recorded by the meaning adjustment agent 122-1. The meaning adjustment agent 122-1 can observe and / or record actions in a variety of different ways depending on the access level of the meaning adjustment agent 122-1 to the computer application running on the client device 120-1 and the permissions granted by the user 118-1. For example, most smartphones include an operating system (OS) interface for providing or revoking permissions to various computer applications (e.g., access to location, camera, etc.). In various embodiments, such an OS interface can be operable to provide / revoke access to the meaning adjustment agent 122 and / or to select a particular level of access that the meaning adjustment agent 122 has to a particular computer application.
[0052] Semantic adjustment agent 122-1 may have varying levels of access to the workings of a computer application, depending on the permissions granted by user 118 and cooperation from the software developer providing the computer application. Some computer applications may provide semantic adjustment agent 122 with “veiled” access to the application's API or to scripts written using a programming language (e.g., macros) embedded in the computer application, for example, with user 118's permission. Other computer applications may not provide as much access. In such cases, semantic adjustment agent 122 may record actions in other ways, such as by capturing screenshots, performing optical character recognition (OCR) on those screenshots to identify menu items, and / or monitoring user input (e.g., interrupts captured by the OS) to determine which graphical elements were manipulated by user 118 and in what order. In some implementations, semantic adjustment agent 122 may intercept actions performed using the computer application from data exchanged between the computer application and the underlying OS (e.g., via system calls). In some implementations, the semantic adjustment agent 122 may intercept and / or access data exchanged between or used by the window manager and / or window system.
[0053] Figure 2 schematically shows an example of how data can be processed and / or used by various components across multiple domains. Starting from the upper left, user 118 operates client device 120 to request or provide permission for meaning adjustment agent 122, which operates at least partially on client device 120, to observe user 118's interaction with application A, which is a first computer application. In various embodiments, meaning adjustment agent 122 may not be able to record actions without receiving this permission. In some embodiments, this permission may be granted for each application in much the same way that an application is granted permission to access GPS coordinates, local files, use of the on-board camera, etc. In other embodiments, this permission may be granted only until user 118 states otherwise, for example, by pressing a "stop recording" button similar to recording a macro, or by providing a voice input such as "stop recording" or "end".
[0054] Upon receiving the request / permission, in some embodiments, meaning adjustment agent 122 may acknowledge (ACK) the request / permission, but this is not essential. Thereafter, user 118 can launch application A and perform a series of actions {A1, A2,...} within domain A using client device 120, and these actions can be captured and stored in edge database 126. These actions {A1, A2,...} can take various forms or combinations of forms, such as interactions with one or more graphical elements (s) of one or more GUIs using various types of input, such as command line input, as well as pointer device (e.g., mouse) input, keyboard input, voice input, eye gaze input, and any other type of input capable of interacting with graphical elements of the GUI.
[0055] In some implementations, a group of actions performed together logically, e.g., within a particular time interval, without interruption, etc., may be grouped together as a semantic task, e.g., by task ID module 110. For example, a user may perform actions A1-A6 during one session, stop interacting with app A for a period of time, and then perform actions A7-A15 thereafter. In various implementations, actions A1-A6 may be grouped together as one semantic task, and actions A7-A15 may be grouped together as another semantic task.
[0056] In various implementations, the domain (A) in which these actions are performed may be identified by the domain module 104 using any number of signals, such as, for example, that the user 118 has launched app A, as well as other signals, if available. These other signals may include, for example, user-provided natural language input (NLI), the user's calendar, the user's electronic communications, the user's social media posts, etc.
[0057] In various implementations, the semantic adjustment agent 122 can observe / record the actions {A1, A2, ...} and pass them (or data indicative of them, such as dimensionality-reduced embeddings) to another component, such as the ML module 108 (not shown in FIG. 2 ). The ML module 108 can then process these actions using all or part of the domain model A to generate a domain-independent action embedding A′ (also referred to as an “intermediate representation”). In various implementations, the domain model A may comprise, for example, the encoder portion of a larger encoder-decoder architecture that can be used to process a sequence of tokens (e.g., actions {A1, A2, ...}) to generate an intermediate representation, e.g., the action embedding A′.
[0058] In some embodiments, as indicated by the dashed line, user 118 may optionally provide NLI-1 to describe what the user 118 is doing when performing actions {A1, A2, ...}. This NLI-1 may be captured by the meaning adjustment agent 122, and the meaning adjustment agent 122 may pass the NLI-1 to the ML module 108 for natural language processing to generate the task embedding T'. The task embedding T' can be used to provide additional context for the actions {A1, A2, ...}. This additional context can be used in various ways, such as additional input for the domain model, or in the future, as an anchor that enables semantically similar actions to be requested (by user 118 or someone else). As indicated by the dashed line, in some embodiments, the task embedding T' and the action embedding A' can be associated with each other, for example, in a database, via a shared / joined embedding space, etc.
[0059] In some embodiments, if user 118 does not provide natural language input describing actions {A1, A2, ...}, the meaning adjustment agent 122 can formulate (or cause to be formulated) a predicted description of the action(s), and then request feedback from user 118 regarding the accuracy or quality of the description. In FIG. 2, for example, the additional dashed arrows show how the meaning adjustment agent 122 uses the action embedding A' to generate a natural language output (NLO) by processing (or causing to be processed) the embedding A' using a meaning decoder machine learning mode trained to convert between, for example, a domain-independent action embedding space and a natural language vocabulary. The meaning adjustment agent 122 can then present this natural language output to user 118 as part of the feedback request ("Feedback Request" in FIG. 2). User 118 may provide feedback (e.g., "Yes, that's right", or "No, you're wrong"). Based on that feedback, the meaning adjustment agent 122 can train or cause to be trained the domain model A.
[0060] As indicated by the horizontal dashed line, some time later, user 118 launches App B, which causes the meaning adjustment agent 122 to identify Domain B as the active domain. In various embodiments, the meaning adjustment agent 122 may use Domain Model B to cause, for example, Action Embedding A' to be processed by ML module 108. For example, Action Embedding A' may be processed using an encoder-decoder network that collectively forms Domain B or the decoder portion of an encoder-decoder network associated with Domain B. This decoding may generate, for example, a probability distribution over the action space of Domain B. Based on these probability distributions, various actions {B1, B2,...} selected from the action space of Domain B may be generated and provided to the meaning adjustment agent 122. The meaning adjustment agent 122 may then cooperate with the interface module 106 (not shown in FIG. 2) to generate HCI guidance for user 118.
[0061] In various implementations, components such as semantic adjustment agent 122 and ML module 108 may continuously train the domain model, for example, to improve the quality of the HCI guidance provided, allow the HCI guidance to be more narrowly tailored to a particular context, etc. Assume that when this HCI guidance is presented, user 118 follows some parts of the guidance but not others, or follows it in a different order. For example, while the HCI guidance was to perform actions {B1, B2, B3, B4, B5, B6, B7} in order, in FIG. 2, assume user 118 performs fewer than all of the actions in a different order {B1, B5, B4, B3}. In some implementations, the difference between the recommended actions {B1, B2, B3, B4, B5, B6, B7} and the actions ultimately taken by the user 118 {B1, B5, B4, B3} can be used as an error, which can then be used, for example, by the ML module 108, to train a domain model B using techniques such as gradient descent, backpropagation, etc.
[0062] 3A-3D illustrate an example of how a component such as semantic adjustment agent 122 may implement, or be made to implement, selected aspects of the present disclosure to provide guidance for performing a semantic task in a new domain. In this example, assume that a user (not shown) is operating HCI 360 in the form of a GUI rendered by or in place of a "virtual CAD" computer application. It should further be assumed that the user previously worked with a different CAD computer application called "FakeCAD" (e.g., as part of the user's employment) to repeatedly perform any number of tasks. Finally, it should be assumed that the user recently transitioned from using FakeCAD to using virtual CAD (e.g., as a result of starting a new job).
[0063] In FIG. 3A, the HCI is in the home state where the "Home" menu item is active. As a result, context-specific menu items such as "New", "Open", "Save", and "Print" are active. These are merely examples of what may be presented and are not intended to be limiting. On the left side, there may be several design tools (1 - 8) including tools commonly seen in a normal CAD program for drawing lines, ellipses, circles, a shape filling tool, different brushes, etc.
[0064] In this example, when operating the previous CAD software, FakeCAD, to perform various semantic tasks, the domain-specific actions previously performed by the user are being processed to generate domain-independent action embeddings. Specifically, a domain model trained to convert to / from the action space associated with FakeCAD is used to process these domain-specific actions into domain-independent action embeddings. Then, for example, one or more of these domain-independent action embeddings are processed using a domain model configured to convert to / from the action space associated with the new software, virtual CAD, by the ML module 108.
[0065] The output of the processing using the domain model of virtual CAD may include a probability distribution(s) over actions within the action space of virtual CAD. Based on these probability distributions, the ML module 108 or the semantic adjustment agent 122 can select one or more actions within the action space of virtual CAD. These selected action(s) can then be used, for example, by the interface module 106 to generate HCI guidance that helps navigate the HCI 360 to perform the tasks previously performed by the user using FakeCAD.
[0066] For example, in Figure 3A, HCI guidance is presented in the form of an overlaid annotation that points to the "View" menu and explains, "Here are the items you've been consistently adjusting." If the user selects the "View" menu as suggested, HCI 360 may transition to the state shown in Figure 3B.
[0067] In Figure 3B, the "View" menu is currently active, revealing several context-specific menu items. "Zoom," "Ruler," "Mode," "Wireframe," and "Connections" are merely illustrative examples of what may be presented and are not intended to be limiting. Here, new HCI guidance is provided, prompting the user to interact with the "Ruler" and "Wireframe" menu items. This may be because, for example, the user had the habit of frequently adjusting similar parameters when using FakeCAD in the past.
[0068] In FIG. 3C, the "Home" menu is again active. The user is now presented with additional HCI guidance related to "Tool 2." Specifically, the user is informed via an overlaid annotation that "Here is the tool that corresponds to tool x, which you frequently used in FakeCAD." For example, because the HCI guidance previously presented in FIGS. 3A-3B corresponded to an action that the user typically performed previously (e.g., setting general parameters), the user may be presented with this HCI guidance at various times, such as after the HCI guidance presented in FIGS. 3A-3B. In contrast, the HCI guidance presented in FIG. 3C may correspond to an action that the user traditionally performed later, for example, after the user set general parameters to their preferences and was ready to work.
[0069] Figure 3D shows the next iteration of the HCI guidance that may be presented, for example, after the user has selected tool 2 (as indicated by the shading in Figure 3D). Again, the HCI guidance refers to the "Format" menu and is an overlaid visual annotation that notifies the user that "when using this tool, you were typically using an 8pt brush stroke with anti-aliasing. Here are those settings." If the user selects the "Format" menu, one or more graphical elements will become available for the user to select the brush stroke size.
[0070] Figure 4 is a flowchart showing an exemplary method 400 for practicing selected aspects of the present disclosure, according to embodiments disclosed herein. For convenience, the operations of the flowchart are described with reference to the system performing the operations. This system can include various components of various computer systems, such as one or more components of the semantic task guidance system 102. Further, the operations of method 400 are shown in a particular order, but this is not intended to be limiting. One or more operations can be reordered, omitted, or added.
[0071] In block 402, the system may identify, for example, by the domain module 104, the first domain of the first computer application that is operable using the first HCI. For example, in Figure 3A, when the user launches virtual CAD, the domain module 104 can identify the virtual CAD domain as active. As described above, the domain module 104 can use any number of signals in addition to or instead of the computer application to identify the current domain. For example, if the user is operating an HMD to explore a virtual reality world, sometimes referred to as the "metaverse," the domain can be identified using the area of the metaverse that the user is currently exploring.
[0072] Based on the domain identified in block 402, in block 404, the system may select, e.g., by the meaning adjustment agent 122 or the ML module 108, a first domain model that converts between the action space of a first computer application and another space. For example, in FIGS. 3A-3D, the ML module 108 selected a virtual CAD domain model.
[0073] Based on the selected first domain model, in block 406, the system may, e.g., by the ML module 108, process domain-independent action embeddings to generate one or more probability distributions over actions within the action space of a first computer program. The action embeddings may represent multiple actions previously performed using a second HCI of a second computer application to perform a meaning task. This is shown in FIGS. 3A-3D, where the domain-independent action embeddings were previously generated based on the user's interaction with the HCI of the previous software, FakeCAD, and the domain-independent action embeddings were processed using the virtual CAD domain model to generate a probability distribution(s) over the action space of virtual CAD.
[0074] Based on the one or more probability distributions generated at block 406, at block 408, the system may identify a second plurality of actions that can be performed using, for example, the first computer application by the ML module 108 or the meaning adjustment agent 122. At block 410, the system may cause the output to be presented at one or more output devices, for example, by the interface module 106. The output may include guidance for navigating the first HCI to perform a meaning task using the first computer application. The guidance may be generated, for example, by the interface module 106, based on the identified second plurality of actions that can be performed using the first computer application. The HCI guidance may be provided in various forms. Visually, this may be presented as overlaid annotations (e.g., as shown in FIGS. 3A - 3D), animations, videos, written natural language output (e.g., text balloons), 3D renderings (e.g., in virtual reality or metaverse settings), etc. Audibly, the HCI guidance may be presented as natural language output, various sounds, or sound effects, etc.
[0075] The examples described herein mainly focus on providing HCI guidance across semantically similar computer applications, such as between the FakeCAD and virtual CAD of a CAD computer application, or between different spreadsheet applications. However, this is not intended to be limiting. As long as a particular meaning task is relatively domain - independent for a particular domain, that task can be used to generate actions in any number of other domains that are not otherwise similar. For example, setting a particular computer application to "dark mode" can be relatively universal and can thus be utilized to provide HCI guidance across various domains, such as other computer applications or even the operating system.
[0076] As another example, an automated assistant (sometimes referred to as a “virtual” assistant or “virtual agent”) can interface with any number of third-party agents, allowing the automated assistant to act as a liaison to perform tasks such as ordering goods or services, making reservations, booking rideshares, etc. Different companies offering similar services (e.g., rideshares) may require users to interact with their respective third-party agents using natural language dialogue. However, the natural language dialogue available for interacting with a first rideshare agent may differ from the dialogue used for interacting with a second rideshare agent. Nevertheless, the final parameters or “slot values” fulfilled to complete a rideshare request may be semantically similar, even if they are named differently, required at different points during the conversation, etc. Thus, the techniques described herein can be used to provide users with HCI guidance, e.g., auditory or visual natural language, graphical elements on a display, etc., that can help a user familiar with a first rideshare agent more efficiently interact with a second rideshare agent. For example, domain-independent action embeddings can be processed to generate a script for an automated agent acting as a liaison. This script can request the necessary parameter or slot values from the user in a domain-independent manner. The automated agent can then use these requested values to engage any rideshare agent without requiring the user to learn the nuances of each agent.
[0077] FIG. 5 is a block diagram of an exemplary computing device 510 that can optionally be utilized to implement one or more aspects of the technology described herein. In some embodiments, one or more of client computing devices 120-1 through 120-P, the semantic task guidance system 102, and / or other component(s) may include one or more components of the exemplary computing device 510.
[0078] The computing device 510 typically includes at least one processor 514 that communicates with several peripheral devices via a bus subsystem 512. These peripheral devices can include, for example, a storage subsystem 524 that includes a memory subsystem 525 and a file storage subsystem 526, a user interface output device 520, a user interface input device 522, and a network interface subsystem 516. The input and output devices enable user interaction with the computing device 510. The network interface subsystem 516 provides an interface to an external network and is coupled to a corresponding interface device within other computing devices.
[0079] The user interface input device 522 may include a pointing device such as a keyboard, mouse, trackball, touchpad, or graphics tablet, a scanner, a touch screen incorporated into a display, an audio input device such as a voice recognition system, microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 510 or a communication network.
[0080] The user interface output device 520 may include a non-visual display such as a display subsystem, a printer, a fax machine, or an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or any other mechanism for creating a visible image. The display subsystem may also provide a non-visual display via, for example, an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 510 to a user or another machine or computing device.
[0081] The storage subsystem 524 stores programming and data structures that provide some or all of the functionality of some of the modules described herein. For example, the storage subsystem 524 may include logic for implementing selected aspects of the method 400 of FIG. 4.
[0082] These software modules are generally executed either alone by processor 514 or in combination with other processors. Memory 525 used within storage subsystem 524 can include several memories, such as main random access memory (RAM) 530 for storing instructions and data during program execution, and read only memory (ROM) 532 in which fixed instructions are stored. File storage subsystem 526 can provide permanent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of a particular embodiment can be stored by file storage subsystem 526 within storage subsystem 524 or on some other machine accessible by processor(s) 514.
[0083] Bus subsystem 512 provides a mechanism for enabling the various components and subsystems of computing device 510 to communicate with each other as intended. Although bus subsystem 512 is shown schematically as a single bus, alternative embodiments of the bus subsystem may use multiple buses.
[0084] Computing device 510 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computing device 510 shown in FIG. 5 is intended only as a specific example for purposes of illustrating some embodiments. Many other configurations of computing device 510 are possible that have more or fewer components than the computing device shown in FIG. 5.
[0085] Although several embodiments have been described and illustrated herein, various other means and / or structures can be utilized to perform the functions and / or to obtain one or more of the results and / or advantages described herein, and each such variation and / or modification is to be regarded as within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications in which the present teachings are used. One of ordinary skill in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. Accordingly, it is to be understood that the foregoing embodiments are presented by way of example only, and that within the scope of the appended claims and their equivalents, embodiments may be practiced otherwise than as specifically described and claimed. Embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure if such features, systems, articles, materials, kits, and / or methods do not mutually conflict.
Claims
Claim 1 A method implemented using one or more processors, comprising: identifying a domain of a first computer application that is operable using a first human-computer interface (HCI); selecting, based on the identified domain, a domain model for converting between an action space of the first computer application and another space; processing action embeddings based on the selected domain model to generate one or more probability distributions over actions within the action space of the first computer application, wherein the action embeddings represent multiple actions previously performed using a second HCI of a second computer application to perform semantic tasks, generating one or more probability distributions; identifying a second plurality of actions that are executable using the first computer application based on the one or more probability distributions; causing an output to be presented on one or more output devices, the output including guidance for navigating the first HCI to perform the semantic task using the first computer application, the guidance being based on the identified second plurality of actions that are executable using the first computer application, causing the output to be presented on one or more output devices. Claim 2 The method of claim 1, wherein the domain model is trained to convert between the action space of the first computer application and a domain-independent action embedding space. Claim 3 The method of claim 1 or 2, wherein the domain model is trained to directly convert between the action space of the first computer application and the action space of the second computer application. Claim 4 The method of any one of claims 1 to 3, wherein the first HCI includes a graphical user interface. Claim 5 The method of claim 4, wherein the guidance for navigating the first HCI includes one or more visual annotations overlaid on the GUI. Claim 6 The method according to claim 5, wherein one or more of the visual annotations are rendered to draw attention to one or more graphical elements of the GUI. **Claim 7** The method according to any one of claims 1 to 6, wherein the guidance for navigating the first HCI includes one or more natural language outputs. **Claim 8** obtaining user input that conveys the semantic task; identifying the action embedding based on the semantic task; The method according to any one of claims 1 to 7, further comprising: **Claim 9** The user input includes natural language input, and the method includes: performing natural language processing (NLP) on the natural language input to generate a first task embedding representing the semantic task; determining a similarity measure between the first task embedding and the action embedding, and The method according to claim 8, wherein the action embedding is processed based on the similarity measure. **Claim 10** A system comprising one or more processors and a memory storing instructions, the instructions, in response to execution of the instructions, causing the one or more processors to: identify a domain of a first computer application that is operable using a first human-computer interface (HCI); select a domain model for conversion between an action space of the first computer application and another space based on the identified domain; processing an action embedding based on the selected domain model to generate one or more probability distributions over actions within the action space of the first computer application, wherein the action embedding represents a plurality of actions previously performed using a second HCI of a second computer application to perform a semantic task; generating one or more probability distributions; identifying a second plurality of actions that are executable using the first computer application based on the one or more probability distributions; causing the output to be presented on one or more output devices, the output including guidance for navigating the first HCI to perform the semantic task using the first computer application, the guidance being based on the identified second plurality of actions that are executable using the first computer application, and causing the output to be presented on one or more output devices; A system for executing. **Claim 11** The system according to claim 10, wherein the domain model is trained to convert between the action space of the first computer application and a domain-independent action embedding space. **Claim 12** The system according to claim 10 or 11, wherein the domain model is trained to directly convert between the action space of the first computer application and the action space of the second computer application. **Claim 13** The system according to any one of claims 10 to 12, wherein the first HCI includes a graphical user interface. **Claim 14** The system according to claim 13, wherein the guidance for navigating the first HCI includes one or more visual annotations overlaid on the GUI. **Claim 15** The system according to claim 14, wherein one or more of the visual annotations are rendered to draw attention to one or more graphical elements of the GUI. **Claim 16** The system according to any one of claims 10 to 15, wherein the guidance for navigating the first HCI includes one or more natural language outputs. **Claim 17** further comprising instructions for obtaining user input that conveys the semantic task and identifying the action embedding based on the semantic task; The system according to any one of claims 10 to 16. **Claim 18** The user input includes natural language input, and the system further comprises instructions for performing natural language processing (NLP) on the natural language input to generate a first task embedding representing the semantic task, and determining a similarity measure between the first task embedding and the action embedding, wherein the action embedding is processed based on the similarity measure. The system according to claim 17. **Claim 19** A non-transitory computer-readable medium containing instructions, which, in response to execution of the instructions by a processor, cause the processor to identify a domain of a first computer application that is operable using a first human-computer interface (HCI), select a domain model for conversion between an action space of the first computer application and another space based on the identified domain, generate one or more probability distributions over actions within the action space of the first computer application by processing action embeddings based on the selected domain model, where the action embeddings represent a plurality of actions previously performed using a second HCI of a second computer application to perform a semantic task, generating one or more probability distributions, identify a second plurality of actions that are executable using the first computer application based on the one or more probability distributions, cause an output to be presented on one or more output devices, the output including guidance for navigating the first HCI to perform the semantic task using the first computer application, the guidance being based on the identified second plurality of actions that are executable using the first computer application, causing an output to be presented on one or more output devices, to execute. A non-transitory computer-readable medium. **Claim 20** The computer-readable medium according to claim 19, wherein the domain model is trained to convert between the action space of the first computer application and a domain-independent action embedding space.
Citation Information
Patent Citations
Information processing system and information processing method
JP2017062647A
Program
JP2019028947A
Image forming apparatus
US20120257904A1
Cited By
Method for supporting program creation, program, and program creation support system
JP7794517B1