User Interface Automation Navigation

The UI automation navigation system enhances real-world UI navigation by generating tokens from UI elements, converting them into task-specific feature representations, and predicting target elements for actions, addressing the limitations of existing technologies in handling complex tasks.

JP2026500078APending Publication Date: 2026-01-06MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025521953
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-07
Filing Date
2023-11-07
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing UI automation navigation technologies struggle with real-world web pages due to the complexity of navigation tasks and the difficulty in designing reward functions for reinforcement learning, and multi-task learning solutions fail to effectively handle diverse tasks on the same web page.

Method used

A UI automation navigation system that generates quantitative representations of UI elements as tokens, converts them into feature representations using task-specific information, and determines target elements for actions through a predictor model, utilizing hybrid embedding layers for efficient navigation across various tasks.

Benefits of technology

Improves the accuracy and effectiveness of UI navigation tasks in real-world environments by accurately identifying target elements and performing actions, allowing for broad applicability across different navigation tasks without requiring large-scale training data or complex reward functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026500078000001_ABST
    Figure 2026500078000001_ABST
Patent Text Reader

Abstract

According to an implementation of the present disclosure, a solution for automatic navigation of a user interface (UI) is provided. According to the solution, for UI elements, tokens representing the UI elements are generated. These UI elements include at least one or more UI elements in a currently presented user interface. The tokens are converted into individual feature representations of the UI elements using at least specific information corresponding to a current navigation task. Based on the feature representations, a target element is determined from the UI elements for the current navigation task. An action associated with the target element is performed. This thereby helps improve performance of various navigation tasks by using the navigation task specific information.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] background In work and life, various websites and applications (APPs) have become common tools for transmitting information and realizing functions. The user interfaces (UIs) of websites and APPs contain rich and diverse graphic and text information. Interaction with the UI is almost essential for browsing websites and using applications. By completing a series of UI navigation tasks, users can perform corresponding actions and realize multiple functions. Summary of the Invention [Means for solving the problem]

[0002] overview According to an implementation of the present disclosure, a solution for UI automation navigation is provided. In the solution, for a set of user interface elements, a set of tokens that individually represent the set of user interface elements is determined. The set of user interface elements includes at least one or more user interface elements in a currently presented user interface. The set of tokens is converted into individual feature representations of the set of user interface elements by using at least certain information corresponding to a current navigation task. Based on the feature representations, a target element is determined for the current navigation task from the one or more user interface elements. An action associated with the target element is performed. According to an implementation of the present disclosure, the quantitative representation of the UI elements is processed with information specific to the navigation task, thereby reflecting the information specific to the navigation task in the quantitative representation of the UI elements. This is useful for improving the performance of each navigation task.

[0003] This Summary is provided to introduce in general form a selection of subject matter that is further described below in specific embodiments. This section is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. [Brief explanation of the drawings]

[0004] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] FIG. 1 illustrates a block diagram of an example environment in which implementations of the present disclosure may be implemented. [Figure 2] 1 illustrates an example architecture for UI automation navigation according to some implementations of the present disclosure. [Figure 3] 1 illustrates an example of a predictor according to some implementations of the present disclosure. [Figure 4] 1 illustrates an example of a hybrid embedding layer in a predictor according to some implementations of the present disclosure. [Figure 5] 1 illustrates a flowchart of a process for UI automation navigation according to some implementations of the present disclosure. [Figure 6] 1 shows a schematic block diagram of an electronic device in which various implementations of the present disclosure may be implemented. DETAILED DESCRIPTION OF THE INVENTION

[0005] Detailed Description Hereinafter, implementations of the present disclosure will be described with reference to some exemplary implementations, which should be understood as being described in order to not imply any limitation on the scope of the present disclosure, but to enable those skilled in the art to better understand and implement the present disclosure accordingly.

[0006] As used herein, the term "comprises" and variations thereof should be interpreted as open, meaning "comprises but is not limited to." The term "based on" should be read as "based at least in part on." The terms "an implementation" and "one implementation" should be interpreted as "at least one implementation." The term "another implementation" should be interpreted as "at least one further implementation." The terms "first," "second," and the like may refer to different or identical entities. Other explicit or implied definitions may also be included below.

[0007] It should be noted that any section / subsection titles provided herein are not intended to be limiting. This disclosure describes various implementations throughout, and any type of implementation may be included under any section / subsection. In addition, an implementation described in any section / subsection may be combined in any manner with one or more other implementations described in the same section / subsection and / or different sections / subsections.

[0008] As used herein, a set of elements, an element set, or similar expressions may include one or more such elements. A set of elements may be ordered or unordered. For example, a "set of UI elements" may include one or more UI elements, and a "set of tokens" may include one or more tokens.

[0009] As used herein, "UI element" means a component of an interface presented to a user on a UI, which may be defined at any appropriate level of granularity for human-computer interaction. For example, UI elements may include, without limitation, images, text, icons, buttons, drop-down menus, search boxes, input boxes, etc. In some implementations, a UI element may be an integral, fundamental part of the UI.

[0010] As used herein, the term "UI navigation" or "navigation" refers to a guide for providing interactive guidance to a user of a UI to assist or replace the user in the corresponding operation and to achieve the required functionality.

[0011] As used herein, the term "model" refers to a model that can learn an association between corresponding inputs and outputs from training data, such that a corresponding output can be generated for a given input after training. Model generation can be based on machine learning techniques. Deep learning (DL) is a machine learning algorithm that processes inputs and provides corresponding outputs by using multi-layer processing units. A neural network model is an example of a model based on deep learning. Also, in this disclosure, a "model" can be referred to as a "machine learning model," a "learning model," a "machine learning network," or a "learning network," and these terms are used interchangeably in this disclosure.

[0012] Generally, machine learning can have three stages: a training stage, a testing stage, and an estimation stage (also called an inference stage). In the training stage, a given model can be trained with a large amount of training data, and iterations can continue until the model can obtain consistent estimates from the training data that meet the predicted goal. Through training, the model can be considered to be able to learn associations between inputs and outputs (also called input-output mappings) from the training data. Parameter values ​​of the trained model are determined. In the testing stage, test inputs are applied to the trained model to test whether the model can provide the correct output to determine the model's performance. In the estimation stage, the model can be used to process actual inputs and determine corresponding outputs based on the parameter values ​​obtained from training.

[0013] Example Environment 1 shows a schematic diagram of an example environment 100 in which implementations of the present disclosure may be implemented. Within the environment 100, a user 101 is interacting with a user device 102 to complete a UI navigation task.

[0014] User device 102 may be any type of mobile, fixed, or portable device, including a mobile phone, desktop computer, laptop computer, notebook computer, netbook computer, tablet computer, media computer, multimedia tablet, personal communication system (PCS) device, personal navigation device, personal digital assistant (PDA), audio / video player, digital / video camera, positioning device, TV receiver, radio receiver, e-book device, gaming device, or any combination of the above, including accessories and peripherals of these devices. Also, in some implementations, user device 102 may support any type of user-specific interface (such as "wearable" circuitry, etc.).

[0015] One or more UIs are presented to the user 101 through the user device 102. These UIs may include website web pages, APP pages, document pages, etc. Each UI includes one or more UI elements. One or more UI navigation tasks (also referred to as one or more navigation tasks) can be completed through interactions with the UI to achieve functionality desired by the user 101. Such navigation tasks may include, without limitation, logging in, changing a password, registering an account, keyword searching, cookie settings, adding to a shopping cart, removing a pop-up window, etc.

[0016] In the example of Figure 1, the presented current UI 110 includes one or more UI elements, also referred to as current elements. As shown in Figure 1, current UI 110 includes current elements 115-1, 115-2, 115-3, 115-4, and 115-5, also referred to individually or collectively as one or more current elements 115.

[0017] It should be understood that the environment 110 shown in FIG. 1 is for illustrative purposes only and is not intended to limit the scope of the present disclosure. The current UI and the number and types of UI elements shown in FIG. 1 are schematic and are not intended to limit the scope of the present disclosure. In implementations of the present disclosure, the processed UIs can be any suitable type of UI and can include any suitable number and types of UI elements. Additionally, while the text in the UI 110 is shown in English, this is for illustrative purposes only and is not intended to limit the scope of the present disclosure. In implementations of the present disclosure, the processed UIs can be provided in any suitable language or languages.

[0018] By correctly identifying, UI understanding and effective navigation of the UI can assist and complete navigation tasks on behalf of the user, thereby accurately and efficiently implementing the functionality required by the user. Therefore, it is expected to effectively navigate the UI and further implement UI automation navigation.

[0019] Some solutions provide UI navigation for web navigation tasks. However, these solutions primarily involve web navigation tasks in simulated website environments. The web pages contained within simulated website environments are typically much simpler than those in the real world. In addition, the predefined navigation tasks within these simulated website environments are typically very simple and involve special commands, such as clicking a button below a text box. Therefore, these simulated navigation tasks are much simpler than those in the real world.

[0020] Among these solutions, the reinforcement learning-based solution requires large-scale training data and reward functions. However, in the case of real-world navigation tasks, it is difficult to design a reward function that can provide accurate feedback, and collecting a large amount of data is expensive. The multi-task learning solution utilizes knowledge shared among multiple related navigation tasks. In this solution, different navigation tasks are associated with different simulation environments. However, in real-world websites, different navigation tasks may be performed on the same web page. As a result, it is difficult to apply this solution to real-world web pages. In addition to web pages, other types of UIs also have similar problems.

[0021] Example implementations of the present disclosure provide a solution for UI automated navigation. According to various implementations of the present disclosure, in response to a current UI being presented, tokens (i.e., quantitative representations of UI elements) representing the UI elements are generated for the UI elements. These UI elements include at least one or more elements in the current UI. The generated tokens are converted into individual feature representations by using at least information specific to the current navigation task. Based on these feature representations, a target element is determined from the UI elements of the current UI. An action associated with the target element, such as clicking the target element, entering one or more texts within the area of ​​the target element, etc., is then performed.

[0022] In an implementation of the present disclosure, quantitative representations of UI elements are converted into UI element features using navigation task-specific information, thereby reflecting the navigation task-specific information in the UI element features. This helps improve the effectiveness and accuracy of each navigation task, and the task-specific information can be specific to a real-world navigation task, allowing the implementation of the present disclosure to effectively process UIs in the real world. Meanwhile, multiple navigation tasks can be automatically processed by using information specific to different navigation tasks. This results in broad applicability of the implementation of the present disclosure.

[0023] Some exemplary implementations of the present disclosure will now be described in more detail with reference to the drawings.

[0024] Example Page Navigation Architecture FIG. 2 illustrates an example architecture 200 for UI automation navigation according to some implementations of the present disclosure. An example operation of architecture 200 is described below with reference to FIG. 1. A tokenizer 210 is used to generate tokens representing UI elements, which may be quantitative representations of the UI elements. Specifically, for a set of UI elements, tokenizer 210 generates a set of tokens that individually represent the set of UI elements. Each token corresponds to a UI element. The set of UI elements includes at least the current element in current UI 110. In the example of FIG. 2, tokens 201-1, 201-2, 201-3, 201-4, and 201-5 correspond to current elements 115-1, 115-2, 115-3, 115-4, and 115-5, respectively, in current UI 110.

[0025] In some implementations, the set of UI elements may further include UI elements used as historical target elements in a historical interaction (i.e., UI elements that are interacted with in a historical interaction), also referred to as historical elements. For example, token 201-6 represents historical element 215. Hereinafter, tokens 201-1, 201-2, 201-3, 201-4, 201-5, and 201-6 may be referred to collectively as set of tokens 201 or individually as tokens 201.

[0026] The historical elements may include individual target elements involved in an interaction trajectory starting from the first UI. Furthermore, in some implementations, if the interaction trajectory is relatively long, one or more historical target elements in the historical interactions whose distance from the current UI is less than a threshold distance may be considered, while one or more historical interactions whose distance is greater than the threshold may not be considered. The reason for this is that a historical interaction that is too far away from the current UI is likely not relevant to the navigation task of the current UI. As an example, the distance between the current UI and a historical interaction may be measured by the number of interactions since the occurrence of the corresponding historical interaction.

[0027] Navigation tasks such as changing a password, registering an account, and adding to a shopping cart actually involve a series of interactions. By considering the historical target elements of the interaction trajectory, it is useful to accurately predict the target element for the current navigation task. As a result, the performance of completing this kind of sequential navigation task can be improved.

[0028] To generate tokens 201 representing UI elements, UI information 250 can be provided to tokenizer 210. UI information 250 may include UI images 251, such as screenshots, which provide visual information. UI information 250 may further include UI metadata 252, such as a hierarchical structure representing relationships between UI elements. In the case where the UI is a web page, metadata 252 may include a document object model (DOM). In the case where the UI is an APP page, metadata 252 may include a view level (VH). It should be understood that images 251 and metadata 252 are merely examples of UI information and are not intended to limit the scope of the present disclosure. In implementations of the present disclosure, tokenizer 210 can utilize any suitable type of UI information.

[0029] The tokenizer 210 can extract element information of the UI element from the UI information 250 to describe one or more aspects of the UI element.

[0030] In some implementations, the extracted element information may include a type of UI element. The type may include, without limitation, “input,” “clickable element,” “plain text,” and “icon.” The type of the UI element may be determined based on UI metadata. In one example, if the label name of a UI element (e.g., a DOM) in the metadata is “input,” the type of the UI element is “input.” If the label name of a UI element in the metadata is “button,” the type of the UI element is “clickable element.” If the text of the UI element is not empty, the type of the UI element is “plain text.” If the text of the UI element is empty, the type of the UI element is “icon.” For the current UI 110 of the example of FIG. 1 , the type of current element 115-1 is “icon,” the type of current element 115-2 is “plain text,” the types of current elements 115-3 and 115-4 are “input,” and the type of current element 115-5 is “clickable element.” The types of UI elements and their classifications described herein are examples only and are not intended to limit the scope of the present disclosure. In implementations of the present disclosure, UI elements can be classified in any suitable manner.

[0031] Alternatively or additionally, in some implementations, the extracted element information may include a text description of the UI element. The UI metadata may include multiple items for a UI element. These items can be concatenated as a string as the text description of the UI element. For UI elements that do not have meaningful items in the metadata (such as elements with an "icon" type), a classifier can be used to classify the UI element. The label of the classified category can be used as the text description of the UI element.

[0032] Alternatively or additionally, in some implementations, the extracted element information may include the location of the UI element within the UI where it is located. In the case of the current element 115, the location is the location of the current element 115 within the current UI 110. In the case of the historical element 215, the location is the location of the historical element 215 within the UI with which the historical element 215 is being interacted with. Such a location may be represented by the two-dimensional position of the UI element within the image 251.

[0033] Alternatively or additionally, in some implementations, the extracted element information may include the time of occurrence of the UI element relative to the current UI 110, also referred to as temporal location. Temporal location may be measured by the distance between the current UI 110 and the past interactions, as described above. For example, the temporal location of the current element 115 may be 0 because the current element 115 is located within the current UI 110. The temporal location of the past element that served as the target element in the last interaction may be 1, and so on. Using temporal location, the current element may be explicitly distinguished from past elements, and different past elements may be distinguished.

[0034] In the above-described implementations, element information of UI elements is extracted from UI information 250 by tokenizer 210. Alternatively, in some implementations, the above-described element information may be extracted by other modules and provided to tokenizer 210.

[0035] A tokenizer 210 then generates tokens based on the element information of the UI elements. For example, for a UI element, one or more aspects of its element information (such as the type, text, position, etc.) can be quantified, and the quantified element information can be concatenated as a token. The generated tokens 201 are provided to a predictor 220.

[0036] In addition to the token 201, the input of the predictor 220 further includes the task type 202 of the current navigation task. The navigation task type may include, without limitation, login, password change, account registration, keyword search, cookie settings, add to shopping cart, remove popup window, etc. It should be understood that these navigation task types are examples only and are not intended to limit the scope of the present disclosure. Implementations of the present disclosure may process any suitable type of navigation task.

[0037] In some implementations, the current navigation task can be defined by the user 101. For example, a selection of multiple predefined navigation tasks can be presented via the user device 102. The user 101 can select a navigation task from these predefined navigation tasks. Furthermore, the current navigation task can be determined based on the user selection. In cases where the predictor 220 is implemented by a machine learning model, the predictor 220 is pre-trained using data related to these predefined navigation tasks.

[0038] Alternatively, in some implementations, the current navigation task may be determined in other ways, for example, the current navigation task may be predicted by another module, and implementations of this disclosure are not limited in this respect.

[0039] In some implementations, the input of the predictor 220 may further include a relative relationship 203 between UI elements. The relative relationship 203 may include a relative relationship between different current elements 115, a relative relationship between the current element 115 and a historical element 205, and a relative relationship between different historical elements. The relative relationship between two UI elements may include the relative encoded position of the two UI elements obtained from UI metadata. Taking a web page as an example, a relative form ID in the DOM can be used to represent the relative encoded position. If the form ID of a UI element is 0, this means that the UI element does not have any form. For any two UI elements, the value of the relative encoded position can be determined based on whether the form IDs of the two elements are 0 and whether the two elements are located within the same UI. It should be understood that the form IDs described herein are merely examples and are not intended to limit the scope of the present disclosure. Depending on the type of UI metadata, any suitable type of relative encoded position can be used.

[0040] Alternatively or additionally, the relative relationship between two UI elements may also include the relative spatial positions of the two UI elements. If two UI elements (such as two current elements 115) are located within the same UI, the relative spatial positions of the two UI elements may be expressed by their relative distance within the UI. If the two UI elements are located within different UIs, the relative spatial positions may be expressed by a default value.

[0041] In this implementation, the relative relationship 203 explicitly represents the relationship between any two UI elements, which is useful for the predictor 220 to find the current target element from a UI element.

[0042] The predictor 220 predicts a target element within the current elements 115 for a current navigation task based on the token 201, the task type 202, and the optional relative relationship 203. Specifically, based on the task type 202, the predictor 220 determines specific information corresponding to the current navigation task. The specific information may inform a transformation to be performed on the token 201 for the current navigation task, which is used to convert the token into a feature representation. This type of information used in the transformation is specific to the current navigation task and is not shared between different navigation tasks. The predictor 220 may have specific information corresponding to multiple predefined navigation tasks. Based on the task type 202, the predictor 220 may select the specific information corresponding to the current navigation task. The predictor 220 then uses this type of specific information to convert the token 201 into a feature representation of a UI element. It should be understood that the feature representation obtained in this manner may reflect features more appropriate for the current navigation task.

[0043] In some implementations, predictor 220 can also use shared information for multiple navigation tasks to convert tokens 201 into feature representations of UI elements. The shared information may inform common transformations for the navigation tasks to be performed on tokens 201. This kind of shared information may reflect common knowledge among different navigation tasks, which is useful for further improving the performance of UI navigation. The shared information may include one or more shared information items. Similarly, the specific information may include one or more specific information items. With reference to FIG. 3 , the following paragraphs will provide an example of converting tokens into feature representations according to a combination of shared information and specific information.

[0044] Based on the individual feature representations of the UI elements, the predictor 220 obtains a prediction result 206. The prediction result 206 may include a probability that the current element 115 will function as the target element. The current element with the highest probability may be determined as the target element for the current navigation task. In the example of FIG. 2 , P1, P2, P3, P4, and P5 in the prediction result 206 are probabilities that the current elements 115-1, 115-2, 115-3, 115-4, and 115-5 are target elements, respectively. P3 is assumed to have the maximum value. Therefore, the current element 115-3 is determined as the target element. Alternatively or additionally, the prediction result 206 may include a probability that the current element 115 is not the target element.

[0045] An action associated with the target element is performed in response to determining the target element. The action performed is related to the type of the target element. For example, if the type of the target element is "clickable element," the action performed is to click the target. In another example, if the type of the target element is "input," the action performed is to input default content. In the example of FIG. 2, the account name "ABCDE" is being input within the area of ​​the current element 115-3.

[0046] Any suitable algorithm may be used to implement the predictor 220. In some implementations, a machine learning model may be used to implement the predictor 220. In this implementation, the specific information and optional shared information described above may be implemented as model parameters. Each specific information item may be implemented as a parameter of a different module of the model. Similarly, the shared information item may be implemented as a parameter of a different module of the model. An example of this will be provided below.

[0047] An example architecture 200 for UI automation navigation has been described above with reference to Figure 2. It should be understood that the UI, UI elements, UI information, and predicted target elements shown in Figure 2 are examples only and are not intended to limit the scope of the present disclosure. In addition, it should be understood that the target elements determined for the current UI 110 can be used as history elements for subsequent UIs.

[0048] Example Predictors As described with reference to FIG. 2 , in some implementations, a machine learning model can be used to implement the predictor 220. An example of the predictor 220 will now be described with reference to FIG. 3 . As shown in FIG. 3 , the predictor 220 generally includes a feature extraction module 310 and an output module 320. Within the feature extraction module 310, the set of tokens 201 is transformed into at least one set of transformed tokens by using specific information corresponding to the current navigation task and shared information for multiple predefined navigation tasks. Each token 201 is transformed using the shared information and the specific information. Then, based on the transformed tokens, individual feature representations of UI elements (including the current element and optional historical elements) can be generated. In the output module 320, a target element is predicted for the current navigation task based on the individual feature representations of the UI elements. The predictor 220 can be implemented using any suitable network structure. In the example of FIG. 3 , a transformer is used to implement the predictor 220. Each token 201 is fed to a hybrid embedding layer 311, a hybrid embedding layer 312, and a task-sharing embedding layer 313 to generate a transformed token as a query (Q), a transformed token as a key (K), and a transformed token as a value (V), respectively. The hybrid embedding layer 311 and the hybrid embedding layer 312 need to use parameters specific to the current task. Therefore, a task type 202 is fed to the hybrid embedding layer 311 and the hybrid embedding layer 312.

[0049] The hybrid embedding layer 400 shown in Figure 4 can be considered as an example implementation of the hybrid embedding layer 311 and the hybrid embedding layer 312 in Figure 3. The task-shared embedding layer 410 has shared parameters for multiple predefined navigation tasks, and the shared parameters can be considered as an example implementation of the shared information items. By using the shared parameters, tokens input to the hybrid embedding layer 400 are converted into intermediate tokens. For example, the task-shared embedding layer 410 can be implemented as a fully connected (FC) layer. The parameters of the fully connected layer are shared by multiple predefined navigation tasks.

[0050] The intermediate tokens are then fed to a task-specific embedding layer 420. The task-specific embedding layer 420 may have specific parameters for multiple predefined navigation tasks, and these specific parameters may be considered as example implementations of specific information items. Based on the task type 202, the task-specific embedding layer 420 applies specific parameters corresponding to the current navigation task to the intermediate tokens to generate transformed tokens. That is, the task-specific embedding layer 420 performs task-adaptive embedding. For example, the task-specific embedding layer 420 may be implemented as a dynamic fully-connected layer. The parameters of the dynamic fully-connected layer dynamically change depending on the task type 202.

[0051] The task-shared embedding layer encodes information shared among different tasks, and the task-specific embedding layer encodes information specific to the current task. This decouples the learning of task-shared knowledge and task-specific knowledge, thereby allowing for the processing of multiple navigation tasks under a common framework. It should be understood that the structure of the hybrid embedding layer shown in FIG. 4 is merely an example. The hybrid embedding layer may have other structures. For example, a task-specific embedding layer may be placed before a task-shared embedding layer. In another example, the hybrid embedding layer may include multiple task-specific embedding layers and / or multiple task-shared embedding layers.

[0052] Referring again to FIG. 3 , the hybrid embedding layer 311 transforms each token 201. The generated first set of transformed tokens is provided to the attention layer 314 as a query. The hybrid embedding layer 312 transforms each token 201. The generated second set of transformed tokens is provided to the attention layer 314 as a key. The task-shared embedding layer 313 is similar to the task-shared embedding layer 410; i.e., the task-shared embedding layer 313 has shared parameters for multiple predefined navigation tasks. The third set of transformed tokens generated by the task-shared embedding layer 313 is provided to the attention layer 314 as a value. The attention layer 314 determines attention information based on the first set of transformed tokens, the second set of transformed tokens, and optional relative relationships 203. The attention information indicates correlations between any two UI elements, such as correlations between any two current elements 115, correlations between the current element and a historical element, and correlations between two historical elements. The attention information may be, for example, an attention weight matrix. The attention layer 314 then weights the third set of transformed tokens based on the attention information. The weighted transformed tokens are passed through a dropout and normalization layer 315, a feedforward layer 316, and the like to generate individual feature representations of the UI elements.

[0053] 3 shows one feature extraction module 310, but this is for example purposes only. In some implementations, the predictor 220 can include multiple cascaded feature extraction modules with identical structures.

[0054] The feature representations of the UI elements are provided to an output module 320, which predicts target elements based on the feature representations. As shown in FIG. 3 , in some implementations, the output module 320 can include a hybrid embedding layer 321. The hybrid embedding layer 321 can have a structure similar to that of the hybrid embedding layer 400. That is, a task-shared embedding layer within the hybrid embedding layer 321 has shared parameters for multiple predefined navigation tasks, and task-specific embedding layers have individual-specific parameters for the multiple predefined navigation tasks. Based on the task type 202, the hybrid embedding layer 321 applies specific parameters corresponding to the current navigation task. Thus, within the hybrid embedding layer 321, the shared parameters and specific parameters are used to transform the feature representations of the UI elements. The head layer 322 determines the probability that each UI element functions as a target element and / or the probability that each UI element does not function as a target element based on the transformed feature representations. For example, the head layer 322 can be implemented as a linear layer.

[0055] An exemplary operation of predictor 220 is now described with reference to Figure 3. For any given token

number

number

number

number

[0056]

number

[0057]

number

number

[0058] By applying a hybrid embedding layer to the tokens used as queries and keys, an adaptive attention mechanism can be implemented for navigation tasks. The attention scores of the jth and ith UI elements may be expressed as:

number

number

number

[0059] The softmax function is used to transform the attention scores to determine the attention weights. Thus, the attention weight α for the j-th UI element relative to the i-th UI element is i,j (T) can be calculated using the following formula:

number

[0060] Therefore, the feature of the i-th UI element output by the attention layer 314 is expressed as follows:

number

[0061] Through the operations expressed by the above equations, the feature representation of the UI element can be obtained through cascaded transformer blocks. The hybrid embedding layer 321 in the output module 320 applies an operation similar to equation (2) to the feature representation. The head layer 322 may be implemented as a linear layer, which has the following properties: z Each feature vector having dimensions is converted into one dimension (corresponding to the probability of being the target element) or two dimensions (corresponding to the probability of being the target element and the probability of not being the target element, respectively).

[0062] In the implementation described above with reference to Figure 3, shared information is implemented as parameters of a task-shared embedding layer, while specific information is implemented as parameters of task-specific embedding layers. In this implementation, a hybrid embedding layer is applied to tokens used as queries and keys, and a task-shared embedding layer is applied to tokens used as values. Thus, in the attention mechanism, tokens used as values ​​reflect unknown features of the task, while tokens used as queries and keys reflect features depending on the task. A predictor implemented in this way can use multi-task data in training and benefit from strategy learning facilitated by different tasks.

[0063] In training the predictor shown in FIG. 3, a set of “successful” human-computer interaction trajectories can be used as training data and supervision information. Human-selected UI elements can be used as ground truth to supervise the training of the predictor 220. Such a predictor 220 is universal because it integrates strategy learning for multiple navigation tasks. By learning common representations across different tasks within a joint framework, the predictor 220 can fully utilize the collected data. Therefore, the predictor 220 has high sample efficiency. Additionally, only “successful” human-computer interaction data is needed for training, thereby avoiding the requirement to collect various interaction trajectories and design complex reward functions.

[0064] It should be understood that the structure of predictor 220 described with reference to Figure 3 is an example. Implementations of the present disclosure are not limited to configurations such as hybrid embedding layers, task-shared embedding layers, and task-specific embedding layers in the manner shown in Figure 3. For example, in some implementations, task-specific embedding layers can be used for queries, keys, and values. In another example, in some implementations, hybrid embedding layers can be used for queries, keys, and values. In another example, in some implementations, task-specific embedding layers can be used for queries and keys, and task-shared embedding layers can be used for values.

[0065] Additionally, although a Transformer is used to implement Predictor 220 in the example of Figure 3, this is by way of example only. In other implementations, Predictor 220 can be implemented by any known or future-developed machine learning model.

[0066] Example Flow 5 shows a flowchart of a process 500 for UI automation navigation according to some implementations of the present disclosure. Process 500 can be implemented in the user device of FIG. 1 or in another computing device, such as a device that provides UI navigation services. At block 510, a set of tokens, each representing a set of UI elements, is generated for a set of UI elements. The set of UI elements includes at least one or more UI elements in the current user interface being presented.

[0067] In some implementations, the set of UI elements can further include a UI element that serves as a history target element in a history interaction.

[0068] In some implementations, generating the set of tokens includes generating a token representing a given user interface element of the set of user interface elements based on at least one of the following items related to the given user interface element: a type of the given user interface element, a text description of the given user interface element, a position of the given user interface element within a user interface in which the given user interface element is located, or a time of occurrence of the given user interface element relative to the current user interface.

[0069] At block 520, the set of tokens is converted into individual feature representations of the set of UI elements using at least certain information corresponding to the current navigation task. At block 525, a target element is determined for the current navigation task from one or more UI elements based on the feature representations. By way of example, the current navigation task may include, without limitation, logging in, changing a password, registering an account, keyword searching, setting cookies, adding to a shopping cart, removing a pop-up window, etc.

[0070] In some implementations, the current navigation task can be defined by a user. A selection for multiple default navigation tasks can be presented. The multiple default navigation tasks include the current navigation task. A user selection can be received for one of the multiple default navigation tasks. The current navigation task can be determined based on the user selection. For example, the navigation task selected by the user is the current navigation task.

[0071] In some implementations, a set of tokens can be transformed into at least one set of transformed tokens using shared information and specific information for multiple predefined navigation tasks. The multiple predefined navigation tasks include a current navigation task. The shared information and specific information can be implemented as parameters of any machine-executable algorithm (such as a machine learning model). In some implementations, a set of tokens can be transformed into a set of intermediate tokens using a first shared information item in the shared information (e.g., parameters of the task-shared embedding layer 410). A set of intermediate tokens can be transformed into at least one set of transformed tokens using a first specific information item corresponding to the current navigation task in the specific information (e.g., parameters of the task-specific embedding layer 420). For example, a first set of transformed tokens to be used as a query can be generated by the hybrid embedding layer 311. As another example, a second set of transformed tones to be used as a key can be generated by the hybrid embedding layer 312.

[0072] Individual feature representations of the set of UI elements can be determined based on the at least one set of transformed tokens. In some implementations, the at least one set of transformed tokens includes a first set of transformed tokens and a second set of transformed tokens corresponding to the set of user interface elements. For example, the first and second sets of transformed tokens can be generated by hybrid embedding layers 311 and 312, respectively.

[0073] A relative relationship between a first user interface element and a second user interface element of the set of user interface elements can be obtained. Attention information can be determined based on the first set of transformed tokens, the second set of transformed tokens, and the relative relationship. The attention information indicates a correlation between the first user interface element and the second user interface element. The set of tokens can be transformed into a third set of transformed tokens corresponding to the set of user interface elements using a second shared information item in the shared information (e.g., a parameter of the task shared embedding layer 313). For example, the third set of transformed tokens as values ​​can be generated through the task shared embedding layer 313. A feature representation can be determined by weighting the third set of transformed tokens based on the attention information.

[0074] In some implementations, the relative relationship between the first and second UI elements can include relative encoded positions of the first and second UI elements obtained from user interface metadata, such as a relative form ID, or alternatively or additionally, the relative relationship between the first and second UI elements can include relative spatial positions of the first and second UI elements.

[0075] Furthermore, the target elements can be determined based on the feature representation. In some implementations, the feature representation can be transformed using a third shared information item in the shared information (e.g., task-shared embedding layer parameters in the hybrid embedding layer 321) and a third specific information item in the specific information corresponding to the current navigation task (e.g., task-specific embedding layer parameters in the hybrid embedding layer 321). Probabilities that one or more UI elements are one or more target elements can be determined based on the transformed feature representation.

[0076] At block 530, an action associated with the target element is performed. For example, depending on the type of target element, the target element may be clicked or predefined content may be entered within the area of ​​the target element.

[0077] Sample Device Figure 6 shows a schematic block diagram of an electronic device capable of implementing various implementations of the present disclosure. It should be understood that the electronic device 600 shown in Figure 6 is merely an example and is not intended to constitute any limitation on the functionality and scope of the implementations described in the present disclosure.

[0078] 6, electronic device 600 has the form of a general-purpose computing device. Components of electronic device 600 may include, without limitation, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660.

[0079] In some implementations, electronic device 600 may be implemented as a computing device, a computing system, a server, a mainframe, and other devices having computing capabilities.

[0080] The processing unit 610 may be a real or virtual processor and may execute various processes according to programs stored in memory 620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of the electronic device 600. The processing unit 610 may include a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor, a controller, and / or a microcontroller.

[0081] Electronic device 600 typically includes multiple computer storage media. Such media may be any available media accessible to electronic device 600, including, without limitation, volatile and nonvolatile media, removable and non-removable media. Memory 620 may include volatile memory (registers, cache, random access memory (RAM)), non-volatile memory (read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 may include removable or non-removable media, such as memory, flash drives, disks, or any other medium that can be used to store information and / or data and that can be accessed within electronic device 600.

[0082] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 6, a disk drive for reading from or writing to a removable non-volatile disk and an optical disk drive for reading from or writing to a removable non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces.

[0083] The communications unit 640 facilitates communication with another computing device over a communications medium. Additionally, the functionality of the components of the electronic device 600 may be implemented within a single computing cluster or multiple computing devices that may communicate over a communications connection. Thus, the electronic device 600 may operate in a networked environment using logical connections to one or more other servers, personal computers (PCs), or other general network nodes.

[0084] The input device(s) 650 may be one or more of a variety of input devices, such as a mouse, a keyboard, a data import device, and the like. The output device(s) 660 may be one or more output devices, such as a display, a data export device, and the like. The electronic device 600 may also communicate with one or more external devices (not shown), such as a storage device, a display device, and the like, through the communication unit 640, as needed, by one or more devices that allow a user to interact with the electronic device 600 or by any device (such as a network card, a modem, and the like) that allows the electronic device 600 to communicate with one or more other computing devices. Such communication may be performed via an input / output (I / O) interface (not shown).

[0085] In addition to being integrated on a single device, in some implementations, some or all components of the electronic device 600 can also be configured in the form of a cloud computing architecture. In a cloud computing architecture, these components can be remotely configured and cooperate to implement the functionality described in this disclosure. In some implementations, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various implementations, cloud computing uses appropriate protocols to provide services over a wide area network such as the Internet. For example, a cloud computing provider may offer applications over a wide area network, which can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and corresponding data can be stored on servers at remote locations. Computing resources within a cloud computing environment can be combined at remote data center locations, or they can be distributed. A cloud computing infrastructure can provide services through a shared data center, even if it represents a single point of access for users. Thus, the components and functionality described herein may be provided from a service provider at a remote location using a cloud computing architecture, or alternatively, they may be provided from a traditional server, or they may be installed directly or otherwise on a client device.

[0086] The electronic device 600 can be used to implement hierarchical relationship analysis in various implementations of the present disclosure. The memory 620 may include one or more modules having one or more program instructions, which can be accessed and executed by the processing unit 610 to implement various implemented functions described herein. For example, the memory 620 may include a hierarchical relationship analysis module 625 for determining the structure of tables in an image. As shown in FIG. 6 , the electronic device 600 can obtain input required for UI navigation through the input device 650 and provide UI navigation output through the output device 660. In some implementations, the electronic device 600 can also receive input from other devices (not shown) via the communication unit 640.

[0087] Example Implementation Below, some example implementations of the present disclosure are listed.

[0088] In one aspect, the present disclosure provides a computer-implemented method that includes: generating, for a set of user interface elements, a set of tokens that individually represent the set of user interface elements, the set of user interface elements including at least one or more user interface elements in a currently presented user interface; converting the set of tokens into individual feature representations of the set of user interface elements using at least certain information corresponding to a current navigation task; determining targets for the current navigation task from the one or more user interface elements based on the feature representations; and performing an action associated with the target elements.

[0089] In some example implementations, the set of user interface elements further includes a user interface element that serves as a history target element in the history interaction.

[0090] In some example implementations, converting the set of tokens into individual feature representations includes converting the set of tokens into at least one set of converted tokens using shared information and specific information for a plurality of predefined navigation tasks, where the plurality of predefined navigation tokens includes the current navigation task, and determining individual feature representations of the set of user interface elements based on the at least one set of converted tokens.

[0091] In some example implementations, converting the set of tokens into at least one set of converted tokens includes converting the set of tokens into a set of intermediate tokens using a first shared information item in the shared information, and converting the set of intermediate tokens into a converted set of tokens of the at least one set of converted tokens using a first specific information item corresponding to the current navigation task in the specific information.

[0092] In some example implementations, at least one set of transformed tokens includes a first set of transformed tokens and a second set of transformed tokens corresponding to the set of user interface elements, and determining an individual feature representation of the set of user interface elements includes obtaining a relative relationship between the first user interface element and a second user interface element of the set of user interface elements; determining attention information based on the first set of transformed tokens, the second set of transformed tokens, and the relative relationship, wherein the attention information indicates a correlation between the first user interface element and the second user interface element; transforming the set of tokens into a third set of transformed tokens corresponding to the set of user interface elements using a second shared information item in the shared information; and determining the feature representation by weighting the third set of transformed tokens based on the attention information.

[0093] In some example implementations, the relative relationship includes at least one of a relative encoded position of the first user interface element and the second user interface element or a relative spatial position of the first user interface element and the second user interface element obtained from user interface metadata.

[0094] In some example implementations, determining the target element includes transforming the feature representation using a third shared information item in the shared information and a third specific information item in the specific information that corresponds to the current navigation task, and determining a probability that one or more user interface elements are the target element based on the transformed feature representation.

[0095] In some example implementations, generating the set of tokens includes generating a token representing a given user interface element of the set of user interface elements based on at least one of the following items related to the given user interface element: a type of the given user interface element, a text description of the given user interface element, a position of the given user interface element within the user interface in which it is located, or a time of occurrence of the given user interface element relative to the current user interface.

[0096] In some example implementations, the method further includes presenting options for a plurality of predefined navigation tasks including a current navigation task, receiving a user selection of a navigation task from the plurality of predefined navigation tasks, and determining the current navigation task based on the user selection.

[0097] In another aspect, the present disclosure provides an electronic device having a processor and a memory coupled to and having instructions stored thereon, the instructions, when executed by the processor, causing the electronic device to perform actions including: generating, for a set of user interface elements, a set of tokens that individually represent the set of user interface elements, the set of user interface elements including at least one or more user interface elements in a currently presented user interface; converting the set of tokens into individual feature representations of the set of user interface elements using at least certain information corresponding to a current navigation task; determining a target element for the current navigation task from the one or more user interface elements based on the feature representations; and performing an action associated with the target element.

[0098] In some example implementations, the set of user interface elements further includes a user interface element that serves as a history target element in the history interaction.

[0099] In some example implementations, converting the set of tokens into individual feature representations includes converting the set of tokens into at least one set of converted tokens using shared information and specific information for a plurality of predefined navigation tasks, where the plurality of predefined navigation tasks includes a current navigation task, and determining individual feature representations of the set of user interface elements based on the at least one set of converted tokens.

[0100] In some example implementations, converting the set of tokens into at least one set of converted tokens includes converting the set of tokens into a set of intermediate tokens using a first shared information item in the shared information, and converting the set of intermediate tokens into a converted set of tokens of the at least one set of converted tokens using a first specific information item corresponding to the current navigation task in the specific information.

[0101] In some example implementations, at least one set of transformed tokens includes a first set of transformed tokens and a second set of transformed tokens corresponding to the set of user interface elements, and determining an individual feature representation of the set of user interface elements includes obtaining a relative relationship between the first user interface element and a second user interface element of the set of user interface elements; determining attention information based on the first set of transformed tokens, the second set of transformed tokens, and the relative relationship, wherein the attention information indicates a correlation between the first user interface element and the second user interface element; transforming the set of tokens into a third set of transformed tokens corresponding to the set of user interface elements using a second shared information item in the shared information; and determining the feature representation by weighting the third set of transformed tokens based on the attention information. In some example implementations, the relative relationship includes at least one of a relative encoded position of the first user interface element and the second user interface element or a relative spatial position of the first user interface element and the second user interface element obtained from user interface metadata.

[0102] In some example implementations, determining the target element includes transforming the feature representation using a third shared information item in the shared information and a third specific information item in the specific information that corresponds to the current navigation task, and determining a probability that one or more user interface elements are the target element based on the transformed feature representation.

[0103] In some example implementations, generating the set of tokens includes generating a token representing a given user interface element of the set of user interface elements based on at least one of the following items related to the given user interface element: a type of the given user interface element, a text description of the given user interface element, a position of the given user interface element within a user interface in which the given user interface element is located, or a time of occurrence of the given user interface element relative to the current user interface.

[0104] In some example implementations, the actions further include presenting options for a plurality of predefined navigation tasks, including the current navigation task, receiving a user selection of a navigation task from the plurality of predefined navigation tasks, and determining the current navigation task based on the user selection.

[0105] In another aspect, the present disclosure provides a computer program product having computer-executable instructions tangibly stored in a computer storage medium that, when executed by an apparatus, causes the apparatus to perform actions including: generating, for a set of user interface elements, a set of tokens that individually represent the set of user interface elements, the set of user interface elements including at least one or more user interface elements in a currently presented user interface; converting the set of tokens into individual feature representations of the set of user interface elements using at least certain information corresponding to a current navigation task; determining a target element for the current navigation task from the one or more interface elements based on the feature representations; and performing an action associated with the target element.

[0106] In some example implementations, the set of user interface elements further includes a user interface element that serves as a history target element in the history interaction.

[0107] In some example implementations, converting the set of tokens into individual feature representations includes converting the set of tokens into at least one set of converted tokens using shared information and specific information for a plurality of predefined navigation tasks, where the plurality of predefined navigation tasks includes a current navigation task, and determining individual feature representations of the set of user interface elements based on the at least one set of converted tokens.

[0108] In some example implementations, converting the set of tokens into at least one set of converted tokens includes converting the set of tokens into a set of intermediate tokens using a first shared information item in the shared information, and converting the set of intermediate tokens into a converted set of tokens of the at least one set of converted tokens using a first specific information item corresponding to the current navigation task in the specific information.

[0109] In some example implementations, at least one set of transformed tokens includes a first set of transformed tokens and a second set of transformed tokens corresponding to the set of user interface elements, and determining individual feature representations of the set of user interface elements includes obtaining a relative relationship between the first user interface element and a second user interface element of the user interface elements; determining attention information based on the first set of transformed tokens, the second set of transformed tokens, and the relative relationship, wherein the attention information indicates a correlation between the first user interface element and the second user interface element; transforming the set of tokens into a third set of transformed tokens corresponding to the set of user interface elements using a second shared information item in the shared information; and determining the feature representation by weighting the third set of transformed tokens based on the attention information.

[0110] In some example implementations, the relative relationship includes at least one of a relative encoded position of the first user interface element and the second user interface element or a relative spatial position of the first user interface element and the second user interface element obtained from user interface metadata.

[0111] In some example implementations, determining the target element includes transforming the feature representation using third shared information in the shared information and a third specific information item in the specific information that corresponds to the current navigation task, and determining a probability that one or more user interface elements are the target element based on the transformed feature representation.

[0112] In some example implementations, generating the set of tokens includes generating a token representing a given user interface element of the set of user interface elements based on at least one of the following items related to the given user interface element: a type of the given user interface element, a text description of the given user interface element, a position of the given user interface element within the user interface in which it is located, or a time of occurrence of the given user interface element relative to the current user interface.

[0113] In some example implementations, the actions further include presenting options for a plurality of predefined navigation tasks, including the current navigation task, receiving a user selection of a navigation task from the plurality of predefined navigation tasks, and determining the current navigation task based on the user selection.

[0114] In another aspect, the present disclosure provides a computer-readable medium having stored thereon computer-executable instructions that, when executed by an apparatus, cause the apparatus to perform one or more example implementations of the methods of the foregoing aspects. The functionality described herein may be performed at least in part by one or more hardware logic units. For example, and without limitation, example types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), etc.

[0115] Program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or another programmable data processing device such that, when executed by the processor or controller, the functions / acts specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the device as a separate software package, partially on the device, partially on the device and partially on a remote device, or entirely on a remote device or server.

[0116] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in combination with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of a machine-readable storage medium would include one or more line-based electrical connections, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EEPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0117] Additionally, while operations are described in a particular order, it should be understood that such operations need not be performed in the particular order or sequential order shown, or that all of the illustrated operations must be performed to achieve desirable results. Under certain circumstances, multitasking and parallel processing may be beneficial. Similarly, while the above description includes several specific implementation details, these should not be construed as limiting the scope of the present disclosure. Furthermore, certain features that are described in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation may also be implemented in multiple implementations individually or in any suitable subcombination.

[0118] Although the subject matter has been described in terms specific to structural features and / or methodological actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely examples of implementing the claims.

Claims

1. 1. A computer-implemented method comprising: generating, for a set of user interface elements, a set of tokens that individually represent the set of user interface elements, the set of user interface elements including at least one or more user interface elements in the current user interface being presented; converting the set of tokens into individual feature representations of the set of user interface elements using at least certain information corresponding to a current navigation task; determining a target element for the current navigation task from the one or more user interface elements based on the feature representation; and performing an action associated with the target element; A method comprising:

2. The method of claim 1 , wherein the set of user interface elements further comprises a user interface element that serves as a history target element in a history interaction.

3. Transforming the set of tokens into the individual feature representations includes: transforming the set of tokens into at least one set of transformed tokens using shared information and the specific information for a plurality of predefined navigation tasks, the plurality of predefined navigation tasks including the current navigation; determining the individual feature representations of the set of user interface elements based on at least one set of the transformed tokens; 2. The method of claim 1, comprising:

4. Transforming the set of tokens into the at least one set of transformed tokens includes: transforming the set of tokens into a set of intermediate tokens using a first shared information item in the shared information; converting the set of intermediate tokens into a set of converted tokens of the at least one set of converted tokens using a first specific information item within the specific information that corresponds to the current navigation task; 4. The method of claim 3, comprising:

5. the at least one set of transformed tokens includes a first set of transformed tokens and a second set of transformed tokens corresponding to the set of user interface elements; and Determining the individual feature representations of the set of user interface elements includes: obtaining a relative relationship between a first user interface element and a second user interface element of the set of user interface elements; determining attention information based on the first set of transformed tokens, the second set of transformed tokens, and the relative relationship, the attention information informing a correlation between the first user interface element and the second user interface element; transforming the set of tokens into a third set of transformed tokens corresponding to the set of user interface elements using a second shared information item in the shared information; determining the feature representation by weighting the third set of transformed tokens based on the attention information; and 4. The method of claim 3, comprising:

6. The relative relationship is: the relative encoded positions of the first and second user interface elements obtained from user interface metadata; or the relative spatial position of the first user interface element and the second user interface element; The method of claim 5 , comprising at least one of:

7. Determining the target element comprises: transforming the feature representation using a third shared information item in the shared information and a third specific information item in the specific information that corresponds to the current navigation task; determining a probability that the one or more user interface elements are the target element based on the transformed feature representation; 4. The method of claim 3, comprising:

8. generating the set of tokens comprises: the type of the given user interface element; a textual description of the given user interface element; the position of the given user interface element within the user interface in which it is located; or a time of occurrence of the given interface element relative to the current user interface; 2. The method of claim 1, further comprising generating a token representing a given user interface element of the set of user interface elements based on at least one of the following terms for the given user interface element:

9. presenting options for a plurality of predefined navigation tasks, including the current navigation task; receiving a user selection of a navigation task from the plurality of predefined navigation tasks; determining the current navigation task based on the user selection; The method of claim 1 further comprising:

10. a processor; a memory coupled to and having instructions stored thereon; wherein the instructions, when executed by the processor, cause the electronic device to perform actions, the actions comprising: generating, for a set of user interface elements, a set of tokens that individually represent the set of user interface elements, the set of user interface elements including at least one or more user interface elements in the current user interface being presented; converting the set of tokens into individual feature representations of the set of user interface elements using at least certain information corresponding to a current navigation task; determining a target element for the current navigation task from the one or more user interface elements based on the feature representation; and performing an action associated with the target element; An electronic device comprising:

11. The electronic device of claim 10 , wherein the user interface set further comprises a user interface element that functions as a historical target element in a historical interaction.

12. Transforming the set of tokens into the individual feature representations includes: transforming the set of tokens into at least one set of transformed tokens using shared information and the specific information for a plurality of predefined navigation tasks, the plurality of predefined navigation tasks including the current navigation task; determining the individual feature representations of the set of user interface elements based on at least one set of the transformed tokens; The electronic device of claim 10, comprising:

13. Transforming the set of tokens into the at least one set of transformed tokens includes: transforming the set of tokens into a set of intermediate tokens using a first shared information item in the shared information; converting the set of intermediate tokens into a set of converted tokens of the at least one set of converted tokens using a first specific information item within the specific information that corresponds to the current navigation task; 13. The electronic device of claim 12, comprising:

14. The at least one set of transformed tokens includes a first set of transformed tokens and a second set of transformed tokens corresponding to the set of user interface elements, and determining the individual feature representations of the set of user interface elements includes: obtaining a relative relationship between a first user interface element and a second user interface element of the set of user interface elements; determining attention information based on the first set of transformed tokens, the second set of transformed tokens, and the relative relationship, the attention information informing a correlation between the first user interface element and the second user interface element; transforming the set of tokens into a third set of transformed tokens corresponding to the set of user interface elements using second shared information in the shared information; determining the feature representation by weighting the third set of transformed tokens based on the attention information; and 13. The electronic device of claim 12, comprising:

15. 1. A computer program product having computer-executable instructions tangibly stored in a computer storage medium and that, when executed by a device, cause the device to perform actions, the actions including: generating, for a set of user interface elements, a set of tokens that individually represent the set of user interface elements, the set of user interface elements including at least one or more user interface elements in the current user interface being presented; converting the set of tokens into individual feature representations of the set of user interface elements using at least certain information corresponding to a current navigation task; determining a target element for the current navigation task from the one or more user interface elements based on the feature representation; and performing an action associated with the target element; 1. A computer program product comprising: