Task execution method and apparatus, electronic device, medium, and program product
Patent Information
- Application Number
- CN202510376849.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]相关技术中,基于AI模型进行任务执行的场景中,任务执行失败的可能性较大
[0024]本公开的实施例提供的技术方案可以包括以下有益效果:由于目标模型是基于任务训练集、界面到达任务训练集以及界面操作任务训练集训练得到。因此,通过基于任务训练集进行训练模型,能够使训练完成的目标模型具有能够执行目标任务的能力,并且训练过程中加入界面到达任务训练集以及界面操作任务训练集,能够使训练完成的目标模型在界面到达能力以及界面操作能力得到增强,使得训练完成的目标模型具有更好的泛化能力。由此,在任务执行过程中,基于训练完成的目标模型执行任务,能够提高任务完成的成功率。
Smart Images

Figure CN122837686A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of equipment control, and more particularly to task execution methods, apparatus, electronic equipment, media, and program products. Background Technology
[0002] With the development of artificial intelligence (AI) technology, AI technology is being widely used in more and more scenarios to perform corresponding tasks.
[0003] In related technologies, the possibility of task execution failure is relatively high in scenarios where tasks are executed based on AI models. Summary of the Invention
[0004] To overcome the problems existing in related technologies, this disclosure provides a task execution method, apparatus, electronic device, medium, and program product.
[0005] According to a first aspect of some embodiments of this disclosure, a task execution method is provided, comprising: determining a target task to be executed, the target task being a task of controlling an electronic device to perform a target operation on a target interface, and the target interface being an interface after the electronic device is triggered to sequentially display m interfaces from the current interface during the execution of the target task, where m is a positive integer; determining, based on a target model, the i-th interface currently triggered and displayed during the execution of the target task, and determining an interface operation subtask to be executed on the i-th interface, the target model being used to determine the interface reached by the execution of the target task and the interface operation subtask to be executed on the reached interface; if the i-th interface is not the target interface, and the interface is not the target interface, the target task is determined to be a target operation. If the operation performed by the interface operation subtask is not the target operation, then the interface operation subtask is executed to display the (i+1)th interface, and the interface operation subtask executed in the (i+1)th interface is determined, until it is determined that the electronic device displays the target interface, and the operation performed by the interface operation subtask is the target operation; if the i-th interface is the target interface, and the operation performed by the interface operation subtask is the target operation, then the target task is completed; wherein, the target model is trained based on the task training set, the interface arrival task training set, and the interface operation task training set, and the interface arrival task training set and the interface operation task training set are determined based on the task training set.
[0006] In some implementations, the target model is trained as follows: An initial model is generated by inputting a task training set, an interface arrival task training set, an interface operation task training set, and the kth interface currently displayed on the electronic device; the interface operation to be performed by the electronic device on the kth interface is determined, where k is a positive integer; based on the corresponding interface operation performed on the kth interface, preference data pairs are determined, and a loss function corresponding to the initial model is determined based on the preference data pairs, where different interface operations on the kth interface correspond to different scores, and the preference data pairs include interface operations corresponding to different scores; the initial model is adjusted based on the loss function to obtain the target model.
[0007] In some implementations, before determining the interface operation that the electronic device needs to perform on the k-th interface, the method further includes: determining action attribute information corresponding to the k-th interface, wherein the action attribute information characterizes the effective response range corresponding to the element of the k-th interface used to respond to the interface operation; selecting a target response range within the effective response range; generating an action space corresponding to the k-th interface based on the effective response range; and inputting the action space into the initial model to determine the interface operation that the electronic device needs to perform on the k-th interface.
[0008] In some embodiments, the task training set includes: a training task, and a first interface stream displayed by an electronic device corresponding to the completion of the training task. The first interface stream includes interface images and first interface operations corresponding to the interface images, wherein adjacent interface images in the first interface stream correspond to at least one first interface operation. The interface arrival task training set and the interface operation task training set are determined based on the task training set in the following manner: based on each interface image and the first interface operation corresponding to each interface image, descriptive information corresponding to each first interface operation is determined; based on the descriptive information, a second interface stream and a third interface stream are determined in the first interface stream, wherein the second interface stream is the interface stream corresponding to the execution of the interface arrival sub-task, and the third interface stream is the interface stream used to execute the interface operation sub-task; the interface arrival task training set is constructed based on the second interface stream, and the interface operation task training set is constructed based on the third interface stream.
[0009] In some implementations, determining a second interface flow in the first interface flow based on the description information includes: determining a first interface image in the first interface flow based on the description information, wherein the first interface image is the interface image that first appears in the first interface flow; determining a second interface image and an interface operation corresponding to the second interface image in the first interface flow, wherein the second interface image is an interface image in the first interface flow preceding the first interface image; and determining the second interface flow based on the second interface image and the interface operation corresponding to the second interface image.
[0010] In some implementations, determining the first interface image in the first interface stream based on the description information includes: in response to the description information including first description information, determining the interface image corresponding to the first description information as the second interface image, wherein the first description information is used to describe the name of the interface image, and the first description information and the interface image have a one-to-one correspondence; in response to the description information including second description information, determining the next interface image corresponding to the description information as the second interface image, wherein the second description information is used to describe the name corresponding to an element in the interface image, and the name corresponding to the element is the name that appears for the first time in the description information.
[0011] In some implementations, determining the third interface flow in the first interface flow based on the description information includes: selecting a second interface operation from the first interface operations in the first interface flow, wherein the second interface operation is an interface operation for constructing a training set of interface operation subtasks; determining the operation type corresponding to the second interface operation based on the description information corresponding to the second interface operation and the second interface operation; and determining the third interface flow in the first interface flow based on the operation type.
[0012] In some implementations, determining a third interface flow in the first interface flow based on the operation type includes: in response to the second interface operation being an interface operation of a first operation type, determining a third interface image in the first interface flow, and determining the interface flow preceding the third interface image as the third interface flow, wherein the interface operation of the first operation type includes a swipe operation and / or a keystroke operation, and the third interface image is the interface image in the first interface flow that is adjacent to the second interface operation after it; in response to the second interface operation being an interface operation of a second operation type, and the similarity between the third interface image and the fourth interface image satisfying a similarity requirement, determining the third interface flow based on the third interface image, the second interface operation, and the fourth interface image, wherein the interface operation of the second operation type includes a click operation, the third interface image is the interface image in the first interface flow that is adjacent to the second interface operation after it, and the fourth interface image is the interface image in the first interface flow that is adjacent to the second interface operation before it.
[0013] According to a second aspect of some embodiments of this disclosure, a task execution apparatus is provided, comprising: an acquisition unit, configured to determine a target task to be executed, the target task being a task of controlling an electronic device to perform a target operation on a target interface, and the target interface being the interface after which the electronic device is triggered to sequentially display m interfaces from the current interface during the execution of the target task, where m is a positive integer; and a processing unit, configured to, based on a target model, determine the i-th interface currently triggered and displayed during the execution of the target task, and determine an interface operation subtask to be executed on the i-th interface; and, if the i-th interface is not the target interface, and the operation corresponding to the interface operation subtask is not the target operation, the processing unit is configured to execute the interface operation subtask and display... The processing unit determines the (i+1)th interface and the interface operation subtask to be executed on the (i+1)th interface, until it is determined that the electronic device displays the target interface, and the operation to be executed corresponding to the interface operation subtask is the target operation; when the i-th interface is the target interface and the operation to be executed corresponding to the interface operation subtask is the target operation, the processing unit is used to complete the target task; wherein, the target model is used to determine the interface reached by executing the target task and the interface operation subtask to be executed on the reached interface, the target model is trained based on the task training set, the interface arrival task training set, and the interface operation task training set, and the interface arrival task training set and the interface operation task training set are determined based on the task training set.
[0014] In some embodiments, the device further includes a training unit for training the target model. The training unit trains the target model in the following manner: inputting a task training set, an interface arrival task training set, an interface operation task training set, and the kth interface currently displayed by the electronic device into an initial model; determining the interface operation that the electronic device needs to perform on the kth interface, where k is a positive integer; determining preference data pairs based on the corresponding interface operations performed on the kth interface; and determining a loss function corresponding to the initial model based on the preference data pairs, where different interface operations on the kth interface correspond to different scores, and the preference data pairs include interface operations corresponding to different scores; and adjusting the initial model based on the loss function to obtain the target model.
[0015] In some implementations, the training unit is further configured to: determine the action attribute information corresponding to the k-th interface, wherein the action attribute information characterizes the effective response range corresponding to the element of the k-th interface used to respond to interface operations; select a target response range within the effective response range; generate an action space corresponding to the k-th interface based on the target response range; and input the action space into the initial model to determine the interface operations that the electronic device needs to perform on the k-th interface.
[0016] In some embodiments, the task training set includes: a training task, and a first interface stream displayed by an electronic device corresponding to the completion of the training task. The first interface stream includes interface images and first interface operations corresponding to the interface images, wherein adjacent interface images in the first interface stream correspond to at least one first interface operation. The interface arrival task training set and the interface operation task training set are determined based on the task training set in the following manner: based on each interface image and the first interface operation corresponding to each interface image, descriptive information corresponding to each first interface operation is determined; based on the descriptive information, a second interface stream and a third interface stream are determined in the first interface stream, wherein the second interface stream is the interface stream corresponding to the execution of the interface arrival sub-task, and the third interface stream is the interface stream used to execute the interface operation sub-task; the interface arrival task training set is constructed based on the second interface stream, and the interface operation task training set is constructed based on the third interface stream.
[0017] In some implementations, the second interface flow is determined as follows: based on the description information, a first interface image is determined in the first interface flow, wherein the first interface image is the interface image that appears for the first time in the first interface flow; in the first interface flow, a second interface image and the interface operation corresponding to the second interface image are determined, wherein the second interface image is the interface image in the first interface flow that precedes the first interface image; the second interface flow is determined based on the second interface image and the interface operation corresponding to the second interface image.
[0018] In some implementations, the first interface image is determined as follows: when the description information includes first description information, the interface image corresponding to the first description information is determined as the second interface image, wherein the first description information is used to describe the name of the interface image, and the first description information and the interface image have a one-to-one correspondence; when the description information includes second description information, the next interface image corresponding to the description information is determined as the second interface image, wherein the second description information is used to describe the name corresponding to an element in the interface image, and the name corresponding to the element is the name that appears for the first time in the description information.
[0019] In some implementations, the third interface flow is determined as follows: a second interface operation is selected from the first interface operations of the first interface flow, wherein the second interface operation is an interface operation used to construct a training set of interface operation subtasks; based on the description information corresponding to the second interface operation and the second interface operation, the operation type corresponding to the second interface operation is determined; based on the operation type, the third interface flow is determined in the first interface flow.
[0020] In some implementations, the third interface flow is determined as follows: When the second interface operation is an interface operation of a first operation type, a third interface image is determined in the first interface flow, and the interface flow preceding the third interface image is determined as the third interface flow. The first operation type includes swiping and / or typing operations, and the third interface image is the interface image in the first interface flow adjacent to the second interface operation. When the second interface operation is an interface operation of a second operation type, and the similarity between the third interface image and the fourth interface image meets the similarity requirement, the third interface flow is determined based on the third interface image, the second interface operation, and the fourth interface image. The second operation type includes clicking operations, the third interface image is the interface image in the first interface flow adjacent to the second interface operation, and the fourth interface image is the interface image in the first interface flow adjacent to the second interface operation.
[0021] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the task execution method described in the first aspect or any embodiment of the first aspect.
[0022] According to a fourth aspect of the present disclosure, a storage medium is provided, the storage medium storing instructions that, when executed by a processor, enable the processor to perform the task execution method described in the first aspect or any embodiment of the first aspect.
[0023] According to a fifth aspect of the present disclosure, a computer program product is provided, the computer program product including a computer program, which, when executed by a processor, implements the task execution method described in the first aspect or any embodiment of the first aspect.
[0024] The technical solutions provided by the embodiments of this disclosure can include the following beneficial effects: Since the target model is trained based on a task training set, an interface arrival task training set, and an interface operation task training set, training the model based on the task training set enables the trained target model to perform the target task. Furthermore, incorporating the interface arrival task training set and the interface operation task training set during training enhances the target model's interface arrival and operation capabilities, resulting in better generalization ability. Therefore, executing tasks based on the trained target model during task execution improves the success rate of task completion.
[0025] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0026] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0027] Figure 1A This is a schematic diagram illustrating an application scenario of a task execution method in the related art, based on some embodiments of the present disclosure.
[0028] Figure 1B This is a schematic diagram of a task execution method in the related art, based on some embodiments of the present disclosure.
[0029] Figure 2A This is a flowchart illustrating a task method according to some embodiments of the present disclosure.
[0030] Figures 2B to 2J These are schematic diagrams one to nine illustrating a task execution method according to some embodiments of this disclosure.
[0031] Figure 3A This is a flowchart illustrating a target model training method according to some embodiments of the present disclosure.
[0032] Figure 3B This is a schematic diagram illustrating the architecture of a target model training method according to some embodiments of the present disclosure.
[0033] Figure 3C This is a schematic diagram of a scenario of the interface flow corresponding to the execution of a target task, based on some embodiments of this disclosure.
[0034] Figure 4 This is a flowchart illustrating a motion space generation method according to some embodiments of the present disclosure.
[0035] Figure 5 This is a flowchart illustrating a method for determining the interface arrival task training set and the interface operation task training set according to some embodiments of the present disclosure.
[0036] Figure 6 This is a flowchart illustrating a method for determining a second interface flow in a first interface flow according to some embodiments of the present disclosure.
[0037] Figure 7 This is a flowchart illustrating a method for determining a first interface image in a first interface stream according to some embodiments of the present disclosure.
[0038] Figure 8This is a flowchart illustrating a method for determining a third interface flow in a first interface flow according to some embodiments of the present disclosure.
[0039] Figure 9 This is a flowchart illustrating a method for determining a third interface flow according to some embodiments of the present disclosure.
[0040] Figure 10 This is a block diagram of a task execution device according to some embodiments of the present disclosure.
[0041] Figure 11 This is a block diagram of an apparatus for task execution according to some embodiments of the present disclosure.
[0042] Figure 12 This is a block diagram 2 illustrating an apparatus for task execution according to some embodiments disclosed in a book. Detailed Implementation
[0043] Some embodiments of this disclosure will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. Various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but can be changed as will become apparent upon understanding this disclosure, except for operations that must be performed in a particular order. Furthermore, for clarity and brevity, descriptions of features known in the art may be omitted.
[0044] With the development of artificial intelligence (AI) technology, AI technology is being widely used in more and more scenarios to perform corresponding tasks.
[0045] Taking electronic device control applications as an example, in related technology applications, users can input task descriptions into AI models, enabling the AI models to control corresponding devices based on these descriptions, thereby completing the user's task. For example, this could involve multimodal agents that automatically execute GUI (Graphical User Interface) interaction tasks. To facilitate understanding, the following will illustrate this further. Figures 1A to 1B The task execution methods mentioned in the related technologies are described exemplarily.
[0046] Figure 1A This is a schematic diagram illustrating an application scenario of a task execution method in the related art, based on some embodiments of this disclosure. For example... Figure 1AAs shown, there is a communication relationship between the AI model and the electronic device (as indicated by the bidirectional arrows in the figure), and the AI model can obtain the task description input by the user (e.g., Figure 1A The system can perform user-input tasks by controlling electronic devices based on task descriptions and communication relationships with them.
[0047] It should be noted that "AI model" can be understood as an electronic device that deploys an AI model, or a module that deploys an AI model. In some scenarios, AI models can be deployed in electronic devices that need to be controlled (e.g., Figure 1A AI models in can be deployed in Figure 1A In some scenarios, the "AI model" can also be deployed in electronic devices that are different from the electronic devices that need to be controlled (e.g., in electronic devices). Figure 1A The AI models in this context can be deployed in other electronic devices, which are distinct from [other devices]. Figure 1A (The electronic device shown). The scenarios involved in the embodiments of this disclosure may include at least one of the two scenarios described above.
[0048] For AI models to control electronic devices to perform tasks, such as Figure 1B As shown, Figure 1B This is a schematic diagram of a task execution method in the related art, based on some embodiments of this disclosure. For example... Figure 1B As shown, continuing Figure 1A In the scenario (where, the relevant content for users and AI models is in Figure 1B The same may also exist in [the context of the previous sentence]. To avoid repetition, please refer to the relevant descriptions. Figure 1A (See related examples). When an AI model receives a task description input by the user, it outputs operation commands to control an electronic device based on the task description. These commands control the electronic device, thus completing the user's task. For example, based on the acquired task description, the AI model outputs operation commands to the electronic device (such as touching a specific icon on the display interface). The electronic device responds to the operation commands, performing a touch operation on the specified icon (e.g., the icon within the circled area in the diagram), thereby causing the electronic device's display interface to switch to the target display interface (e.g., the icon within the circled area in the diagram). Figure 1B The interface displays the weather for region A to complete the user's task.
[0049] Understandably, in the above scenario, the AI model can complete the user's task by outputting operation instructions for single-step operations on the electronic device (such as operation instructions for touching a specified icon on the display interface).
[0050] It is understandable that, in order to improve the accuracy of AI models in single-step control of electronic devices, related technologies may train AI models based on corresponding training sets (e.g., training sets specifically designed for controlling electronic devices to perform single-step operations) and / or training methods (e.g., building a model training framework based on a greedy algorithm). This allows the trained AI models in related technologies to focus more on the content of the current display interface of the electronic device during task execution (i.e., the application phase) and control the electronic device to perform tasks based on the content of the current display interface.
[0051] However, in other scenarios, the task descriptions that users input into the AI model may be more complex, requiring the AI model to perform multi-step control of the electronic device (for example, the AI model outputs operation instructions to perform multi-step operations on the electronic device) in order to control the electronic device to complete the user's task.
[0052] As can be seen, in the above scenarios, the AI models in the relevant technologies focus more on the content of the current display interface of the electronic device during task execution and control the electronic device to perform tasks based on the content of the current display interface, without paying attention to the overall impact of the electronic device's operations on the current display interface on the task execution process (e.g., the impact of the task failing due to getting stuck in a local optimum). This makes the task execution methods in the relevant technologies unable to perform well in scenarios where more complex tasks are performed, thereby increasing the possibility of task execution failure and reducing the user experience.
[0053] In view of this, embodiments of this disclosure propose a task execution method. Based on a target model, a target task, and the currently displayed interface of the electronic device (e.g., the i-th interface), the method determines the operations (e.g., processing operations corresponding to the i-th interface) that the electronic device needs to perform on the currently displayed interface (e.g., the i-th interface). It also determines whether the interface after performing the processing operation corresponding to the i-th interface (e.g., the (i+1)-th interface) is an interface used to execute the target task. If the (i+1)-th interface is an interface used to execute the target task, then the corresponding operation is performed on the (i+1)-th interface to complete the target task. If the (i+1)th interface is not the interface used to perform the target task, then the (i+1)th interface is taken as the currently displayed interface of the electronic device, and the processing operation corresponding to the currently displayed interface is determined again in a loop until the target interface is reached, and the corresponding operation is completed on the target interface to complete the target task. The target model is trained based on a first training set, a second training set, and a third training set. The first training set is used to enable the trained target model to have the logic to execute the complete target task, the second training set is used to enhance the target model's ability to reach the target interface, and the third training set is used to enhance the target model's interface operation capabilities. In this embodiment, by using the target model based on the content displayed on the current interface of the electronic device, the operation that the electronic device needs to perform on the currently displayed interface is determined. After the electronic device performs the operation, the interface displayed by the electronic device after the operation is further determined and the corresponding operation is performed until the task ends. Since the target model is a model with enhanced capabilities based on the second and third training sets, executing the target task based on the target model can reduce the possibility of task failure due to the AI model getting stuck in a local optimum during task execution, thus improving the user experience.
[0054] It should be noted that the task execution method provided in this disclosure can be applied to electronic devices. Electronic devices may include, for example, terminals or servers. Terminals include, for example, mobile phones, wearable devices, IoT devices, automobiles with communication capabilities, smart cars, tablets, computers with wireless transceiver capabilities, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminal devices in industrial control, wireless terminal devices in self-driving, wireless terminal devices in remote medical surgery, wireless terminal devices in smart grids, wireless terminal devices in transportation safety, wireless terminal devices in smart cities, and wireless terminal devices in smart homes, but are not limited thereto. Servers may include, but are not limited to, independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0055] The embodiments disclosed herein can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.
[0056] It should be noted that all actions involving the acquisition of signals, information, or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with authorization from the owner of the relevant device.
[0057] For ease of understanding, some technical terms involved in the embodiments of this disclosure will be explained by way of example below:
[0058] A graphical user interface (GUI) is an interface that uses graphical elements to enable interaction between a user and a computer system. Users can manipulate these graphical elements using input devices such as a mouse and keyboard, or they can use gestures, voice commands, or other non-input device methods to complete various tasks, such as opening files, running programs, and adjusting settings. In this embodiment, the GUI interface may be simply referred to as the "interface."
[0059] The components of a GUI typically include at least one of the following: window, icon, menu, button, text box, or scroll bar.
[0060] In the relevant embodiments of this disclosure, a GUI image can be understood, for example, as a screenshot of the GUI. A GUI stream can be understood, for example, as an image stream comprising a series of GUI images, which are associated based on processing operations performed by the electronic device.
[0061] Specialized Fine-Tuning: This process involves further training the pre-trained model so that it can better perform specific tasks or follow specific instructions. During the SFT (or SFT fine-tuning) phase, the model is trained on one or more very specific tasks, which helps improve the model's accuracy and adaptability on those tasks.
[0062] Direct Preference Optimization (DPO) is primarily used in Reinforcement Learning (RL) to optimize an agent's behavioral policies. Unlike traditional reinforcement learning methods, DPO does not rely on a reward function; instead, it guides the agent's policy learning by learning a preference model. In DPO, the agent attempts to implement different policies and gathers preference information about these policies through some means (such as consulting human experts or using historical data). Based on the collected preference information, the agent updates its preference model. The updated preference model is then used to evaluate and adjust the policy, making the agent's behavior more aligned with its preferences.
[0063] Action space refers to the set of all possible actions an agent can take in a given environment. It defines the rules and constraints that govern an agent's interaction within a specific environment. The agent's goal is to find the optimal strategy within its action space to maximize cumulative reward.
[0064] The embodiments described in the following examples of this disclosure are not representative of all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0065] The task execution methods provided in some embodiments of this disclosure can be applied, for example, to electronic device control scenarios.
[0066] Figure 2A This is a flowchart illustrating a task method according to some embodiments of the present disclosure, such as... Figure 2A As shown, the task method is used for electronic devices (e.g., terminals or servers) and includes the following steps.
[0067] In step S11, the target task to be executed is determined.
[0068] In step S12, based on the target model, the i-th interface currently triggered and displayed during the execution of the target task is determined, and the interface operation sub-task executed in the i-th interface is determined.
[0069] In step S13-1, if the i-th interface is not the target interface and the operation corresponding to the interface operation subtask is not the target operation, then the interface operation subtask is executed, the (i+1)-th interface is displayed, and the interface operation subtask executed in the (i+1)-th interface is determined, until it is determined that the electronic device displays the target interface and the operation corresponding to the interface operation subtask is the target operation.
[0070] In step S13-2, if the i-th interface is the target interface and the operation corresponding to the interface operation subtask is the target operation, then the target task is completed.
[0071] The target task is to control an electronic device to perform a target operation on a target interface, and the target interface is the interface after the electronic device sequentially displays m interfaces from the current interface during the execution of the target task, where m is a positive integer. The target model is used to determine the interface reached when the target task is executed and the interface operation sub-tasks performed on the reached interface. The target model is trained based on the task training set, the interface arrival task training set, and the interface operation task training set, which are subsets of the task training set.
[0072] In this embodiment, the target model is trained based on task training sets, interface arrival task training sets, and interface operation task training sets. Therefore, by training the model based on the task training sets, the trained target model can be equipped with the ability to perform the target task. Furthermore, the inclusion of the interface arrival task training sets and interface operation task training sets during training enhances the target model's interface arrival and operation capabilities, resulting in better generalization ability. Consequently, during task execution, the target model exhibits better performance, reducing the possibility of the model getting trapped in local optima and increasing the success rate of task completion.
[0073] In some embodiments, determining the i-th interface currently triggered and displayed during the execution of the target task can be understood as determining the interface image corresponding to the i-th interface, as well as the action attribute information of the elements in the i-th interface.
[0074] In this context, elements within the interface can be understood as such as icons, text boxes, scroll bars, etc. The action attributes of these elements can be understood as the type of interface operation they can respond to (e.g., touch input or swipe input) and their effective response range (e.g., maximum effective touch range or maximum swipe range).
[0075] Understandably, by determining the image corresponding to the i-th interface, the differences between the i-th interface and the target interface in the image can be determined, thereby confirming whether the i-th interface is the target interface. By determining the attribute information of the elements, the electronic device can operate more accurately during interface operations on the i-th interface, reducing the possibility of invalid or erroneous operations.
[0076] To facilitate understanding, the following will be explained through... Figures 2B to 2J Regarding the above Figure 2A The proposed task execution method is described and illustrated by example.
[0077] Figure 2B This is a schematic diagram illustrating a task execution method according to some embodiments of this disclosure. For example... Figure 2B As shown, the electronic device to be controlled is, for example... Figure 2B The first page of the electronic device shown could be, for example, the page currently displayed on the electronic device. For Figure 2B The "target model" can be understood as, for example, an electronic device on which the target model is deployed, or a module on which the target model is deployed. In some scenarios, the target model can be deployed in an electronic device that needs to be controlled (e.g., Figure 2B AI models in can be deployed in Figure 2BIn some scenarios, the "target model" may also be deployed in an electronic device that is different from the electronic device that needs to be controlled (e.g., in an electronic device). Figure 2B The target model in the model can be deployed in other electronic devices, which are distinguished from other electronic devices. Figure 2B (The electronic device shown). For the target task, for example, it can make... Figure 2B The description states: "Add product A2 to my shopping cart." Therefore, the target interface can be understood as "the interface that allows adding product A2 to the shopping cart," and the target operation can be understood as the electronic device adding product A2 to its shopping cart from the interface that allows adding product A2 to the shopping cart.
[0078] Understandably, to complete a target task, electronic devices may require different interface operations depending on the currently displayed interface. For example, Figure 2C as well as Figure 2D The interface displayed by an electronic device can be understood as the different interfaces currently displayed by the electronic device. Therefore, in order to complete the target task, different interface operations need to be performed on different interfaces.
[0079] To facilitate understanding, the following will be based on Figure 2C The following example illustrates how the interface displayed by the electronic device is "the i-th interface currently triggered and displayed during the execution of the target task by the electronic device".
[0080] Figure 2C This is a second schematic diagram illustrating a task execution method according to some embodiments of this disclosure. The interface currently displayed by the electronic device is a "desktop" interface, which includes icons for multiple applications. Before the electronic device displays the target interface, a series of operations may need to be performed, such as searching for the corresponding application and launching the corresponding application. Therefore, in this case, based on the target model and the "i-th interface" currently displayed by the electronic device, the interface operation to be performed can be determined as: launching the corresponding application (e.g., Figure 2C (The icons within the dashed circle represent the applications, etc.)
[0081] In some implementations, for an electronic device to perform interface operations, such as launching a corresponding application, the target model can generate control commands for controlling the electronic device based on the target task and the interface currently displayed on the electronic device. In response to receiving the control commands, the electronic device launches the application based on the control commands.
[0082] It should be noted that in this embodiment, the process of the electronic device performing interface operations can be implemented by referring to this implementation method, and will not be described in detail below.
[0083] It is also understandable that after performing corresponding interface operations on the currently displayed interface of the electronic device, the displayed interface of the electronic device may change (for example, the i-th interface may jump to the (i+1)-th interface, etc.).
[0084] like Figure 2D As shown, Figure 2D This is a schematic diagram of a task execution method according to some embodiments of this disclosure, shown in scenario three. (Continued) Figure 2C and related embodiments described, Figure 2D It can be understood as the interface displayed by the electronic device after the corresponding interface operation is performed on the "i-th interface". For example, the interface displayed by the electronic device after the application is launched can be referred to as the "i+1-th interface" of the electronic device in the process of performing the target task.
[0085] according to Figure 2D It is known that the "(i+1)th screen" is not the target screen corresponding to the target task; that is, the "(i+1)th screen" is not the screen that "can add product A2 to the shopping cart." Therefore, the target model still needs to perform corresponding screen operations on the "(i+1)th screen," such as looping. Figure 2C The action execution logic (excluding specific interface operations) is used to complete the target task.
[0086] For example, in some implementations, if the "(i+1)th interface" is not the target interface corresponding to the target task, the target model will use the "(i+1)th interface" as the currently displayed interface of the electronic device and perform interface operations based on the "(i+1)th interface" and the target model. For example, for Figure 2D In other words, the interface operation can be a touch control labeled "Product A".
[0087] Figure 2E This is a schematic diagram of a task execution method according to some embodiments of the present disclosure. Figure 4 , Figure 2F This is a schematic diagram of a task execution method according to some embodiments of the present disclosure. Figure 5 Continuing Figure 2D and related implementation methods, such as Figure 2E As shown, after the electronic device performs an interface operation on the "i+1th interface", the interface displayed by the electronic device changes again (for example, it can be understood as a display interface jump). For ease of understanding, we can... Figure 2E The interface is called the "i+2nd interface". According to... Figure 2EIt is clear that the "(i+2)th screen" is not the target screen corresponding to the objective task, meaning that the "(i+2)th screen" is not the screen that allows adding product A (item A2) to the shopping cart. Therefore, the objective model still needs to perform corresponding interface operations on the "(i+2)th screen," such as looping. Figure 2C The execution logic (excluding specific interface operations) is used to complete the target task. For the interface operations performed, for example, it could be... Figure 2F The touch-sensitive control labeled "Add to Cart" is shown (i.e., the control within the dashed circle).
[0088] Figure 2G This is a schematic diagram of a task execution method according to some embodiments of the present disclosure. Figure 6 , Figure 2H This is a schematic diagram of a task execution method according to some embodiments of the present disclosure. Figure 7 2I is a scenario illustration of a task execution method according to some embodiments of the present disclosure. Figure 8 . Figure 2J This is a schematic diagram of a task execution method according to some embodiments of the present disclosure. Figure 9 Continuing Figure 2F The relevant example describes how an electronic device arrives at a certain screen after performing the corresponding interface operation on the "i+2nd screen". Figure 2G The interface shown can be referred to as, for example, "the (i+3)th interface". According to... Figure 2G It can be seen that the "i+3rd screen" can be the target screen corresponding to the target task, that is, the "i+3rd screen" is the screen that "can add product A (item A2) to the shopping cart". Therefore, the target model can perform the corresponding target operation on the "i+3rd screen" to complete the target task. For example, it can be like this: Figure 2H As shown, touch the control labeled "A2 style" (i.e., the control within the dotted circle) to select the item to add to the shopping cart. And, as... Figure 2I As shown, you can touch the control marked "OK" (i.e., the control within the dashed circle) in the interface to complete the target operation. The completion of the target operation could be, for example, in... Figure 2I After performing the interface operation, confirm in the displayed interface, for example... Figure 2J The interface shown.
[0089] It is also understandable that, for some target tasks, electronic devices only need to reach the target interface. For example, continuing from the above... Figures 2B to 2J In related implementation methods, assuming the target task is "help me find product A", this target task may be... Figure 2EThe displayed interface is sufficient. Therefore, in this case, the target operation can be understood as the operation of displaying a specified interface on the display screen of an electronic device, rather than the interactive operations (e.g., touch) listed in the above examples.
[0090] It is understandable that, in the above Figures 2B to 2J In the relevant example descriptions, electronic devices can complete target tasks through interface operations. For example, in Figure 2D In the example, Figure 2D The display interface shown includes a search bar (such as the display area corresponding to "Search" in the figure). In some scenarios, electronic devices can also use the search bar to access other interfaces. For example, entering "Product A" in the search bar will lead to... Figure 2E The interface shown. For example, entering "Product A, Style A2" in the search bar will lead to... Figure 2H The interface shown is an example of a user interface. Regardless of the method used, all methods can be considered as ways to perform the target task. Therefore, during model training, appropriate reward methods can be set to enhance the model's adaptability to different scenarios during task execution, thereby improving the model's generalization ability.
[0091] Therefore, the target model can be trained in the following way, for example. Figure 3A This is a flowchart illustrating a target model training method according to some embodiments of this disclosure. Figure 3A As shown, the method includes the following steps.
[0092] In step S21, the task training set, the interface arrival task training set, the interface operation task training set, and the kth interface currently displayed by the electronic device are input into the initial model to determine the interface operation that the electronic device needs to perform on the kth interface.
[0093] In step S22, based on the corresponding interface operation performed in the k-th interface, the preference data pair is determined, and the loss function corresponding to the initial model is determined based on the preference data pair.
[0094] In step S23, the initial model is adjusted based on the loss function to obtain the target model.
[0095] Where k is a positive integer, and different interface operations in the k-th interface correspond to different scores. The preference data pairs include interface operations corresponding to different scores.
[0096] In this embodiment, the task training set, the interface arrival task training set, and the interface operation task training set are input into the initial model so that the initial model outputs the interface operation corresponding to the k-th interface. By setting corresponding scores for different interface operations, a clear reward signal can be set for the model output results. Furthermore, by setting preference data pairs based on the interface operations corresponding to different scores, the model can more accurately distinguish between suggested and discouraged interface operations when outputting diverse content. This reduces the likelihood of the model outputting a "disadvantaged interface operation," thereby improving the success rate of the trained target model in executing tasks.
[0097] To facilitate understanding, the following will be explained through... Figure 3B An exemplary illustration of the target model training architecture is provided. Figure 3B This is a schematic diagram illustrating the architecture of a target model training method according to some embodiments of this disclosure. Figure 3B As shown, the training architecture of the target model can be divided into, for example, fine-tuning architecture and reinforcement architecture.
[0098] For fine-tuning the architecture, the initial model selected can be, for example, a Large Language Model (LLM) base, which may only have basic text output capabilities. Therefore, in order to obtain a target model that can be used to perform the target task, it is necessary to fine-tune it based on the corresponding training set (e.g., the task training set, the interface accessing the task training set, and the operation of the task training set), the action space corresponding to the display interface, etc. (e.g., fine-tuning based on the SFT policy).
[0099] It is also understandable that, in the aforementioned task execution scenarios, the target model needs to perform inference in conjunction with the display interface corresponding to the electronic device. Therefore, during the target model training phase, relevant information about the display interface corresponding to the electronic device (such as the display interface image and its corresponding attribute information) and the corresponding training set (such as the task training set, the interface arrival task training set, and the interface operation task training set) need to be used as input to the model. It is also understandable that, since the model may output multiple rounds during the fine-tuning phase, historical output results and the training set can be input together into the initial model. For example, in the Nth round of input, based on the display interface of the (N-1)th round, the action history of the (N-1)th round, the action space corresponding to the display interface of the Nth round, and the training set (such as the initial input training set), the interface operation output of the Nth round is obtained, where N is a positive integer.
[0100] Therefore, the fine-tuning architecture can include, for example, a graphics encoder, an image adapter, and a fully connected layer, as shown in the figure. The display image of the electronic device's corresponding display interface can be input to the graphics encoder, then via the image adapter to the fully connected layer. This fully connected layer concatenates features with the input training set (e.g., the task training set, the interface arrival task training set, or the interface operation task training set). The concatenated features are then used as the input to the initial model to obtain the corresponding output result (e.g., the interface operation corresponding to the current display interface in the model's inference).
[0101] For reinforcement architectures, designs can be based on DPO strategies. For example, different reward scores can be assigned to different interface operations output by the initial model at the same time step, and preference data pairs can be constructed based on the interface operations corresponding to different reward scores at the same time step. Then, the corresponding loss function can be applied based on the preference data pairs.
[0102] The mechanism for assigning reward points can be implemented, for example, through the following example.
[0103] Understandably, based on Figures 2B to 2J As relevant examples show, during the execution of a target task by an electronic device, performing different interface operations on the same display screen may lead to different results in subsequent interface transitions. For example... Figure 3C For example, Figure 3C This is a schematic diagram illustrating a scenario of the interface flow corresponding to the execution of a target task, based on some embodiments of this disclosure. For example... Figure 3C As shown, P0, P1, P2, and P3 can be understood as the "optimal" (e.g., the optimal one listed in this example) interface flow corresponding to completing the target task. Figure 3C The solid arrows in the diagram represent the paths corresponding to the interface images, where P0 can be understood as the time step before the target model performs inference, and the interface image corresponding to the display interface of the electronic device. P0 indicates that the electronic device performs corresponding interface operations on the interface, and P1 indicates that after the electronic device performs corresponding interface operations on the interface, the corresponding interface image displayed on the electronic device is displayed. P1 indicates that the electronic device performs corresponding interface operations on the interface, and P2 indicates that after the electronic device performs corresponding interface operations on the interface in P1, the electronic device displays the corresponding interface image. P2 indicates that the electronic device performs corresponding interface operations on the interface, and P3 indicates the interface image displayed by the electronic device after performing corresponding interface operations on the interface.
[0104] It should also be noted that, for Figure 3CSubscript numbers represent the time steps of inference for the target model, where "0" indicates that inference has not started, "1" indicates the first time step, "2" indicates the second time step, and "3" indicates the third time step. Superscript numbers represent the corresponding interface operations performed. The same superscript number indicates that the same interface operation was performed, and different superscript numbers indicate that different interface operations were performed.
[0105] And for Figure 3C The other paths shown may affect the completion of the target task through the corresponding interface operations at each time step. For ease of understanding, the following examples (A1) to A4) will be used to illustrate this.
[0106] A1) For example, the path indicated by the solid arrow represents the "optimal" interface flow for completing the target task.
[0107] A2) For example, the path represented by the dashed arrow, although it corresponds to a higher number of interface operations (i.e., an increase in the part represented by "a") compared to the path represented by the solid arrow, can still... The interface returns (or is redirected) to the P2 interface, and then the process is performed within the P2 interface. The operation causes the electronic device to display the P3 interface to complete the target task.
[0108] A3) For example, the path indicated by the dotted-line arrow in the figure does not ultimately lead to the P3 interface. That is, the path indicated by the dotted-line arrow may have performed an incorrect interface operation, resulting in the inability to complete the target task.
[0109] A4) For example, the target model may determine an interface operation in the first time step that cannot be executed in the current interface. Figure 3C (omitted). For example, the interface may only contain controls triggered by a "swipe" action, but the target model determines that the electronic device needs to perform a "click" action on that interface.
[0110] It is understandable that there may be other similar paths in the above path representation. These other paths can also be classified into the cases A1) to A4) above, which will not be elaborated here.
[0111] In summary, four different levels of reward points can be set based on the four scenarios mentioned above. For example, A1) corresponds to reward point B1, A2) corresponds to reward point B2, A3) corresponds to reward point B3, and A4) corresponds to reward point B4. Among them, B1 is greater than B2, B2 is greater than B3, and B3 is greater than B4.
[0112] Therefore, based on different reward scores, the interface operations that can be used to set preference data pairs can be determined, so that the set data preference pairs can form positive and negative samples of interface operations with different reward scores (for example, interface operations with higher reward scores are determined as positive samples, and interface operations with lower reward scores are determined as negative samples). This allows the model to be given a clear reward signal during model training, and the corresponding loss function can be determined based on the above reward scores to adjust the model parameters and obtain the target model.
[0113] The loss function can be expressed, for example, by the following mathematical expression:
[0114]
[0115] In the above mathematical expression, π θ This represents the current policy, defined by the parameter θ. It indicates that the model is in a given task X and the current display interface, so P is the appropriate policy. t The probability distribution of choosing an action under certain circumstances.
[0116] π SFT : This represents the policy fine-tuned based on SFT, used to guide the optimization of the current policy.
[0117] X: Represents the task description, indicating the task that the agent needs to complete.
[0118] P t : Represents the current display interface, indicating the display interface obtained by the model at time step t.
[0119] This represents the positive sample in the preference data pair. This represents a negative sample in a pair of preference data.
[0120] D: Represents the preference dataset, including task X and the current page P. t Positive samples and negative samples The data pairs.
[0121] β: Represents the balancing parameter, used to control the weight difference between positive and negative sample actions, and is used to adjust the sensitivity of the loss function to positive and negative sample actions.
[0122] σ: represents the sigmoid function, which maps input values to the (0,1) interval and represents the probability.
[0123] Therefore, the loss function determined by the above implementation method can be used to adjust the parameters of the model during the training phase in order to obtain the target model.
[0124] Understandably, the action space of a display interface is usually determined based on the attribute information of the elements within the interface. For example, if the element is an icon, its size in the display interface is represented as {[a, b], [c, d]}, where [a, b] represents the coordinates of the top-left corner of the icon, and [c, d] represents the coordinates of the bottom-right corner. Therefore, the effective response range of an icon to interface operations usually does not exceed its corresponding size limit. Furthermore, it is understandable that the larger the effective response range of an icon, the larger the amount of data required to determine the action space of the corresponding display interface. To simplify the determination of the amount of data required for the action space of the display interface, in some scenarios, the action space of the display interface can be preprocessed.
[0125] For example, Figure 4 This is a flowchart illustrating a motion space generation method according to some embodiments of this disclosure. Figure 4 As shown, the method includes the following steps.
[0126] In step S31, the action attribute information corresponding to the k-th interface is determined.
[0127] In step S32, the target response range is selected from the valid response range corresponding to the element used to respond to interface operations in the k-th interface.
[0128] In step S33, the action space corresponding to the k-th interface is generated based on the target response range.
[0129] In step S34, the action space is input into the initial model to determine the interface operation that the electronic device needs to perform on the k-th interface.
[0130] Among them, the action attribute information represents the effective response range of the element used by the k-th interface to respond to interface operations.
[0131] In this embodiment, since action attribute information characterizes the effective response range corresponding to the element on the k-th interface that responds to interface operations, the element on the k-th interface that can respond to interface operations, and the effective response range corresponding to the element that can respond to interface operations, can be determined by knowing the action attribute information. This allows for the selection of a smaller target response range from the effective response range to generate the action space corresponding to the k-th interface, thereby reducing the computational load required to generate the action space.
[0132] In some embodiments, determining the action attribute information corresponding to the k-th interface can be achieved, for example, by obtaining the attribute document (e.g., an XML document) corresponding to the interface.
[0133] It should be noted that elements used to respond to interface operations can include at least one of the following: "icons", "text boxes", "scroll bars", etc. Action attribute information can be understood as the type of interface operation that can be responded to (e.g., whether it can respond to touch operations or swipe operations) and the effective response range (e.g., the maximum effective touch range or the maximum swipe range).
[0134] For example, within the valid response range corresponding to the element used to respond to interface operations in the k-th interface, a target response range can be selected. For instance, assuming the valid response range of an icon in the k-th display interface is represented as {[e, f], [g, h]}, where [e, f] represents the coordinates of the top-left corner of the icon within the valid response range, and [g, h] represents the coordinates of the bottom-right corner of the icon within the valid response range. In this case, the target response range can be selected as {[h, i], [j, k]} within the valid response range.
[0135] In summary, the above-described implementation methods can be used to determine the training framework for the target model training process and the related training procedures.
[0136] It is understood that, based on the above description of the relevant embodiments, the interface arrival task training set and the interface operation task training set used by the target model can be determined based on the task training set. For example, in some scenarios, the interface arrival task training set and the interface operation task training set can be obtained by decomposing the task training set accordingly. Therefore, obtaining the data used for training the target model can be achieved by obtaining the task training set.
[0137] In some implementation scenarios, the acquisition of task datasets may include steps C1) to C4).
[0138] C1) Obtain the interface flow with the required path length from the pre-trained dataset.
[0139] C2) Based on the preset first model, generate corresponding interface descriptions for the interface images in the interface flow.
[0140] C3) Based on the preset second model, generate corresponding descriptive information and task summary for each interface operation corresponding to each image. The task summary describes the task corresponding to the interface flow, and the descriptive information describes the display operation corresponding to each display interface.
[0141] C4) Based on image descriptions and descriptive information, the interface streams are filtered, and the first interface stream that meets the quality requirements is retained to determine the task training set.
[0142] For example, the pre-trained dataset can be some open-source interface streams.
[0143] For a pre-defined first model, such as a multimodal large language model, it can generate a corresponding interface description based on the input interface image. For example, input Figure 2C The interface image can generate a description of the "electronic device desktop display interface". For the preset second model, such as a large language model, it can generate description information corresponding to each interface operation based on the interface operation corresponding to the display interface, and can generate an overall task summary of the interface flow based on all interface descriptions and description information.
[0144] To facilitate understanding, the following table (1) will provide an example of some of the relevant elements obtained above for determining the task training set.
[0145] Table 1(1)
[0146]
[0147]
[0148] Therefore, given the availability of the task training set, we can base our decision on the following: Figure 5 In related embodiments, the interface arrival task training set and the interface operation task training set are determined.
[0149] Figure 5 This is a flowchart illustrating a method for determining the interface arrival task training set and the interface operation task training set according to some embodiments of this disclosure. Figure 5 As shown, the method includes the following steps.
[0150] In step S41, based on each interface image and the first interface operation corresponding to each interface image, the description information corresponding to each first interface operation is determined.
[0151] In step S42, based on the description information, the second interface stream and the third interface stream are determined in the first interface stream.
[0152] In step S43, an interface is constructed based on the second interface stream to reach the task training set, and an interface operation task training set is constructed based on the third interface stream.
[0153] The task training set includes: training tasks, and a first interface stream displayed by an electronic device corresponding to the completion of the training tasks. The first interface stream includes interface images and first interface operations corresponding to the interface images. In the first interface stream, there is at least one first interface operation between adjacent interface images. The second interface stream is the interface stream corresponding to the execution interface reaching the sub-task. The third interface stream is the interface stream used to execute the interface operation sub-task.
[0154] In this embodiment, by determining a second interface stream related to interface arrival and a third interface stream related to interface operation within the first interface stream based on the description information corresponding to the first interface stream, the interface arrival task training set constructed based on the second interface stream can be correlated with the task training set, the interface operation task training set constructed based on the third interface stream can be correlated with the task training set, and the interface arrival task training set and the interface operation task training set can be correlated. This allows the model, when trained using the task training set, the interface arrival task training set, and the interface operation task training set, to enhance its ability to execute tasks corresponding to each step and its processing capabilities between steps.
[0155] For ease of understanding, some features in the above embodiments will be explained by way of example in Table (1):
[0156] For example, the training task can be understood as “help me find detailed information about product A and add a white one to the shopping cart” in Table (1) above. For the first interface operation, it can be understood as the corresponding interface operation in the interface operation list in Table (1) above. For the first interface stream, it is understood as an ordered data stream consisting of the interface operation performed by the electronic device and the corresponding interface image displayed during the execution of the interface operation, and the interface images are connected through the interface operation.
[0157] In some embodiments, for constructing an interface arrival task training set based on the second interface stream, for example, a task describing the second interface stream can be generated based on the corresponding task generation template or large language text, and an interface arrival task training set can be constructed based on the second interface stream and the task corresponding to the second interface stream.
[0158] In some embodiments, for constructing a training set of interface operation tasks based on a third interface stream, for example, a task describing the third interface stream can be generated based on a corresponding task generation template or large language text, and an interface operation task training set can be constructed based on the third interface stream and the task corresponding to the third interface stream.
[0159] To facilitate understanding, the task training set in Table (1) above will be split into Table (2) below, and the interface arrival task training set and interface operation training set will be explained by example.
[0160] Table (2)
[0161]
[0162]
[0163] It can be seen that the interface operations in Table (2) can be understood as the first interface operations based on the interface flow in Table (1).
[0164] Among them, the first interface operation and the corresponding interface image for completing the interface arrival task are used to construct the second interface flow. For example, “accessing the search results interface” in Table (2) can be understood as an interface arrival task. The corresponding first interface operation (such as “clicking “search”, entering “product A”, clicking “search icon” and “task completed” in the table) and the interface image corresponding to the interface operation can be used to construct the second interface flow, and thus form an interface arrival task training set with “accessing the search interface”.
[0165] The first interface operation and the corresponding interface image for completing the interface operation task are used to construct the third interface flow. For example, "In the search interface, enter "Product A" to search" in Table (2) can be understood as an interface operation task. The corresponding first interface operation (click "search", enter "Product A" and complete the task) and the interface image corresponding to the interface operation can be used to construct the third interface flow, and thus form an interface operation task training set with "In the search interface, enter "Product A" to search".
[0166] For the second interface flow, it can be understood that if an electronic device presents an interface different from the previous interface after performing an interface operation (for example, a jump occurs), it can be determined that a new interface has been reached. That is, the electronic device displaying this interface can be considered as completing an "interface arrival task". Therefore, the determination of the interface arrival task flow can be made by whether a "new interface" appears compared to the previous interface image.
[0167] For example, Figure 6 This is a flowchart illustrating a method for determining a second interface flow in a first interface flow according to some embodiments of the present disclosure, such as... Figure 6 As shown, the method includes the following steps.
[0168] In step S51, a first interface image is determined in the first interface stream based on the description information, wherein the first interface image is the interface image that appears for the first time in the first interface stream.
[0169] In step S52, in the first interface stream, a second interface image and the interface operation corresponding to the second interface image are determined, wherein the second interface image is the interface image preceding the first interface image in the first interface stream.
[0170] In step S53, the second interface flow is determined based on the second interface image and the interface operation corresponding to the second interface image.
[0171] In this embodiment of the disclosure, since the description information includes relevant descriptions of interface operations, the description information can be used to determine the relevant description of the display interface corresponding to the interface operation performed by the electronic device. Furthermore, based on the relevant description of the display interface, it can be determined whether this display interface is the first appearance of a display interface compared to previous display interfaces, and thus whether this display interface can be used to construct a second interface stream. Therefore, when it is necessary to construct a second interface stream, corresponding construction rules can be set in this way to achieve automated construction in the first interface stream, improving data construction efficiency.
[0172] As for determining the first appearance of the interface image based on the description information, for example, if it is determined that the description information includes the first description of the interface name, then the interface image corresponding to that name can be considered as the first appearance of the interface image. For example, continuing from the description of the relevant embodiments in Table (1), the "details interface" in description information 4) does not appear in description information 1) to 3), so it can be considered that the "details interface" is the first appearance of the interface.
[0173] For example, if it is determined that the description information includes the first description of the name of an element in the interface, then the interface image of the element corresponding to that name can be considered as the first appearance of the interface image. Continuing from the description of the relevant embodiments in Table (1), the "search icon" in description information 2) does not appear in description information 1), so the "search results interface" in description information 3) can be considered as the first appearance of the interface image.
[0174] Therefore, the determination of the first image (i.e., the first interface image in the above-mentioned related embodiments) can be achieved through the following... Figure 7 The disclosed implementation method is determined. Figure 7 This is a flowchart illustrating a method for determining a first interface image in a first interface stream according to some embodiments of the present disclosure, such as... Figure 7 As shown, the method includes the following steps.
[0175] In step S61, the description information is determined.
[0176] In step S62-1, in response to the inclusion of first description information in the description information, the interface image corresponding to the first description information is determined as the second interface image.
[0177] In step S62-2, in response to the inclusion of second description information in the description information, the next interface image corresponding to the description information is determined as the second interface image.
[0178] The first descriptive information describes the name of the interface image, and there is a one-to-one correspondence between the first descriptive information and the interface image. The second descriptive information describes the names of the elements in the interface image, and the names of the elements are the names that appear for the first time in the descriptive information.
[0179] In this embodiment of the disclosure, the first interface image can be determined in the first interface stream by determining the content included in the description information. Therefore, when a second interface stream needs to be constructed, corresponding construction rules can be set in this way to automatically determine the first interface image in the first interface stream, thereby achieving automated construction of the second interface stream and improving data construction efficiency.
[0180] Understandably, to complete a user interface operation task, the model needs to reach the specified display interface and correctly perform the specified interface operation. However, the impact of different types of interface operations performed by electronic devices on the display interface may vary. Therefore, to ensure accurate segmentation of the third interface flow, the operation type corresponding to the interface operation can be taken into consideration.
[0181] For example, Figure 8 This is a flowchart illustrating a method for determining a third interface flow in a first interface flow according to some embodiments of the present disclosure. Figure 8 As shown, the method includes the following steps.
[0182] In step S71, a second interface operation is selected from the first interface operation of the first interface stream, wherein the second interface operation is an interface operation used to construct a training set of interface operation subtasks.
[0183] In step S72, the operation type corresponding to the second interface operation is determined based on the description information corresponding to the second interface operation and the second interface operation.
[0184] In step S73, a third interface flow is determined in the first interface flow based on the operation type.
[0185] In this embodiment, since the second interface operation is selected based on the first interface operation, and the first interface operation has descriptive information, the descriptive information corresponding to the second interface operation is obtained when the second interface operation is determined, and then the operation type corresponding to the second interface operation is determined based on the descriptive information. Therefore, accurate division of the third interface flow within the first interface flow can be achieved based on different operation types.
[0186] It is understandable that, as described in the related embodiments in Table (1), the interface operation actions are divided into various types (e.g., scrolling, clicking or typing, etc.), and the electronic device performs different types of interface operation actions on the current display interface. Before and after the interface operation is performed, the display interface may not change significantly.
[0187] For example, as described in information 1) in Table (1), it can be understood that after inputting "Product A", the display interface of the electronic device may not change (e.g., it still corresponds to the "search interface"), but the content displayed on the interface may change significantly, such as the addition of the text description "Product A". Therefore, when "inputting Product A" is used as the second interface operation to construct the third interface flow, in order to ensure the complete execution of the interface operation task, it is necessary to determine whether the interface operation task has been completed by checking the display interface after the interface operation is performed. That is, it is necessary to construct the third interface flow based on the display interface corresponding to the second interface operation performed by the electronic device.
[0188] For example, as described in information 5) in Table (1), it can be understood that after inputting "swipe up to view more parameter information", the display interface of the electronic device may not change (e.g., it still corresponds to the "parameter interface"), but the content displayed on the interface may change significantly, such as the text description changing. Therefore, when "swipe up to view more parameter information" is used as the second interface operation to construct the third interface flow, in order to ensure that the interface operation task is executed completely, it is necessary to determine whether the interface operation task has been executed completely by checking the display interface after the interface operation is executed. That is, it is necessary to construct the third interface flow based on the display interface corresponding to the second interface operation performed by the electronic device.
[0189] For example, as described in information 6) in Table (1), it can be understood that there may be other similar parameter descriptions such as "black" and "blue" in the parameter interface. After "selecting white", the display interface of the electronic device may not change (e.g., it is still on the "parameter interface"), and the content displayed on the interface may not change significantly. For example, the parameter descriptions in the interface may still include "black" and "blue". Therefore, when "selecting white" is used as the second interface operation to construct the third interface flow, the third interface flow can be constructed based only on the interface images adjacent to the second interface operation and the second interface operation. In this way, when the model learns how to "select white" attribute, it can share the same prior operation logic (e.g., the prior logic on how to reach the parameter interface) and only need to learn each specific color parameter separately, thereby improving the model training efficiency.
[0190] For example, Figure 9 This is a flowchart illustrating a method for determining a third interface flow according to some embodiments of this disclosure. Figure 9 As shown, the method includes the following steps.
[0191] In step S81, the operation type is determined.
[0192] In step S82-1, in response to an interface operation of the first operation type, a third interface image is determined in the first interface stream, and the interface stream preceding the third interface image is determined as the third interface stream.
[0193] In step S82-2, in response to the second interface operation being an interface operation of the second operation type, and the similarity between the third interface image and the fourth interface image meeting the similarity requirement, the third interface flow is determined based on the third interface image, the second interface operation, and the fourth interface image.
[0194] The interface operations of the first operation type include swiping operations and / or typing operations, and the third interface image is the interface image in the first interface flow that follows the second interface operation. The interface operations of the second operation type include clicking operations, and the third interface image is the interface image in the first interface flow that follows the second interface operation, and the fourth interface image is the interface image in the first interface flow that precedes the second interface operation.
[0195] In this embodiment, on the one hand, by determining the interface images used to construct the third interface flow based on different interface operation types, a correlation is established between the training sets of different interface operation tasks and different interface operation types. This allows the model trained on the training sets of different interface operation tasks to pay attention to the impact of interface operations of different interface operation types on the display interface. On the other hand, by combining the similarity of the display interface before and after the interface operation to determine the third interface flow, the third interface flow can be simplified, thereby simplifying the training set of interface operation tasks constructed based on the third interface flow and improving the model training efficiency.
[0196] Therefore, the above implementation method can be used to determine the task training set, the interface arrival task training set, and the interface operation training set.
[0197] In summary, the task execution method provided by the above-described embodiments of this disclosure has a higher probability of success because the target model is jointly trained based on the task training set, the interface arrival task training set, and the interface operation task training set. Therefore, the trained target model is better at reaching the target interface and performing target operations within the target interface during the execution of the target task. It is also less likely to get trapped in local optima and fail to complete the task, thus increasing the probability of task completion.
[0198] Figure 10This is a block diagram illustrating a task execution apparatus according to some embodiments of the present disclosure. (Refer to...) Figure 10 The device 100 includes an acquisition unit 101 and a processing unit 102.
[0199] The acquisition unit 101 is used to determine the target task to be executed. The target task is to control the electronic device to perform the target operation on the target interface. The target interface is the interface after the electronic device is triggered to display m interfaces sequentially from the current interface during the execution of the target task, where m is a positive integer.
[0200] The processing unit 102 is used to determine the i-th interface currently triggered and displayed during the execution of the target task based on the target model, and to determine the interface operation sub-task to be executed in the i-th interface;
[0201] If the i-th interface is not the target interface and the operation corresponding to the interface operation subtask is not the target operation, the processing unit is used to execute the interface operation subtask, display the (i+1)-th interface, and determine the interface operation subtask executed in the (i+1)-th interface until it is determined that the electronic device displays the target interface and the operation corresponding to the interface operation subtask is the target operation.
[0202] When the i-th interface is the target interface and the operation corresponding to the interface operation subtask is the target operation, the processing unit is used to complete the target task.
[0203] The target model is used to determine the interface reached when the target task is executed and the interface operation sub-tasks to be executed on the interface reached. The target model is trained based on the task training set, the interface arrival task training set, and the interface operation task training set. The interface arrival task training set and the interface operation task training set are determined based on the task training set.
[0204] In some embodiments, the device 100 further includes a training unit 103, which is used to train a target model. The training unit 103 trains the target model in the following manner: inputting a task training set, an interface arrival task training set, an interface operation task training set, and the kth interface currently displayed by the electronic device into an initial model; determining the interface operation that the electronic device needs to perform in the kth interface, where k is a positive integer; determining preference data pairs based on the corresponding interface operations performed in the kth interface; and determining the loss function corresponding to the initial model based on the preference data pairs, where different interface operations in the kth interface correspond to different scores, and the preference data pairs include interface operations corresponding to different scores; and adjusting the initial model based on the loss function to obtain the target model.
[0205] In some implementations, the training unit 103 is further configured to: determine the action attribute information corresponding to the k-th interface, wherein the action attribute information characterizes the effective response range corresponding to the element of the k-th interface used to respond to the interface operation; select a target response range from the effective response range corresponding to the element of the k-th interface used to respond to the interface operation; generate the action space corresponding to the k-th interface based on the target range; and input the action space into the initial model to determine the interface operation that the electronic device needs to perform in the k-th interface.
[0206] In some implementations, the task training set includes: a training task, and a first interface stream displayed by an electronic device corresponding to the completion of the training task. The first interface stream includes interface images and first interface operations corresponding to the interface images, wherein adjacent interface images in the first interface stream correspond to at least one first interface operation. The interface arrival task training set and the interface operation task training set are determined based on the task training set in the following manner: based on each interface image and the first interface operation corresponding to each interface image, descriptive information corresponding to each first interface operation is determined; based on the descriptive information, a second interface stream and a third interface stream are determined in the first interface stream, wherein the second interface stream is the interface stream corresponding to the execution of the interface arrival sub-task, and the third interface stream is the interface stream used to execute the interface operation sub-task; the interface arrival task training set is constructed based on the second interface stream, and the interface operation task training set is constructed based on the third interface stream.
[0207] In some implementations, the second interface flow is determined as follows: based on description information, a first interface image is determined in the first interface flow, wherein the first interface image is the interface image that first appears in the first interface flow; in the first interface flow, a second interface image and the interface operation corresponding to the second interface image are determined, wherein the second interface image is the interface image preceding the first interface image in the first interface flow; and the second interface flow is determined based on the second interface image and the interface operation corresponding to the second interface image.
[0208] In some implementations, the first interface image is determined as follows: when the description information includes first description information, the interface image corresponding to the first description information is determined as the second interface image, wherein the first description information is used to describe the name of the interface image, and the first description information and the interface image have a one-to-one correspondence; when the description information includes second description information, the next interface image corresponding to the description information is determined as the second interface image, wherein the second description information is used to describe the name corresponding to the element in the interface image, and the name corresponding to the element is the name that appears for the first time in the description information.
[0209] In some implementations, the third interface flow is determined as follows: a second interface operation is selected from the first interface operations in the first interface flow, wherein the second interface operation is an interface operation used to construct a training set of interface operation subtasks; the operation type corresponding to the second interface operation is determined based on the description information corresponding to the second interface operation and the second interface operation; and the third interface flow is determined in the first interface flow based on the operation type.
[0210] In some implementations, the third interface flow is determined as follows: When the second interface operation is an interface operation of the first operation type, a third interface image is determined in the first interface flow, and the interface flow preceding the third interface image is determined as the third interface flow. The first operation type interface operation includes a swipe operation and / or a keystroke operation. The third interface image is the interface image in the first interface flow that follows the second interface operation. When the second interface operation is an interface operation of the second operation type, and the similarity between the third interface image and the fourth interface image meets the similarity requirement, the third interface flow is determined based on the third interface image, the second interface operation, and the fourth interface image. The second operation type interface operation includes a click operation. The third interface image is the interface image in the first interface flow that follows the second interface operation, and the fourth interface image is the interface image in the first interface flow that precedes the second interface operation.
[0211] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0212] Figure 11 This is a block diagram of an apparatus for task execution according to some embodiments of the present disclosure. For example, apparatus 200 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0213] Reference Figure 11 The device 200 may include one or more of the following components: processing component 202, memory 204, power component 206, multimedia component 208, audio component 210, input / output (I / O) interface 212, sensor component 214, and communication component 216.
[0214] Processing component 202 typically controls the overall operation of device 200, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 202 may include one or more processors 220 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 202 may include one or more modules to facilitate interaction between processing component 202 and other components. For example, processing component 202 may include a multimedia module to facilitate interaction between multimedia component 208 and processing component 202.
[0215] Memory 204 is configured to store various types of data to support the operation of device 200. Examples of such data include instructions for any application or method operating on device 200, contact data, phonebook data, messages, pictures, videos, etc. Memory 204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0216] The power supply component 206 provides power to the various components of the device 200. The power supply component 206 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 200.
[0217] Multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 208 includes a front-facing camera and / or a rear-facing camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0218] Audio component 210 is configured to output and / or input audio signals. For example, audio component 210 includes a microphone (MIC) configured to receive external audio signals when device 200 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 204 or transmitted via communication component 216. In some embodiments, audio component 210 also includes a speaker for outputting audio signals.
[0219] I / O interface 212 provides an interface between processing component 202 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0220] Sensor assembly 214 includes one or more sensors for providing status assessments of various aspects of device 200. For example, sensor assembly 214 may detect the on / off state of device 200, the relative positioning of components such as the display and keypad of device 200, changes in the position of device 200 or a component of device 200, the presence or absence of user contact with device 200, the orientation or acceleration / deceleration of device 200, and temperature changes of device 200. Sensor assembly 214 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 214 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 214 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0221] Communication component 216 is configured to facilitate wired or wireless communication between device 200 and other devices. Device 200 can access wireless networks based on communication standards, such as WiFi, 3G, 4G, 5G, other communication standards, or combinations thereof. In some embodiments of this disclosure, communication component 216 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In some embodiments of this disclosure, communication component 216 further includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0222] In some embodiments of this disclosure, the apparatus 200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0223] In some embodiments of this disclosure, a storage medium including instructions is also provided, such as a memory 204 including instructions, which can be executed by the processor 220 of the device 200 to perform the above-described method. For example, the storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0224] In some embodiments of this disclosure, a storage medium is provided, which may be a non-transitory computer-readable storage medium.
[0225] In some embodiments of this disclosure, when instructions in the storage medium are executed by the processor of the device 200, the device 200 is able to perform the methods described above.
[0226] Figure 12 This is a block diagram two illustrating an apparatus for task execution according to some embodiments disclosed in a publication. For example, apparatus 300 can be provided as a server. See also... Figure 12 The device 300 includes a processing component 322, which further includes one or more processors, and memory resources represented by memory 332 for storing instructions, such as application programs, that can be executed by the processing component 322. The application programs stored in memory 332 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 322 is configured to execute instructions to perform the methods described above.
[0227] Device 300 may also include a power supply component 326 configured to perform power management of device 300, a wired or wireless network interface 350 configured to connect device 300 to a network, and an input / output (I / O) interface 358. Device 300 may operate on an operating system stored in memory 332, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.
[0228] In some embodiments of this disclosure, a storage medium including instructions is also provided, such as a memory 204 including instructions, which can be executed by the processing component 322 of the device 300 to perform the above-described method. For example, the storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0229] In some embodiments of this disclosure, a storage medium is provided, which may be a non-transitory computer-readable storage medium.
[0230] In some embodiments of this disclosure, when instructions in the storage medium are executed by the processor of device 300, device 300 is able to perform the methods described above.
[0231] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this disclosure can be implemented by electronic hardware, computer software, or a combination of both. Whether such functionality is implemented in hardware or software depends on the specific application and the overall system design requirements. Those skilled in the art can implement the described functionality using various methods for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this disclosure.
[0232] Although terms such as “first,” “second,” and “third” may be used herein to describe various components, parts, regions, layers, or sections, these components, parts, regions, layers, or sections are not limited to these terms. Rather, these terms are used only to distinguish one component, part, region, layer, or section from another. Therefore, without departing from the teachings of the examples described herein, a first component, part, region, layer, or section mentioned in the examples may also be referred to as a second component, part, region, layer, or section. Furthermore, the terms “first” and “second” are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as “first” or “second” may explicitly or implicitly include at least one of that feature.
[0233] It is further understood that the terms "first," "second," etc., are used to describe various types of information, but this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another, and do not indicate a specific order or degree of importance. In fact, the expressions "first," "second," etc., are completely interchangeable. For example, without departing from the scope of this disclosure, first information can also be referred to as second information, and similarly, second information can also be referred to as first information.
[0234] In this description, "multiple" means at least two, referring to two or more, such as two, three, etc., unless otherwise explicitly specified. Other quantifiers are similar. The singular forms "a," "the," and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, unless otherwise specified or clearly indicated from the context, the articles "a" and "an" as used in this application and the appended claims are generally understood to mean "one or more."
[0235] In this description, the terms “in response to…”, “in response to determining…”, “in the case of…”, “when…”, “if…”, “if…”, etc. can be used interchangeably.
[0236] In this description, the terms “greater than”, “greater than or equal to”, “not less than”, “more than”, “more than or equal to”, “not less than”, “higher than”, “higher than or equal to”, “not lower than”, and “above” can be used interchangeably. The terms “less than”, “less than or equal to”, “not greater than”, “less than”, “less than or equal to”, “not more than”, “lower than”, “lower than or equal to”, “not higher than”, and “below” can be used interchangeably.
[0237] It should be understood that, unless otherwise specifically indicated, features of various embodiments of this disclosure described herein can be combined with each other. As used herein, the term "and / or" includes any one of the related listed items and any combination of two or more; "and / or" describes the association relationship between related objects, indicating that three relationships may exist, for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Similarly, "at least one of..." includes any one of the related listed items and any combination of two or more.
[0238] It is further understood that the terms "first," "second," etc., are used to describe various types of information, but this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another, and do not indicate a specific order or degree of importance. In fact, the expressions "first," "second," etc., are completely interchangeable. For example, without departing from the scope of this disclosure, first information can also be referred to as second information, and similarly, second information can also be referred to as first information.
[0239] Furthermore, the term "exemplary" is used herein to indicate that it serves as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous compared to other aspects or designs. Rather, the use of the term "exemplary" is intended to present concepts in a concrete manner. As used herein, the term "or" is intended to indicate an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, "X applies A or B" is intended to indicate any of the natural inclusive permutations. That is, if X applies A; X applies B; or X applies both A and B, then applying A or B satisfies the condition under any of the foregoing instances.
[0240] Similarly, although this disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding the specification and drawings. This disclosure includes all such modifications and variations and is limited only by the scope of the claims. In particular, with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terminology used to describe such components is intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if it is not structurally equivalent to the disclosed structure. Furthermore, although specific features of this disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations, as may be desired and advantageous to any given or particular application. Moreover, with regard to the terms “comprising,” “owning,” “having,” “having,” or variations thereof as used in this disclosure, such terms are intended to be inclusive in a manner similar to the term “including.”
[0241] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein.
[0242] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A task execution method, characterized in that, include: The target task to be executed is determined, wherein the target task is to control the electronic device to perform a target operation on the target interface, and the target interface is the interface after the electronic device is triggered to display m interfaces sequentially from the current interface during the execution of the target task, where m is a positive integer; Based on the target model, the i-th interface currently triggered and displayed during the execution of the target task is determined, and the interface operation sub-tasks executed on the i-th interface are determined. The target model is used to determine the interface reached by the execution of the target task and the interface operation sub-tasks executed on the reached interface. If the i-th interface is not the target interface, and the operation corresponding to the interface operation subtask is not the target operation, then the interface operation subtask is executed, the (i+1)-th interface is displayed, and the interface operation subtask executed in the (i+1)-th interface is determined, until it is determined that the electronic device displays the target interface and the operation corresponding to the interface operation subtask is the target operation; If the i-th interface is the target interface, and the operation performed by the interface operation subtask is the target operation, then the target task is completed. The target model is trained based on the task training set, the interface arrival task training set, and the interface operation task training set, and the interface arrival task training set and the interface operation task training set are determined based on the task training set.
2. The method according to claim 1, characterized in that, The target model is trained in the following manner: Input the task training set, the interface arrival task training set, the interface operation task training set, and the kth interface currently displayed by the electronic device into the initial model, and determine the interface operation that the electronic device needs to perform on the kth interface, where k is a positive integer; Based on the corresponding interface operation performed in the k-th interface, a preference data pair is determined, and the loss function corresponding to the initial model is determined based on the preference data pair. Here, different interface operations in the k-th interface correspond to different scores, and the preference data pair includes interface operations corresponding to different scores. The initial model is adjusted based on the loss function to obtain the target model.
3. The method according to claim 2, characterized in that, Before determining the interface operation that the electronic device needs to perform on the k-th interface, the method further includes: Determine the action attribute information corresponding to the k-th interface, wherein the action attribute information characterizes the effective response range of the element of the k-th interface used to respond to interface operations; Select the target response range from the effective response range; Based on the target response range, the action space corresponding to the kth interface is generated; The action space is input into the initial model to determine the interface operation that the electronic device needs to perform on the k-th interface.
4. The method according to claim 1 or 2, characterized in that, The task training set includes: a training task, and a first interface stream displayed by an electronic device corresponding to the completion of the training task. The first interface stream includes interface images and first interface operations corresponding to the interface images. In the first interface stream, there is at least one first interface operation between adjacent interface images. The interface arrival task training set and the interface operation task training set are determined based on the task training set in the following manner: Based on each interface image and the first interface operation corresponding to each interface image, determine the description information corresponding to each first interface operation. Based on the description information, a second interface flow and a third interface flow are determined in the first interface flow, wherein the second interface flow is the interface flow corresponding to the subtask reached by the execution interface, and the third interface flow is the interface flow used to execute the interface operation subtask. The interface arrival task training set is constructed based on the second interface stream, and the interface operation task training set is constructed based on the third interface stream.
5. The method according to claim 4, characterized in that, Determining a second interface flow in the first interface flow based on the described information includes: Based on the description information, a first interface image is determined in the first interface stream, wherein the first interface image is the interface image that appears for the first time in the first interface stream; In the first interface stream, a second interface image and the interface operation corresponding to the second interface image are determined, wherein the second interface image is the interface image preceding the first interface image in the first interface stream. The second interface flow is determined based on the second interface image and the interface operation corresponding to the second interface image.
6. The method according to claim 5, characterized in that, Determining the first interface image in the first interface stream based on the description information includes: In response to the inclusion of first description information in the description information, the interface image corresponding to the first description information is determined as the second interface image, wherein the first description information is used to describe the name of the interface image, and the first description information and the interface image have a one-to-one correspondence. In response to the inclusion of second description information in the description information, the next interface image corresponding to the description information is determined as the second interface image, wherein the second description information is used to describe the name corresponding to an element in the interface image, and the name corresponding to the element is the name that appears for the first time in the description information.
7. The method according to claim 4, characterized in that, Based on the description information, determining the third interface flow in the first interface flow includes: In the first interface operation of the first interface stream, a second interface operation is selected, wherein the second interface operation is an interface operation used to construct a training set of interface operation subtasks; Based on the description information corresponding to the second interface operation and the second interface operation, determine the operation type corresponding to the second interface operation; Based on the operation type, a third interface flow is determined in the first interface flow.
8. The method according to claim 7, characterized in that, Determining the third interface flow in the first interface flow based on the operation type includes: In response to an interface operation of the first operation type, a third interface image is determined in the first interface flow, and the interface flow preceding the third interface image is determined as the third interface flow. The interface operation of the first operation type includes a swipe operation and / or a keystroke operation, and the third interface image is the interface image adjacent to the second interface operation in the first interface flow. In response to the second interface operation being a second type of interface operation, and the similarity between the third interface image and the fourth interface image meeting the similarity requirement, the third interface flow is determined based on the third interface image, the second interface operation, and the fourth interface image. The second type of interface operation includes a click operation; the third interface image is the interface image in the first interface stream that is adjacent to the second interface operation after it; and the fourth interface image is the interface image in the first interface stream that is adjacent to the second interface operation before it.
9. A task execution device, characterized in that, include: The acquisition unit is used to determine the target task to be executed, wherein the target task is a task to control the electronic device to perform a target operation on a target interface, and the target interface is the interface after the electronic device is triggered to display m interfaces sequentially from the current interface during the execution of the target task, wherein m is a positive integer; The processing unit is used to determine, based on the target model, the i-th interface currently triggered and displayed during the execution of the target task, and to determine the interface operation sub-task to be executed on the i-th interface; If the i-th interface is not the target interface and the operation corresponding to the interface operation subtask is not the target operation, the processing unit is used to execute the interface operation subtask, display the (i+1)-th interface, and determine the interface operation subtask executed in the (i+1)-th interface, until it is determined that the electronic device displays the target interface and the operation corresponding to the interface operation subtask is the target operation. When the i-th interface is the target interface, and the operation performed by the interface operation subtask is the target operation, the processing unit is used to complete the target task. The target model is used to determine the interface reached by executing the target task and the interface operation sub-tasks executed on the reached interface. The target model is trained based on the task training set, the interface arrival task training set, and the interface operation task training set, which are determined based on the task training set.
10. The apparatus according to claim 9, characterized in that, The device further includes a training unit for training the target model. The training unit trains the target model in the following manner: Input the task training set, the interface arrival task training set, the interface operation task training set, and the kth interface currently displayed by the electronic device into the initial model, and determine the interface operation that the electronic device needs to perform on the kth interface, where k is a positive integer; Based on the corresponding interface operation performed in the k-th interface, a preference data pair is determined, and the loss function corresponding to the initial model is determined based on the preference data pair. Here, different interface operations in the k-th interface correspond to different scores, and the preference data pair includes interface operations corresponding to different scores. The initial model is adjusted based on the loss function to obtain the target model.
11. The apparatus according to claim 10, characterized in that, The training unit is also used for: Determine the action attribute information corresponding to the k-th interface, wherein the action attribute information characterizes the effective response range of the element of the k-th interface used to respond to interface operations; Select the target response range from the effective response range; Based on the target response range, the action space corresponding to the kth interface is generated; The motion space is input into the initial model to determine the interface operation that the electronic device needs to perform on the k-th interface.
12. The apparatus according to claim 9 or 10, characterized in that, The task training set includes: a training task, and a first interface stream displayed by an electronic device corresponding to the completion of the training task. The first interface stream includes interface images and first interface operations corresponding to the interface images. In the first interface stream, there is at least one first interface operation between adjacent interface images. The interface arrival task training set and the interface operation task training set are determined based on the task training set in the following manner: Based on each interface image and the first interface operation corresponding to each interface image, determine the description information corresponding to each first interface operation. Based on the description information, a second interface flow and a third interface flow are determined in the first interface flow, wherein the second interface flow is the interface flow corresponding to the subtask reached by the execution interface, and the third interface flow is the interface flow used to execute the interface operation subtask. The interface arrival task training set is constructed based on the second interface stream, and the interface operation task training set is constructed based on the third interface stream.
13. The apparatus according to claim 12, characterized in that, The second interface flow is determined in the following way: Based on the description information, a first interface image is determined in the first interface stream, wherein the first interface image is the interface image that appears for the first time in the first interface stream; In the first interface stream, a second interface image and the interface operation corresponding to the second interface image are determined, wherein the second interface image is the interface image preceding the first interface image in the first interface stream. The second interface flow is determined based on the second interface image and the interface operation corresponding to the second interface image.
14. The apparatus according to claim 13, characterized in that, The first interface image is determined in the following way: When the description information includes first description information, the interface image corresponding to the first description information is determined as the second interface image, wherein the first description information is used to describe the name of the interface image, and the first description information and the interface image have a one-to-one correspondence. When the description information includes second description information, the next interface image corresponding to the description information is determined as the second interface image, wherein the second description information is used to describe the name corresponding to the element in the interface image, and the name corresponding to the element is the name that appears for the first time in the description information.
15. The apparatus according to claim 12, characterized in that, The third interface flow is determined in the following manner: In the first interface operation of the first interface stream, a second interface operation is selected, wherein the second interface operation is an interface operation used to construct a training set of interface operation subtasks; Based on the description information corresponding to the second interface operation and the second interface operation, determine the operation type corresponding to the second interface operation; Based on the operation type, a third interface flow is determined in the first interface flow.
16. The apparatus according to claim 15, characterized in that, The third interface flow is determined in the following manner: When the second interface operation is an interface operation of the first operation type, a third interface image is determined in the first interface flow, and the interface flow preceding the third interface image is determined as the third interface flow. The interface operation of the first operation type includes a swipe operation and / or a keystroke operation, and the third interface image is the interface image adjacent to the second interface operation in the first interface flow. If the second interface operation is a second type of interface operation, and the similarity between the third interface image and the fourth interface image meets the similarity requirement, then the third interface flow is determined based on the third interface image, the second interface operation, and the fourth interface image. The second type of interface operation includes a click operation; the third interface image is the interface image in the first interface stream that is adjacent to the second interface operation after it; and the fourth interface image is the interface image in the first interface stream that is adjacent to the second interface operation before it.
17. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the task execution method according to any one of claims 1 to 8.
18. A storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the task execution method according to any one of claims 1 to 8.
19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the task execution method as described in any one of claims 1 to 8.