Data processing method and device, electronic equipment and storage medium
By performing text and icon annotation and operation reasoning on the GUI interface, combining multimodal model and environmental knowledge base, the problem of insufficient flexibility in GUI automation processing is solved, and more efficient task execution and adaptability is achieved.
Patent Information
- Application Number
- CN202510510843.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the automation processing flexibility of GUI is poor and it is difficult to adapt to dynamic and complex application scenarios.
Through the object detection model, the user graphical interface before and after the task operation is marked with text and icons, combined with task instructions and recorded data for operation reasoning, and the multimodal model and environmental knowledge base are used to improve the accuracy and applicability of task execution.
It improves the accuracy of detection and labeling, enhances the universality of each processing scenario, ensures the context consistency of long-term tasks, and improves the success rate of task execution and the ability to adapt to complex GUI environments.
Smart Images

Figure CN120343333A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a data processing method, apparatus, electronic device, storage medium, and program product. Background Art
[0002] With the development of computer technology, a graphical user interface (GUI) has emerged. The GUI is the core technology of human-computer interaction and is also a relatively intuitive and visually driven way. Users can access digital systems, network systems, etc. through the GUI and interact with the system. To further improve the efficiency of human-computer interaction, the user's needs have gradually shifted to the automated processing of the GUI interface. Usually, the automated processing of the GUI is driven by preset scripts and rules. However, this preset script-driven GUI processing method is only applicable to fixed processes, resulting in poor flexibility in GUI interface processing. Summary of the Invention
[0003] The present disclosure provides a data processing method, apparatus, electronic device, storage medium, and program product to at least solve the problem of poor flexibility in the automated processing of the GUI in the related art. The technical solutions of the present disclosure are as follows:
[0004] According to a first aspect of an embodiment of the present disclosure, a data processing method is provided, including:
[0005] Obtaining a target task and a task instruction of the target task, where the target task includes at least one task operation;
[0006] Obtaining a first graphical user interface before performing the task operation and a second graphical user interface after performing the task operation, and performing annotation processing through a target detection model to obtain an annotated graphical user interface annotated with text elements and image elements;
[0007] Performing operation reasoning based on the task instruction, the recorded data corresponding to the task operation, and the annotated graphical user interface to obtain a task operation to be executed, and executing the task operation to be executed until a preset task completion condition is met, to obtain a graphical user interface for completing the execution of the target task.
[0008] In one embodiment, the method further includes:
[0009] Performing task decomposition on the target task to obtain a plurality of subtasks of the target task and task instructions of each subtask;
[0010] Encoding the task instructions of each subtask to obtain a plurality of task instruction vectors;
[0011] In a preset task knowledge base, queries are respectively performed based on each of the task instruction vectors to obtain a plurality of similar vectors that meet the first similarity condition. Based on the planning knowledge of each of the similar vectors and the task instruction of the target task, a task planning suggestion is obtained, and the task planning suggestion is used to determine each task operation of the target task.
[0012] In one embodiment, the method further includes:
[0013] Determine the element images associated with the target task of the first graphical user interface, where the elements include text elements and / or icon elements;
[0014] Determine the text description data of each of the elements, and encode the element images and the text description data through a multimodal model to obtain the multimodal vector of the first graphical user interface.
[0015] In a preset environment knowledge base, a query is performed based on the multimodal vector of the first graphical user interface to obtain a multimodal similar vector that meets the second similarity condition, and determine the environment knowledge corresponding to the multimodal similar vector; multiple multimodal vectors and corresponding environment knowledge are stored in the environment knowledge base.
[0016] In one embodiment, the recorded data includes action records and information records for executing the task operation; the operation reasoning based on the task instruction, the recorded data corresponding to the task operation, and the labeled graphical user interface to obtain the task operation to be executed includes:
[0017] Based on the task instruction, the action record and information record corresponding to the task operation, and the labeled graphical user interface, perform operation reasoning to obtain the execution result of the task operation;
[0018] If the execution result is correct execution, then based on the task instruction, the action memory, information memory, the labeled graphical user interface, and the environment knowledge corresponding to the task operation, determine that the task operation to be executed is the next task operation of the task operation;
[0019] If the execution result is incorrect execution, then determine that the task operation to be executed is to re-execute the task operation.
[0020] In one embodiment, the preset task completion condition includes that the step length of the task operation reaches a preset step length threshold, or the target task has been completed.
[0021] In one embodiment, the target task is determined in the current task pool; the method further includes:
[0022] After all tasks in the current task pool are completed, a first task list of successfully completed tasks and a second task list of unsuccessfully completed tasks are obtained;
[0023] Combine and / or modify each task in the first task list to obtain multiple new tasks, and split each task in the second task list to obtain multiple subtasks. The complexity of the new tasks is higher than that of the tasks in the first task list, and the complexity of the subtasks is lower than that of the tasks in the second task list;
[0024] Based on each of the new tasks and each of the subtasks, an updated task pool is obtained, and according to the complexity of each of the new tasks and each subtask and the dependency relationship between tasks, the execution priority and execution order of each task in the updated task pool are obtained.
[0025] In one embodiment, the method further includes:
[0026] Obtain the task planning text information of each task, respectively encode the task planning text information of each task to obtain a planning vector, and add the planning vector corresponding to each task to a preset task knowledge base.
[0027] In one embodiment, the method further includes:
[0028] Based on the first graphical user interface and the second graphical user interface corresponding to the task operation, determine the graphical user interface difference data after executing the task operation;
[0029] Based on the graphical user interface difference data and the task operation, obtain environmental knowledge;
[0030] Take screenshots of each element on the first graphical user interface to obtain element images and obtain text description data corresponding to each of the element images; encode the element images and the text descriptions through the multimodal model to obtain the multimodal vector of the first graphical user interface;
[0031] Add the multimodal vector of the first graphical user interface corresponding to each task operation and the environmental knowledge to the environmental knowledge base.
[0032] According to a second aspect of the embodiments of the present disclosure, a data processing device is provided, including:
[0033] A first acquisition unit configured to execute acquiring a target task and a task instruction of the target task, where the target task includes at least one task operation;
[0034] A second acquisition unit, configured to acquire a first graphical user interface before performing the task operation and a second graphical user interface after performing the task operation, and perform annotation processing through a target detection model to obtain an annotated graphical user interface with text elements and image elements annotated thereon;
[0035] An inference unit, configured to perform operation inference based on the task instruction, the recorded data corresponding to the task operation, and the annotated graphical user interface to obtain a task operation to be executed, and execute the task operation to be executed until a preset task completion condition is met, obtaining a graphical user interface for completing the target task.
[0036] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0037] A processor;
[0038] A memory for storing executable instructions of the processor;
[0039] Wherein, the processor is configured to execute the instructions to implement the data processing method according to any one of the above first aspects.
[0040] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the data processing method according to any one of the above first aspects.
[0041] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, when the instructions are executed by a processor of an electronic device, enabling the electronic device to execute the data processing method according to any one of the above first aspects.
[0042] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0043] By using a target detection model to separately annotate text and icons on the user graphical interface before and after performing a task operation, the accuracy of detection and annotation can be improved, and the generality for various processing scenarios can be enhanced. By performing data processing on the simply acquired visual data and avoiding acquiring underlying data, the convenience of obtaining annotated data is improved; by performing operation inference based on the task instruction, the recorded data corresponding to the task operation, and the annotated graphical user interface to obtain a task operation to be executed, the context consistency during the execution of long-time sequence tasks can be ensured. By recording data to ensure that key information related to the task is recorded, the applicability to various dynamically changing GUI application scenarios is further ensured, and the success rate of task execution is improved.
[0044] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. Brief Description of the Drawings
[0045] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an undue limitation to the present disclosure.
[0046] Figure 1 It is a flowchart of a data processing method shown according to an exemplary embodiment.
[0047] Figure 2 It is a flowchart of the step of obtaining a task planning suggestion in an exemplary embodiment.
[0048] Figure 3 It is a flowchart of the step of obtaining environmental knowledge in an exemplary embodiment.
[0049] Figure 4 It is a flowchart of the step of determining a task operation to be executed in an exemplary embodiment.
[0050] Figure 5 It is a flowchart of the step of updating a task pool in an exemplary embodiment.
[0051] Figure 6 It is a flowchart of the step of obtaining an environmental knowledge base in an exemplary embodiment.
[0052] Figure 7 It is the processing flow of an observation processing module in a data processing method shown according to an exemplary embodiment.
[0053] Figure 8 It is the execution flowchart of a basic agent in a data processing method shown according to an exemplary embodiment.
[0054] Figure 9 It is a flowchart of obtaining environmental knowledge in a data processing method shown according to an exemplary embodiment.
[0055] Figure 10 It is a flowchart of obtaining task knowledge in a data processing method shown according to an exemplary embodiment.
[0056] Figure 11 It is the execution flowchart of a self-driven exploration module in a data processing method shown according to an exemplary embodiment.
[0057] Figure 12 It is a flowchart of a data processing method shown according to another exemplary embodiment.
[0058] Figure 13 It is a block diagram of a data processing device shown according to an exemplary embodiment.
[0059] Figure 14 It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0060] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0061] It should be noted that the terms "first", "second", etc. in the specification, claims and drawings of the present disclosure are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0062] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.
[0063] The graphical user interface (GUI) is the core of human-computer interaction, which is a relatively intuitive and visually driven way, facilitating users to access and interact with digital systems. The GUI automation methods in related technologies rely on scripts or rules to drive, lacking flexibility and adaptability in dynamic and complex actual application scenarios.
[0064] The data processing system corresponding to the data processing method provided in this embodiment includes at least an observation processing module, a basic agent module, an environment knowledge base module, a task knowledge base module, and a self-driven exploration module, which includes an outer loop, a middle loop, and an inner loop. The outer loop means that the agent module continuously learns through the self-driven exploration module and continuously updates the task pool. The task pool stores various types of tasks, further improving the task learning ability and task execution ability of the model. The middle loop means that for a given task pool, a target task is selected from multiple tasks / task instructions included in the task pool, and the task knowledge base module is scheduled to retrieve task knowledge based on the task instructions of the target task, providing task planning knowledge for the agent model to execute the target task. Based on the task planning knowledge, the agent model can enter the inner loop. When the target task and the task instructions corresponding to the target task are determined, the agent module can complete a single target task by repeatedly calling the observation processing module, the basic agent module, and the environment knowledge base module and executing multiple task operations in the target task. Based on this, through the collaborative operation of each module and the self-driven exploration learning process of multiple nested loops, the agent model realizes efficient and intelligent autonomous learning and task execution in the GUI environment, and the agent model in this data processing method also has strong generality and applicability, and can be applied to various GUI environments. The agent model therein can be an intelligent model running on an electronic device.
[0065] Specifically, the inner loop repeatedly calls the observation processing module, the basic agent module, and the environment knowledge base module to execute a single task when the task is determined. For example, after selecting the target task, the above three modules are called to execute the target task, which can handle various complex situations, such as the dynamic change of the GUI interface and the ambiguity of the task goal. The agent model can continuously evaluate the execution situation of the target task in real time. When the task end condition is met, that is, the task is completed or the preset maximum number of steps is reached, the inner loop ends. The setting of the maximum number of steps in the task end condition can avoid getting stuck in an infinite loop for a certain task and ensure the efficiency and stability of the system.
[0066] The middle loop selects a task instruction from the current task pool, that is, selects the target task, and calls the task knowledge base module once. The task knowledge base module retrieves knowledge according to the task instruction and provides task planning knowledge for the agent. Based on these planning knowledge, the agent enters the inner loop to execute a single task. When selecting the task instruction, the system may comprehensively consider factors such as the priority and complexity of the task to ensure that the agent can complete the task efficiently.
[0067] The outer loop continuously repeats the middle loop. After completing the task learning of a batch, the self-driven exploration module is called once to update the task pool. The update of the task pool is to enable the agent to be exposed to more different types and difficulties of tasks, thereby continuously improving its learning ability and task execution ability. The updated task pool will contain new task instructions. Based on these new tasks, the agent re-enters the middle loop and starts a new round of learning and task execution.
[0068] Figure 1 is a flowchart of a data processing method shown according to an exemplary embodiment, as Figure 1 shown, the data processing method is used in an electronic device and includes the following steps.
[0069] In step S110, the target task and the task instruction of the target task are obtained.
[0070] Specifically, the target task includes at least one task operation; the target task is a task selected from the current task pool; the task instruction is used to describe the goal of the task. That the target task includes at least one task operation means that completing the target task includes at least one action performed on the GUI interface. For example, the target task can be a form filling task, and the corresponding task instruction can be to fill in form data in the current GUI interface; the task operation can include a click operation on the form input box displayed on the GUI interface, an information input operation for the form input box, and so on. The electronic device can execute each task operation in the execution order of each task operation of the target task.
[0071] In step S120, the first graphical user interface before performing the task operation and the second graphical user interface after performing the task operation are obtained, and annotation processing is performed through the target detection model to obtain an annotated graphical user interface annotated with text elements and image elements.
[0072] Among them, the target detection model can be an expert model. For example, the expert model can include a text detection model DBNet and an icon detection model Grounding-DINO; the text detection model can accurately detect the text area on the GUI interface, and the icon detection model can accurately detect the icon area on the GUI interface and identify different types of icons; the text element is the text area on the GUI interface, and the image element is each icon on the GUI interface. The annotated graphical user interface includes the annotated graphical user interface corresponding to the first GUI interface and the annotated graphical user interface corresponding to the second GUI interface.
[0073] Specifically, before performing the task operation of the target task, the electronic device can obtain the current GUI interface. After performing the task operation, the electronic device can collect the current GUI interface. The GUI interface collected before performing the task operation is used as the first graphical user interface, that is, the first GUI interface, and the GUI interface collected after performing the task operation is used as the second graphical user interface, that is, the second GUI interface. In this way, the electronic device can perform text annotation processing on the first GUI interface through a text detection model and perform icon annotation processing on the first GUI interface through an icon detection model to obtain a first GUI interface marked with text elements and icon elements, that is, the marked graphical user interface corresponding to the first GUI interface. Through a similar process, the electronic device can obtain a second GUI interface marked with text elements and icon elements, that is, the marked graphical user interface corresponding to the second GUI interface.
[0074] In step S130, based on the task instruction, the recorded data corresponding to the task operation, and the marked graphical user interface, operation inference is performed to obtain the task operation to be executed, and the task operation to be executed is executed until the preset task completion condition is met, and a graphical user interface for completing the target task is obtained.
[0075] Among them, the recorded data corresponding to the task operation can be data related to the task operation recorded by the electronic device after performing the task operation. For example, it can include action records and information records. The action record can include, for example, the execution order of the task operation and the execution result of the task operation. The information record can include the information added or reduced on the GUI interface when performing the task operation. The task operation to be executed can be the task operation that needs to be executed after performing the task operation. The preset task completion condition is used to determine whether the current operation to be executed needs to be executed and whether the target task is completed.
[0076] Specifically, the target task includes at least one task operation. After the electronic device executes one of the task operations through the agent model, the electronic device can obtain the first GUI interface before executing the task operation and the second GUI interface after executing the task operation, and perform annotation processing on the first GUI interface and the second GUI interface to obtain an annotated user graphical interface. In this way, the electronic device can reflect and perform operation planning based on the task instruction, the recorded data after executing the task operation, and the annotated user graphical interface, so as to obtain the optimal and most suitable operation to be executed for the current situation. In this way, when the preset task completion condition is not met, the electronic device can convert the task operation to be executed into an interactive operation on the second GUI interface through a pre-configured underlying actuator, and re-execute the method in the above embodiment until it is determined that the current preset task completion condition is met, and obtain a graphical user interface for executing the target task, which can be a graphical user interface for correctly executing the target task (the target task is successfully completed), or a graphical user interface for not correctly executing the target task (the target task is not successfully completed).
[0077] In the above data processing method, by using the target detection model to separately perform text and icon annotation on the user graphical interface before and after executing the task operation, the accuracy of detection and annotation can be improved, and the generality for each processing scenario can be enhanced. By performing data processing on the simply obtained visual data and avoiding obtaining underlying data, the convenience of obtaining annotated data is improved; by performing operation reasoning through the task instruction, the recorded data corresponding to the task operation, and the annotated graphical user interface to obtain the task operation to be executed, the context consistency during the execution of long-time sequential tasks can be ensured. By recording data, it is ensured that all key information related to the task is recorded, further ensuring the applicability to a variety of dynamically changing GUI application scenarios and improving the success rate of task execution.
[0078] In an exemplary embodiment, as Figure 2 shown, the method further includes:
[0079] In step 210, disassemble the target task to obtain multiple subtasks of the target task.
[0080] Among them, the complexity of each subtask of the target task is lower than the complexity of the target task.
[0081] Specifically, the electronic device can disassemble the target task to obtain multiple subtasks, or disassemble the task instruction of the target task to obtain multiple subtasks.
[0082] In step 220, encode the task instructions of each subtask to obtain multiple task instruction vectors.
[0083] Among them, the task instruction of the subtask can be the task planning text information of the subtask.
[0084] Specifically, the electronic device can encode the task instructions / task planning text information of each subtask through a preset encoding model to obtain task instruction vectors corresponding to each subtask, that is, text vectors. That is to say, the preset encoding model can be a text encoding model, and the electronic device can perform text encoding on the task planning text information / task instructions of each subtask through the text encoding model to obtain task instruction vectors corresponding to each subtask.
[0085] In step 230, in the preset task knowledge base, queries are respectively performed based on each task instruction vector to obtain multiple similar vectors that meet the first similarity condition, and based on the planning knowledge of each similar vector and the task instruction of the target task, a task planning suggestion is obtained.
[0086] Among them, the task planning suggestion is used to determine each task operation of the target task. For example, it can be to determine each task operation required to execute the target task, as well as the logical association and execution order between each task operation, etc. The preset task knowledge base can be a Faiss vector library, which is an efficient vector search library. Multiple vectors and the task planning knowledge corresponding to each vector are stored in this task knowledge base. The first similarity condition can be that the similarity is higher than a preset similarity threshold or the similarity is the top k similarities with the largest similarity, etc.
[0087] Specifically, after the electronic device obtains the task instruction vectors of each subtask, for the task instruction vector of each subtask, the similarity between the task instruction vector of the subtask and each vector in the preset task knowledge base is calculated respectively. In this way, the electronic device can obtain the similarities between the task instruction vectors of each subtask and each vector in the task knowledge base. The electronic device can sort them in descending order of similarity to obtain a similarity sequence. For example, the similarity greater than the preset threshold can be determined as the similarity that meets the first similarity condition, or the first k similarities can also be used as the similarities that meet the first similarity condition. Based on this, the electronic device can obtain each similarity that meets the first similarity condition and the vector corresponding to each similarity. This vector is the similar vector. The electronic device can obtain the task planning knowledge of each similar vector, and comprehensively combine the task instruction of the original target task and the task planning knowledge of each similar vector to obtain a task planning suggestion for the target task. The electronic device can obtain the next task after completing the target task in the current task pool based on this task planning suggestion, or can also determine each task operation required to complete the target task, as well as the logical association and execution order between each task operation based on this task planning suggestion.
[0088] Based on the above solution, by querying in the task knowledge base, the task planning experience associated with the target task is obtained, the probability of incorrect task execution is reduced, the task execution process is optimized, and the user experience of using the intelligent agent model for GUI interface interaction is further improved.
[0089] In one exemplary embodiment, as Figure 3 shown, the method further includes:
[0090] In step 310, determine the element image associated with the target task of the first graphical user interface.
[0091] Among them, the elements include text elements and / or icon elements.
[0092] Specifically, the electronic device can perform text analysis and icon analysis in the first image user interface to obtain multiple text elements and multiple icon elements. In this way, the electronic device can determine the degree of association between each element and the target task, and determine the element / element image associated with the target task.
[0093] In step 320, determine the text description data of each element, and encode the element image and the text description data through a multimodal model to obtain the multimodal vector of the first graphical user interface.
[0094] Among them, the multimodal model can be a pre-configured multimodal encoding model, such as a CLIP encoding model.
[0095] Specifically, the electronic device can determine the text description data of the text element and the text description data of the icon element, and perform multimodal encoding processing on the text description data of each text element and the text description data of each icon element through the CLIP encoding model to obtain the multimodal vector corresponding to the first graphical user interface. Optionally, the text description vector of the text element can be the text represented by the text element. For example, the text element can be the table name displayed on the first graphical user interface, etc.; the text description data of the icon element is used to represent the function of the icon. For example, the icon element can be an input box, and the corresponding text description data can be "text box for user input", etc.
[0096] In step 330, in the preset environment knowledge base, query based on the multimodal vector of the first graphical user interface to obtain the multimodal similarity vector that meets the second similarity condition, and determine the environment knowledge corresponding to the multimodal similarity vector.
[0097] Among them, the environment knowledge base stores multiple multimodal vectors and the corresponding environment knowledge. The second similarity condition can be that the similarity is higher than the preset similarity threshold or the similarity is the top k similarities with the largest similarity, etc.
[0098] Specifically, after the multi-modal vector of the first graphical user interface, the electronic device can calculate the similarity between the multi-modal vector of the first graphical user interface and each multi-modal vector in the preset environmental knowledge base. In this way, the electronic device can obtain the similarities between the multi-modal vector of the first graphical user interface and each multi-modal vector in the environmental knowledge base, arrange them in descending order of similarity to obtain a similarity sequence. For example, the similarities greater than a preset threshold can be determined as the similarities that meet the second similarity condition, or the first k similarities can also be used as the similarities that meet the second similarity condition. Based on this, the electronic device can obtain each similarity that meets the second similarity condition and the multi-modal vector corresponding to each similarity. This multi-modal vector is the multi-modal similarity vector, and the electronic device can obtain the environmental knowledge of each multi-modal similarity vector and determine it as the environmental knowledge of the multi-modal vector.
[0099] Based on the above solution, environmental knowledge can be queried through the multi-modal vector of the interface, which is beneficial for the intelligent agent model to quickly understand and adapt to new GUI interfaces and new icons, and further improve the adaptation degree of the intelligent agent model to each GUI interface.
[0100] In an exemplary embodiment, the recorded data includes action records and information records of performing task operations. As Figure 4 shown, based on the task instruction, the recorded data corresponding to the task operation, and the annotated graphical user interface for operation reasoning, the task operation to be executed can be obtained through the following steps:
[0101] In step 410, based on the task instruction, the action record and information record corresponding to the task operation, and the annotated graphical user interface for operation reasoning, the execution result of the task operation is obtained.
[0102] In step 420, if the execution result is a correct execution, then based on the task instruction, the action memory, information memory, annotated graphical user interface, and environmental knowledge corresponding to the task operation, the task operation to be executed is determined as the next task operation of the task operation.
[0103] In step 430, if the execution result is an incorrect execution, then the task operation to be executed is determined as re-executing the task operation.
[0104] Specifically, the electronic device can determine whether the current task operation has achieved a preset task effect, that is, obtain the execution result of the task operation, based on the task instruction of the current task operation, the action record of the current task operation, as well as the information record and the labeled GUI interface; optionally, the current task operation can be a pop-up window closing operation, and the execution result of the task operation can include correct execution and incorrect execution. Correspondingly, if the electronic device determines, based on the action record, information record, and labeled graphical user interface corresponding to the task operation, that "the pop-up window is not displayed", it can determine that the execution result of the current task operation is a correct execution; correspondingly, if the electronic device determines, based on the action record, information record, and labeled graphical user interface corresponding to the task operation, that "the pop-up window is displayed", it can determine that the execution result of the current task operation is an incorrect execution.
[0105] If it is determined that the current execution result is a correct execution, then the electronic device can plan the task operation based on the task instruction, the action memory, information memory, labeled graphical user interface, and environmental knowledge corresponding to the task operation, and generate the next task operation that is optimal and most in line with the task objective of the target task, that is, the electronic device can determine that the to-be-executed operation is the next task operation of the already-executed task operation; for example, if the current task operation is step i, the corresponding next task operation is step i + 1; if it is determined that the current execution result is an incorrect execution, then the electronic device can plan the task operation based on the task instruction, the action memory, information memory, labeled graphical user interface, and environmental knowledge, and the electronic device can determine that the task operation needs to be re-executed, that is, determine the already-executed task operation as the to-be-executed operation; for example, if the current task operation is step i, the corresponding next task operation is step i.
[0106] Optionally, the electronic device can determine the to-be-executed task operation through a basic agent module, which includes a reflection module, a recording module (memory module), and a decision-making module. The electronic device can evaluate the effect of the previous task operation through the reflection module, determine whether the previous task operation was successful, and whether it needs to be re-executed; if the electronic device determines that the previous task operation did not achieve the preset effect, the electronic device can determine that the to-be-executed task operation is to re-execute the previous task operation.
[0107] The memory module is divided into an action memory module and an information memory module. The action memory module records historical operations and their impacts, and is used to assist the agent in maintaining context consistency in long-term tasks. For example, in a task that requires multiple click and input operations, the action memory module can record the order and results of each operation, ensuring that the agent can execute the task according to the correct steps. The information memory module records key observation information, and when performing complex tasks, it can help the agent accurately remember key task information. For example, in a task that requires filling out a form, the information memory module can record the information in the form, preventing the agent model from repeatedly obtaining the same information; the decision-making module plans the next action based on the existing information, and it will generate an optimal operation plan according to factors such as task instructions, memory information, and observation results, and determine the task operation to be executed.
[0108] Based on the above solution, through task instructions, recorded data, labeled GUI interfaces, environmental knowledge, and task knowledge, it is possible to reflect before determining the operation to be executed, evaluate the execution effect of the previous operation. By recording the relevant data of task operations, it can help the agent model maintain context consistency in long-term tasks, avoid time sequence chaos, ensure the correctness of the execution of each task operation of the task, and determine the operation to be executed by integrating multiple pieces of information, further improving the adaptability of the agent model to each GUI interface and increasing the success execution probability of long-term tasks.
[0109] In an exemplary embodiment, the preset task completion conditions include that the step length of the task operation reaches a preset step length threshold, or the target task has been completed.
[0110] Specifically, the step length of the task operation can be the number of times of executing multiple tasks, or the number of task operations that have been executed; the preset step length threshold can be the maximum number of task operations that can be executed configured, and the present disclosure does not limit the specific value of the preset step length threshold, which can be determined based on the requirements of the actual application scenario. That the target task has been completed can mean that the task goal of the target task has been achieved or reached. For example, if the target task is a pop-up window closing task, then when the pop-up window on the current GUI interface of the electronic device has been closed, it can be determined that the target task has been completed.
[0111] Based on the above solution, by limiting the preset task completion conditions, it is possible to terminate the execution of the task in a timely manner when the task goal is reached, and avoid the agent model from falling into a loop in a certain task operation, ensuring the reliability and stability of task execution, and also ensuring the execution efficiency of the task.
[0112] In an exemplary embodiment, the target task is determined in the current task pool. As Figure 5 shown, the method further includes:
[0113] In step 510, after all tasks in the current task pool are completed, a first task list of successfully completed tasks and a second task list of unsuccessfully completed tasks are obtained.
[0114] Among them, the current task pool can be the task pool of the task batch to which the target task belongs. The current task pool contains multiple tasks. The electronic device selects the target task in the current task pool based on the execution order and complexity of each task.
[0115] Specifically, after the electronic device completes the execution of the target task, it can obtain the execution result of the target task, and then re-select a new target task in the current task pool where the target task is located, and based on the new target task, re-execute the method described in the above embodiments until all tasks in the current task pool are completed. In this way, the electronic device can distinguish each task based on the execution results of each task, and obtain a first task list of successfully completed tasks and a second task list of unsuccessfully completed tasks. Optionally, after the electronic device completes the execution of the target task, it can obtain the execution result of the target task. If the execution result is successfully completed, the target task is added to the first task list. If the execution result indicates that it is not successfully completed, the target task is added to the second task list; at the same time, in the current task pool where the target task is located, a new target task is selected, and based on the new target task, the method described in the above embodiments is re-executed until all tasks in the current task pool are completed.
[0116] In step 520, each task in the first task list is combined and / or modified to obtain multiple new tasks, and each task in the second task list is split to obtain multiple subtasks.
[0117] Among them, the complexity of the new tasks is higher than that of the tasks in the first task list, and the complexity of the subtasks is lower than that of the tasks in the second task list.
[0118] Specifically, the electronic device can analyze the characteristics and rules of each task in the first task list, and combine the processing level of the current agent model to modify each task, or combine each task, or modify and combine each task to obtain multiple new tasks with higher complexity. Similarly, the electronic device can split and / or decompose each task in the second task list to obtain subtasks corresponding to each task, reducing the execution difficulty of the agent model.
[0119] Optionally, if the first task list is multiple form-filling tasks that have been successfully completed, then after the electronic device modifies and combines them, the new task obtained can be a new task that includes multiple forms to be filled and data association; if the tasks included in the second task list can be tasks that include multiple form-filling and submissions, then the multiple subtasks obtained by disassembling can be the form-filling tasks and form-submission tasks for each individual form.
[0120] In step 530, based on each new task and each subtask, an updated task pool is obtained, and based on the complexity of each new task and each subtask and the dependency relationships between the tasks, the execution priorities and execution orders of the tasks in the updated task pool are obtained.
[0121] Specifically, the electronic device can combine each new task and each subtask to obtain an updated task pool. In this way, the electronic device can, through the task sorting unit in the self-driven exploration module, based on the dependency relationships and the complexity of each task included in each updated task pool, obtain the execution priorities and execution orders of each task. For example, the electronic device can obtain the degree of association between each task and the original task respectively, adjust the execution order of each task according to each degree of association to obtain an initial task sequence, and adjust the initial sequence based on the complexity of each task to obtain a target task sequence to be executed.
[0122] Based on the above solution, by automatically updating the task pool, the transformation from simple tasks to complex tasks can be gradually completed, realizing the transformation from novice to expert. The unfinished tasks are gradually disassembled into smaller subtasks to increase the probability of successful task completion, and the successfully completed tasks are combined to generate new and more complex tasks to be executed, that is, to generate more challenging tasks, enabling the intelligent agent model to come into contact with various types of tasks, gradually improving the task processing level of the intelligent agent model, and realizing the autonomous learning and task execution of the intelligent agent model in a complex environment.
[0123] In an exemplary embodiment, the method further includes:
[0124] Obtain the task planning text information of each task, encode the task planning text information of each task respectively to obtain a planning vector, and add the planning vector corresponding to each task to a preset task knowledge base.
[0125] Among them, the task planning text information of a task can be historical experience data for executing the task or task description information. For example, if the task is for the user to close the advertisement information on the GUI interface, the task planning text information of this task can be to find the virtual option to close the advertisement and click on the virtual option, and the virtual option is below the advertisement content, etc.
[0126] Specifically, after the electronic device finishes executing a task, it can obtain the execution trajectory of completing the task, divide the execution trajectory into multiple sub-trajectories, and inversely calculate the possible execution targets of each sub-trajectory based on each sub-trajectory. Based on the possible execution targets of each sub-trajectory, each sub-trajectory, and the execution trajectory of the task, task planning experience is obtained, that is, the task planning text information of the task is obtained. In this way, the electronic device can perform encoding processing on the task planning text information of the task through a preset encoding model to obtain a task instruction vector corresponding to the task, that is, a text vector; where the preset encoding model can be a text encoding model, and the electronic device can perform text encoding on the task planning text information of each sub-task through the text encoding model to obtain the task instruction vector of the task. The electronic device adds the task planning vectors corresponding to the obtained multiple tasks to a preset database to obtain a preset task knowledge base.
[0127] Based on the above solution, by converting the task planning text information into a text vector, it is convenient for data storage and data retrieval. The vectorized storage method can effectively utilize the efficient retrieval ability of the vector database, further improve the storage efficiency and management efficiency of the task planning knowledge of each task, provide a reliable data basis for the subsequent automated and intelligent processing of the GUI interface, and realize knowledge sharing.
[0128] In an exemplary embodiment, as Figure 6 shown, the method further includes:
[0129] In step 610, based on the first graphical user interface and the second graphical user interface corresponding to the task operation, determine the graphical user interface difference data after executing the task operation.
[0130] Specifically, the electronic device can obtain the GUI interface before the underlying actuator executes the task operation, that is, the first graphical user interface, and the GUI interface after executing the task operation, that is, the second graphical user interface. Based on the first graphical user interface and the second graphical user interface, determine different text areas, icon areas, etc. to obtain the graphical user interface difference data.
[0131] In step 620, based on the graphical user interface difference data and the task operation, obtain environmental knowledge.
[0132] Specifically, the electronic device can obtain environmental knowledge based on the graphical user interface difference data and the task operations that generate the difference data. For example, the task operation can be a pop-up window closing option, and if the first graphical user interface is the same as the second graphical user interface, the corresponding environmental knowledge can be "this closing option cannot close the pop-up window"; optionally, the task operation can be a pop-up window closing option, and the difference data between the first graphical user interface and the second graphical user interface can be that the pop-up window is displayed on the first graphical user interface and not displayed on the second graphical user interface, then the corresponding environmental knowledge can be "this closing option can close the pop-up window".
[0133] In step 630, take screenshots of each element on the first graphical user interface to obtain element images and the text description data corresponding to each element image. Encode the element images and the text descriptions through a multimodal model to obtain the multimodal vector of the first graphical user interface.
[0134] Among them, each element on the first graphical user interface can include text elements within the text area and each icon element. The multimodal model can be a pre-configured multimodal encoding model, such as a CLIP encoding model.
[0135] Specifically, the electronic device can perform image analysis and text analysis on the first graphical user interface to obtain each element. For example, it can obtain each text element and each icon element. The electronic device can obtain the text description data corresponding to each element, and perform multimodal encoding processing on the text description data of each text element and the text description data of each icon element through the CLIP encoding model to obtain the multimodal vector corresponding to the first graphical user interface. Optionally, the text description vector of the text element can be the text represented by the text element. For example, the text element can be the table name displayed on the first graphical user interface, etc.; the text description data of the icon element is used to represent the function of the icon. For example, the icon element can be an input box, and the corresponding text description data can be "text box for user input", etc.
[0136] In step 640, add the multimodal vector of the first graphical user interface corresponding to each task operation and the environmental knowledge to the environmental knowledge base.
[0137] Specifically, the electronic device can obtain the multi-modal vectors corresponding to each task operation and the environmental knowledge obtained after inductive processing. In this way, the electronic device can combine the first graphical user interface of each task with the corresponding multi-modal vector to form a data pair, and uniformly add multiple data pairs to a preset environmental knowledge base. The multi-modal vector maps the text information and icon information of the first graphical user interface to the same feature space. The data pair is a key-value structure, and the preset environmental knowledge base can be a Python dictionary structure.
[0138] Based on the above solution, by using the CLIP encoding model to extract multi-modal features, data in different dimensions can be mapped to the same feature space, and the functional descriptions of various elements on the graphical user interface can be accurately extracted. Storing knowledge through a dictionary structure facilitates subsequent data retrieval, improves the storage efficiency and management efficiency of environmental knowledge, and provides a reliable data basis for more intelligent GUI automation processing in the future.
[0139] The following describes the specific implementation process of the above data processing method in detail in combination with a specific embodiment:
[0140] The data processing system corresponding to the data processing method provided in this embodiment includes at least an observation processing module, a basic agent module, an environmental knowledge base module, a task knowledge base module, and a self-driven exploration module.
[0141] The observation processing module is used to process the information of the original screen screenshot of the GUI. In practical applications, the underlying document information (such as HTML, XML, etc.) of the GUI interface may vary due to different application programs and different versions, which brings great difficulties to information processing. To improve the generality and accuracy of processing, this module adopts a pure vision solution to avoid relying on the underlying document information.
[0142] The object detection model (expert model) of this module includes lightweight models, namely a text detection model (DBNet) and an icon detection model (Grounding-DINO). DBNet is a text detection model that can quickly and accurately detect the text areas in the screen screenshot. Grounding-DINO is used for icon detection and can identify various types of icons. Based on this expert model, the observation processing module can accurately detect and label the text and icon elements in the screen.
[0143] The input of the observation processing module is the unprocessed screen screenshot, and the output is the annotated screenshot in the format of SoM (Set-of-Mark). The SoM format contains the position information and serial number tags of the elements, which are beneficial to the reasoning and operation of the basic agent module. The observation processing module is processed by an expert model, and the positioning accuracy is significantly improved.
[0144] As Figure 7 shown, the processing flow of the observation processing module can be: obtaining the GUI screen screenshot, parsing the image elements through Grounding-DINO, parsing the text elements through the text detection model, merging, removing duplicates, sorting, and annotating the parsed elements to obtain the GUI screen screenshot annotated with SoM. In the actual processing process, the original screen screenshot will first undergo a series of preprocessing operations, such as image enhancement and noise reduction, to improve the quality of the image. Then, the text detection model and the icon detection model will respectively detect the preprocessed image, and label the detected text and icon elements. Finally, the annotation information is integrated into the screenshot in the SoM format and output.
[0145] The basic agent module is responsible for reasoning and executing specific operations based on the task instructions, memory, observation results (SoM screen screenshot), task knowledge, and environmental knowledge to determine the operation to be executed. The functional architecture of this module consists of a reflection module, a memory module, and a decision-making module. The electronic device can evaluate the effect of the previous task operation through the reflection module, judge whether the previous task operation is successful, and whether it needs to be executed again; if the electronic device determines that the previous task operation does not meet the preset effect, the electronic device can determine that the task operation to be executed is to re-execute the previous task operation.
[0146] The memory module is divided into an action memory module and an information memory module. The action memory module records historical operations and their impacts, which is used to assist the agent in maintaining context consistency in long-term sequential tasks. For example, in a task that requires multiple click and input operations, the action memory module can record the order and results of each operation to ensure that the agent can execute the task according to the correct steps. The information memory module records key observation information, which can help the agent accurately remember key task information when performing complex tasks. For example, in a task that requires filling out a form, the information memory module can record the information in the form to prevent the agent model from repeatedly obtaining the same information; the decision-making module plans the next action by synthesizing the existing information. It will generate the optimal operation plan according to factors such as task instructions, memory information, and observation results to determine the task operation to be executed.
[0147] As Figure 8As shown, it can be the execution flow chart of the basic agent: a GUI screenshot marked by SOM, reflecting the previous action and plan based on the new and old screenshots, judging whether the steps are correct. If the steps are correct, update the memory information (record data), record the execution status of the latest step, and infer the next action (continue execution / re-execution); if the steps are incorrect, directly infer the next action (continue execution / re-execution); after obtaining the next action (the task operation to be executed), judge whether to terminate the action. If the action is not terminated, interact with the screen through the action executor to obtain a new screenshot, that is, execute the task operation to be executed through the action executor to obtain the GUI interface after executing the task operation to be executed.
[0148] The environmental knowledge base module endows the agent with the ability to learn and utilize knowledge in a new environment and helps it adapt to unknown GUI elements. This module is designed based on the CLIP model and Python dictionary structure and has knowledge storage and retrieval functions.
[0149] In terms of storage operations, the multimodal large language model extracts the function information of GUI elements as "values". The multimodal large language model can comprehensively analyze image and text information and extract the functional descriptions of GUI elements. Then, the CLIP encodes the image and text description to generate a "key"; maps the image and text to the same feature space to obtain a multimodal vector. In this way, the key-value pair can be stored in the Python dictionary structure to obtain a preset environmental knowledge base. That is to say, the action executor interacts with the screen to obtain a new screenshot, and based on the new and old screenshots and the previous execution action, summarizes environmental knowledge; intercepts the element image, generates the element text description, uses CLIP to encode the image and text representations to obtain a multimodal vector, and adds each multimodal vector to the environmental knowledge base.
[0150] In the retrieval process of the environmental knowledge base, such as in the retrieval operation, according to the input image or text description, recall K entries with the highest similarity from the knowledge base. The calculation of similarity is based on the feature vectors generated by the CLIP model, and the similarity is determined by calculating the distance between the vectors. The recalled entries will be used by the basic agent module to help the agent better understand and process unknown GUI elements. For example Figure 9As shown in the figure, infer the element tags related to the task, intercept the element images, and generate text descriptions; use the CLIP vision model to encode the element image representations, that is, encode each intercepted element image to obtain the first encoding result; use the CLIP text model to encode the text description representations, that is, encode the text description data of each element image to obtain the second encoding result; based on the first encoding result and the second encoding result, obtain the multimodal vector. In the environmental knowledge base, retrieve the top K multimodal environmental knowledge with the highest similarity; obtain the environmental knowledge related to the current GUI screen.
[0151] The task knowledge base module improves the agent's task planning ability by summarizing task execution experience from trial and error. This module consists of a Faiss vector database and a text encoding model, and supports the storage and retrieval of knowledge related to task planning.
[0152] In terms of storage operations, encode the task planning text information into vectors and store them in the Faiss vector database to achieve fast vector retrieval and matching. The text encoding model is responsible for converting the task planning text information into vector representations for storage and retrieval, and can effectively utilize the efficient retrieval ability of the vector database to improve the storage and management efficiency of task planning knowledge. The storage process of the task knowledge base can include: dividing the execution trajectory into several sub-trajectories, backtracking the possible task goals of the sub-trajectories, summarizing the planning experience according to the sub-trajectories and the speculated tasks, summarizing the planning experience according to the total trajectory and the task instructions, and adding the obtained planning information and vectors to the task knowledge base together.
[0153] In terms of retrieval operations, after decomposing the task into subtasks, encode and retrieve them through the text encoding model. Generate planning suggestions by combining the retrieved information with the original task instructions. In practical applications, tasks may be relatively complex. By decomposing the task into subtasks, the complexity of the task can be reduced and the accuracy of task planning can be improved. As Figure 10 shown in the figure, based on the task instructions, decompose the task instructions to generate multiple-step subtasks, use the embedding model to generate a vector list of the tasks and subtasks, and in the task knowledge base, use the distance calculation function of Faiss to retrieve the top K task planning knowledge with the highest similarity to obtain the task planning knowledge related to the current task.
[0154] The self-driven exploration module is responsible for dynamically adjusting the exploration goals of the agent and promoting it to gradually complete the task learning from simple to complex. The operation logic of this module consists of a task decomposition module, a task combination module, and a task sorting module.
[0155] The task decomposition module decomposes unfinished tasks into smaller subtasks. In practical applications, some complex tasks may contain multiple steps and subtargets. Through the task decomposition module, these complex tasks can be decomposed into more manageable subtasks, reducing the learning difficulty of the agent. For example, a task that requires filling out and submitting multiple forms can be decomposed into multiple individual form-filling tasks and submission tasks.
[0156] The task combination module generates new and more complex exploration goals based on the completed tasks. It analyzes the characteristics and patterns of the completed tasks and combines them with the current ability level of the agent to generate challenging new tasks. For example, if the agent has completed multiple simple form-filling tasks, the task combination module can generate a complex task that includes multiple form-filling and data association operations.
[0157] The task sorting module adjusts the task priorities and execution order according to logical associations. It considers factors such as the dependencies and complexity between tasks to ensure that the agent can execute tasks in a reasonable order. For example, if a task needs to depend on the result of another task to execute, the task sorting module will arrange these two tasks in the correct order. As Figure 11 shown, it can be the execution process of the self-driven exploration module, determining whether to complete a batch of exploration tasks, that is, determining whether to complete the tasks in the current task pool. If the batch of exploration tasks is completed, a list of successfully and unsuccessfully completed tasks is obtained. Based on the unsuccessfully completed tasks, new subtasks are obtained through decomposition, and the successfully completed tasks are modified or combined to obtain more difficult tasks, that is, new tasks. The tasks are selected and sorted according to their inclusion relationships, and each subtask and each new task are selected and sorted to obtain an updated exploration task pool, that is, an updated task pool.
[0158] The data processing method provided in this embodiment can achieve the execution process through three nested loops in the self-driven exploration learning process of the agent model, as Figure 12 shown. During the actual operation process, the system will execute in the order of these three loops in sequence, continuously enabling the agent to learn and execute tasks.
[0159] The inner loop is the basis of the entire learning process, ensuring that the agent can efficiently execute a single task under a definite task instruction. The middle loop is responsible for selecting appropriate tasks from the task pool and providing task planning knowledge to the agent. The outer loop promotes the agent to continuously learn and explore. By updating the task pool, the agent is exposed to more different types of tasks, gradually enhancing its learning ability and task execution ability.
[0160] Through this nested loop design, the agent can gradually improve its task execution ability through continuous learning and practice. Starting from simple tasks, it gradually transitions to complex tasks, realizing the transformation from novice to expert. The self-driven exploration learning method provides an effective solution for the agent's autonomous learning and task execution in complex environments and has broad application prospects.
[0161] The data processing method provided in this embodiment directly analyzes GUI screenshots by combining multiple lightweight expert models, avoiding the limitations of traditional methods that rely on underlying document information (such as HTML, XML, etc.). At the same time, the expert model provides more refined element positioning information, and has higher action accuracy compared to the prediction method that directly uses multimodal large language models. By adopting a dual memory mechanism to record historical action information and key screen observation information respectively, the success rate of the multi-hop GUI screen information retrieval task is significantly improved. In addition, the reflection module dynamically checks and adjusts the agent's short-term actions and long-term plans, reducing the cumulative error phenomenon that occurs in long-time tasks and avoiding the dead loop problem caused by incorrect actions or plans. By adopting a multimodal storage and retrieval mechanism, the agent can gradually summarize the function information of new elements through interaction with the GUI screen even when encountering a new environment it has never seen before. For example, the agent can learn and record the meaning and function of new icons. Using the reverse inference mechanism, the agent can summarize task planning experience from action trajectories and failure trajectories. This mechanism greatly improves the efficiency of the agent's learning from interaction data, that is, under the same amount of trajectory data, it can summarize more valuable task planning experience. Through the two-way task automation generation and sorting mechanism, the agent can reasonably split complex tasks according to uncompleted tasks and preferentially learn subtasks, and at the same time gradually generate more complex task goals according to completed tasks. This module dynamically adjusts the learning order of tasks according to the dependency relationship of tasks, ensuring that the task difficulty is neither too low resulting in low learning efficiency nor too high resulting in difficult learning, thus realizing the adaptive adjustment of task difficulty.
[0162] It should be understood that although Figures 1 - 12 the steps in the flowchart of Figures 1 - 12 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0163] It is understandable that the same / similar parts among the various embodiments of the above methods in this specification can be referred to each other. Each embodiment focuses on the differences from other embodiments. For the related parts, refer to the descriptions of other method embodiments.
[0164] Figure 13 FIG. is a block diagram of a data processing apparatus 1300 shown according to an exemplary embodiment. Referring to Figure 13 , it includes:
[0165] A first acquisition unit 1302, configured to execute acquiring a target task and a task instruction of the target task, where the target task includes at least one task operation;
[0166] A second acquisition unit 1304, configured to execute acquiring a first graphical user interface before executing a task operation and a second graphical user interface after executing the task operation, and performing annotation processing through a target detection model to obtain an annotated graphical user interface annotated with text elements and image elements;
[0167] An inference unit 1306, configured to execute performing operation inference based on the task instruction, recorded data corresponding to the task operation, and the annotated graphical user interface to obtain a task operation to be executed, and executing the task operation to be executed until a preset task completion condition is satisfied, to obtain a graphical user interface for completing the execution of the target task.
[0168] In one embodiment, the apparatus further includes:
[0169] A first disassembling unit, configured to execute disassembling the target task to obtain a plurality of subtasks of the target task and task instructions of each of the subtasks;
[0170] A first encoding unit, configured to execute encoding the task instructions of each of the subtasks to obtain a plurality of task instruction vectors;
[0171] A first query unit, configured to execute querying in a preset task knowledge base respectively based on each of the task instruction vectors to obtain a plurality of similar vectors that meet a first similarity condition, and obtaining a task planning suggestion based on the planning knowledge of each of the similar vectors and the task instruction of the target task, where the task planning suggestion is used to determine each task operation of the target task.
[0172] In one embodiment, the apparatus further includes:
[0173] A first determination unit, configured to execute determining an element image associated with the target task of the first graphical user interface, where the element includes a text element and / or an icon element;
[0174] A second determination unit, configured to determine the text description data of each of the elements, and encode the element images and the text description data through a multimodal model to obtain a multimodal vector of the first graphical user interface;
[0175] A third determination unit, configured to perform a query in a preset environment knowledge base based on the multimodal vector of the first graphical user interface, obtain a multimodal similarity vector that meets the second similarity condition, and determine the environment knowledge corresponding to the multimodal similarity vector; multiple multimodal vectors and corresponding environment knowledge are stored in the environment knowledge base.
[0176] In one embodiment, the recorded data includes action records and information records of performing the task operation; the inference unit is specifically configured to perform:
[0177] Based on the task instruction, the action record and information record corresponding to the task operation, and the annotated graphical user interface, perform operation inference to obtain an execution result of the task operation;
[0178] If the execution result is correct execution, then based on the task instruction, the action memory, information memory, the annotated graphical user interface, and the environment knowledge corresponding to the task operation, determine that the task operation to be executed is the next task operation of the task operation;
[0179] If the execution result is not correctly executed, then determine that the task operation to be executed is to re - execute the task operation.
[0180] In one embodiment, the preset task completion condition includes that the step length of the task operation reaches a preset step length threshold, or the target task has been completed.
[0181] In one embodiment, the target task is determined in the current task pool; the apparatus further includes:
[0182] A fourth determination unit, configured to obtain a first task list that is successfully completed and a second task list that is not successfully completed after all tasks in the current task pool are completed;
[0183] A combination unit, configured to combine and / or modify each task in the first task list to obtain multiple new tasks, and split each task in the second task list to obtain multiple subtasks, the complexity of the new tasks is higher than the complexity of the tasks in the first task list, and the complexity of the subtasks is lower than the complexity of the tasks in the second task list;
[0184] An update unit, configured to execute to obtain an updated task pool based on each of the new tasks and each of the subtasks, and obtain the execution priority and execution order of each task in the updated task pool according to the complexity of each of the new tasks and each subtask and the dependency relationships between tasks.
[0185] In one embodiment, the apparatus further includes:
[0186] A third acquisition unit, configured to execute to acquire the task planning text information of each of the tasks, encode the task planning text information of each task respectively to obtain a planning vector, and add the planning vectors corresponding to each of the tasks to a preset task knowledge base.
[0187] In one embodiment, the apparatus further includes:
[0188] A fifth determination unit, configured to execute to determine the graphical user interface difference data after executing the task operation based on the first graphical user interface and the second graphical user interface corresponding to the task operation;
[0189] A sixth determination unit, configured to execute to obtain environmental knowledge based on the graphical user interface difference data and the task operation;
[0190] A second encoding unit, configured to execute to perform image screenshots on each element on the first graphical user interface to obtain element images and obtain text description data corresponding to each of the element images; encode the element images and the text description through the multimodal model to obtain the multimodal vector of the first graphical user interface;
[0191] An addition unit, configured to execute to add the multimodal vector of the first graphical user interface corresponding to each task operation and the environmental knowledge to an environmental knowledge base.
[0192] Regarding the apparatus in the above embodiments, the specific manners in which each module executes operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0193] Figure 14 It is a block diagram of an electronic device 1400 for a data processing method shown according to an exemplary embodiment. For example, the electronic device 1400 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0194] Refer to Figure 14, the electronic device 1400 may include one or more of the following components: a processing component 1402, a memory 1404, a power component 1406, a multimedia component 1408, an audio component 1410, an input / output (I / O) interface 1412, a sensor component 1414, and a communication component 1416.
[0195] The processing component 1402 generally controls the overall operation of the electronic device 1400, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 1402 may include one or more processors 1420 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 1402 may include one or more modules to facilitate the interaction between the processing component 1402 and other components. For example, the processing component 1402 may include a multimedia module to facilitate the interaction between the multimedia component 1408 and the processing component 1402.
[0196] The memory 1404 is configured to store various types of data to support the operation of the electronic device 1400. Examples of such data include instructions for any application or method operating on the electronic device 1400, contact data, phone book data, messages, pictures, videos, etc. The memory 1404 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, optical disks, or graphene memory.
[0197] The power component 1406 provides power to various components of the electronic device 1400. The power component 1406 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 1400.
[0198] The multimedia component 1408 includes a screen that provides an output interface between the electronic device 1400 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1408 includes a front camera and / or a rear camera. When the electronic device 1400 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0199] The audio component 1410 is configured to output and / or input audio signals. For example, the audio component 1410 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 1400 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1404 or transmitted via the communication component 1416. In some embodiments, the audio component 1410 further includes a speaker for outputting audio signals.
[0200] The I / O interface 1412 provides an interface between the processing component 1402 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0201] The sensor component 1414 includes one or more sensors for providing an assessment of the status of various aspects of the electronic device 1400. For example, the sensor component 1414 can detect the on / off state of the electronic device 1400, the relative positioning of components, such as the display and keypad of the electronic device 1400. The sensor component 1414 can also detect a change in the position of the electronic device 1400 or an electronic device 1400 component, the presence or absence of user contact with the electronic device 1400, the orientation or acceleration / deceleration of the device 1400, and a change in the temperature of the electronic device 1400. The sensor component 1414 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 1414 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 1414 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0202] The communication component 1416 is configured to facilitate communication between the electronic device 1400 and other devices in a wired or wireless manner. The electronic device 1400 can access a communication standard-based wireless network, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 1416 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1416 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0203] In an exemplary embodiment, the electronic device 1400 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0204] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as the memory 1404 including instructions, and the above instructions can be executed by the processor 1420 of the electronic device 1400 to complete the above method. For example, the computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0205] In an exemplary embodiment, a computer program product is also provided, and the computer program product includes instructions that can be executed by the processor 1420 of the electronic device 1400 to complete the above method.
[0206] It should be noted that the above-mentioned device, electronic device, computer-readable storage medium, computer program product, etc. may also include other implementation manners according to the description of the method embodiments. The specific implementation manners can refer to the description of the relevant method embodiments and will not be elaborated herein one by one.
[0207] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only to be considered exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0208] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A data processing method, characterized in that, Including: Obtain a target task and a task instruction for the target task, where the target task includes at least one task operation; Obtain a first graphical user interface before performing the task operation and a second graphical user interface after performing the task operation, and perform annotation processing through a target detection model to obtain an annotated graphical user interface annotated with text elements and image elements; Based on the task instruction, the recorded data corresponding to the task operation, and the annotated graphical user interface, perform operation reasoning to obtain a task operation to be executed, and execute the task operation to be executed until a preset task completion condition is met, obtaining a graphical user interface for successfully executing the target task.
2. The data processing method according to claim 1, wherein The method further includes: Decompose the target task into multiple subtasks of the target task; Encode the task instructions of each subtask to obtain multiple task instruction vectors; In a preset task knowledge base, query respectively based on each task instruction vector to obtain multiple similar vectors that meet the first similarity condition, and based on the planning knowledge of each similar vector and the task instruction of the target task, obtain a task planning suggestion, where the task planning suggestion is used to determine each task operation of the target task.
3. The data processing method according to claim 2, wherein The method further includes: Determine an element image of the first graphical user interface associated with the target task, where the element includes a text element and / or an icon element; Determine the text description data of each element, and encode the element image and the text description data through a multimodal model to obtain a multimodal vector of the first graphical user interface; In a preset environment knowledge base, query based on the multimodal vector of the first graphical user interface to obtain a multimodal similar vector that meets the second similarity condition, and determine the environment knowledge corresponding to the multimodal similar vector; multiple multimodal vectors and corresponding environment knowledge are stored in the environment knowledge base.
4. The data processing method according to claim 3, wherein The recorded data includes an action record and an information record for executing the task operation; the performing operation reasoning based on the task instruction, the recorded data corresponding to the task operation, and the annotated graphical user interface to obtain a task operation to be executed includes: Based on the task instruction, the action record and information record corresponding to the task operation, and the annotated graphical user interface, perform operation reasoning to obtain an execution result of the task operation; If the execution result is correct execution, then based on the task instruction, the action memory, information memory, the annotated graphical user interface, and the environment knowledge corresponding to the task operation, determine that the task operation to be executed is the next task operation of the task operation; If the execution result is incorrect execution, then determine that the task operation to be executed is to re-execute the task operation.
5. The data processing method according to claim 1, wherein The preset task completion condition includes that the step length of the task operation reaches a preset step length threshold, or the target task has been completed.
6. The data processing method according to claim 1, wherein The target task is determined in the current task pool; the method further includes: After each task in the current task pool is completed, obtain a first task list of successfully completed tasks and a second task list of unsuccessfully completed tasks; Combine and / or modify each task in the first task list to obtain multiple new tasks, and split each task in the second task list to obtain multiple subtasks. The complexity of the new tasks is higher than that of the tasks in the first task list, and the complexity of the subtasks is lower than that of the tasks in the second task list; Based on each of the new tasks and each of the subtasks, obtain an updated task pool, and according to the complexity of each of the new tasks and each subtask, and the dependency relationships between tasks, obtain the execution priorities and execution orders of each task in the updated task pool.
7. The method according to claim 2, wherein The method further includes: Obtain the task planning text information of each task, and encode the task planning text information of each task respectively to obtain a planning vector, and add the planning vector corresponding to each task to a preset task knowledge base.
8. The method according to claim 3, characterized in that, The method further includes: Based on the first graphical user interface and the second graphical user interface corresponding to the task operation, determine the graphical user interface difference data after executing the task operation; Based on the graphical user interface difference data and the task operation, obtain environmental knowledge; Take screenshots of each element on the first graphical user interface to obtain element images and obtain text description data corresponding to each of the element images; encode the element images and the text descriptions through the multimodal model to obtain the multimodal vector of the first graphical user interface; Add the multimodal vector of the first graphical user interface corresponding to each task operation and the environmental knowledge to the environmental knowledge base.
9. A data processing device, characterized in that, Includes: A first acquisition unit configured to execute acquiring a target task and a task instruction of the target task, where the target task includes at least one task operation; A second acquisition unit configured to execute acquiring a first graphical user interface before executing the task operation and a second graphical user interface after executing the task operation, and perform annotation processing through a target detection model to obtain an annotated graphical user interface annotated with text elements and image elements; An inference unit configured to execute performing operation inference based on the task instruction, the recorded data corresponding to the task operation, and the annotated graphical user interface to obtain a task operation to be executed, and execute the task operation to be executed until a preset task completion condition is met, to obtain a graphical user interface for completing the execution of the target task.
10. An electronic device, characterized in that, Includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the data processing method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the data processing method according to any one of claims 1 to 8.
12. A computer program product, comprising instructions, characterized in that, When the instructions are executed by the processor of the electronic device, the electronic device is enabled to execute the data processing method according to any one of claims 1 to 8.
Citation Information
Cited By
Interaction guiding method and device, storage medium and electronic equipment
CN120669887A
Interaction guiding method and device, storage medium and electronic device
CN120669887B
GUI task planning method, system and device and storage medium
CN121008725A