Method and system for improving energy efficiency of intelligent agent energy efficiency operation task

By building a UI operation agent framework and a diversified graphical interface operation dataset, training a multi-modal model to form a UI operation agent, and performing cross-system application cycle training and tuning, the problem of limited generalization in the existing technology of UI operation agents in multiple platforms and cross-application scenarios is solved, and the efficient positioning and navigation capabilities of the agent in benchmark tests are realized.

CN120010987AInactive Publication Date: 2025-05-16BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510475236.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology lacks sufficient data when developing general UI operation agents, resulting in limited generalization of the agents in multi-platform and cross-application scenarios, and it is impossible to effectively train powerful UI operation agents.

Method used

By building a UI operation agent framework, automatically collecting and processing multimodal UI operation tutorials, combining big data capture and processing, building a diversified graphical interface operation data set, training multimodal models, forming UI operation agents, and performing cross-system application cycle training and tuning to achieve the generalization ability of the agent.

Benefits of technology

The positioning and navigation energy efficiency of UI operation agents in benchmark tests has been significantly improved, with an improvement of about 10% compared to baseline agents, and its adaptability and energy efficiency in different operating systems and applications has been significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010987A_ABST
    Figure CN120010987A_ABST
Patent Text Reader

Abstract

The invention provides a method and a system for improving the energy efficiency of an intelligent agent effective operation task. A UI operation intelligent agent framework is constructed, a UI operation track is automatically formed to train a multi-modal model, and a multi-modal UI operation course is automatically collected and processed; in combination with a UI operation agent framework and big data capturing and processing, multi-mode UI operation courses are automatically collected and integrated, big data is automatically analyzed and mined, a UI operation track data set is formed in parallel, and a diversified graphical interface operation data set is constructed; according to the UI operation track data set and the diversified graphical interface operation data set, training a multi-modal model to form a UI operation agent; the cross-system application loop training and tuning UI operation agent is generalized across different operating systems and applications; and constructing a UI operation intelligent experience certificate framework, testing a UI operation agent on the UI operation evaluation test set, and synchronously verifying the positioning energy efficiency, the navigation energy efficiency and the multi-task parallel operation energy efficiency of the agent in the benchmark test.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of UI operation intelligent body positioning and navigation testing, and more specifically, to a method and system for improving the energy efficiency of intelligent body energy efficiency operation tasks. Background Art

[0002] The progress of AI intelligent big models, such as large language models (LLM) and multimodal big models (vision-language model, VLM), makes it possible to complete tasks based on model-based operation of UI interface; UI operation agents based on AI intelligent big models generally complete UI operation automation tasks by imitating human interaction with various applications; In reality, the lack of sufficient trajectory data is the main challenge in developing general UI operation agents; To solve the challenge of lack of data, existing methods mainly rely on manually annotated interaction trajectories or synthetic data from large open source or closed source models, but these methods have problems such as high cost and lack of diversity; For example, a large amount of manual annotation and synthetic data are used in ShowUI to verify the feasibility of such methods, but the effect is limited; There are abundant UI operation tutorials on the Internet, which provide detailed step-by-step instructions on how to control computer browsers, desktop applications, and smartphone applications, which is an underutilized resource; However, in reality, there is a lack of systems and technologies for collecting and processing such UI operation tutorials; UI operation agents mainly complete UI operation tasks by learning the form of human operation of UI; The primary challenge in developing UI operation agents is the lack of sufficient UI operation data. I operation data; most existing technologies solve this problem through manual labeling; although the manually labeled data has high quality, the labeling cost is too high so that the number of data sets cannot meet the standard of sufficient training; another way is to synthesize data through open source or closed source LLM and VLM; this method can generate a large enough amount of data; however, the diversity and accuracy of the data are reduced; in the work ShowUI by Lin et al., the above-mentioned multiple data collection methods are applied, and a relatively powerful model is developed; however, after analysis, the model energy efficiency of ShowUI is limited by the training data, and it cannot produce reasonable generalization in multi-platform and cross-application scenarios, which in turn limits the availability of this technology; powerful UI operation agents often need to work in cross-platform and cross-application scenarios, and the existing data sets have no diversity and inaccurate content, resulting in the inability to train such agents. Problems remain to be solved; therefore, it is necessary to propose a method and system for improving the energy efficiency of intelligent agent operation tasks to at least partially solve the problems existing in the prior art. Summary of the invention

[0003] A series of simplified concepts are introduced in the content of the invention, which will be further described in detail in the specific implementation method section; the content of the invention of the present invention does not mean to attempt to limit the key features and essential technical features of the technical solution claimed for protection, nor does it mean to attempt to determine the scope of protection of the technical solution claimed for protection.

[0004] To at least partially solve the above problems, the present invention provides a method for improving the energy efficiency of an intelligent agent energy efficiency operation task, comprising: S10, build a UI operation agent framework, automatically generate UI operation trajectories to train multimodal models, and automatically collect and process multimodal UI operation tutorials; S20, combined with the UI operation agent framework and big data capture and processing, automatically collects and integrates multimodal UI operation tutorials and processes big data analysis and mining, parallelizes them into UI operation trajectory data sets, and constructs a diverse graphical interface operation data set; S30, training a multimodal model based on the UI operation trajectory data set and the diversified graphical interface operation data set to form a UI operation agent; cross-system application loop training and tuning the UI operation agent to generalize across different operating systems and applications; S40, builds a UI operation agent verification architecture, tests the UI operation agent on the UI operation evaluation test set, and simultaneously verifies the agent's positioning energy efficiency, navigation energy efficiency, and multi-task parallel operation energy efficiency in the benchmark test.

[0005] Preferably, S10 includes: S101, through supervised precision adjustment, trains AI intelligent large models to learn operation trajectories; based on incremental learning, it continuously updates model parameters by capturing real-time features of data streams; S102, in the model update phase, uses real-time features to update model parameters in real time, and cyclically updates and optimizes the AI ​​intelligent big model; this provides a significantly enhanced data foundation for the AI ​​intelligent big model in solving UI operation type problems; The diverse graphical interface data set also includes actions: HotKey (keyboard shortcuts), Drag (hold down the left button and drag), Input (input characters), etc.; it provides a significantly enhanced data foundation for the AI ​​intelligent model in solving UI operation type problems.

[0006] Preferably, S20 includes: S201, select a UI operation task query word, and use the UI operation task query word in combination with a search interface to perform data expansion on the website; obtain a website link containing UI operation tutorial articles and videos; S202, by combining the UI operation agent framework and the multimodal tutorial for big data capture and processing, a diverse graphical interface operation dataset covering multiple operating systems and multiple applications under the operating systems was constructed.

[0007] Preferably, S30 includes: S301, based on the UI operation trajectory data set and the diversified graphical interface operation data set, use supervised precision to adjust the large-scale parameter AI intelligent large model; the large-scale parameters include: 3B and 7B large-scale parameters; 3B and 7B represent 3bilion approximately 3 billion parameters and 7bilion approximately 7 billion parameters respectively; S302, continuously train and tune the AI ​​big model to form a UI operation agent, and perform cross-system application cyclic training and tuning of the UI operation agent to perform operation tasks and generalize applications across different operating systems.

[0008] Preferably, S40 includes: S401, constructing a UI operation agent verification framework, testing the UI operation agent on a UI operation evaluation test set, and simultaneously verifying the positioning efficiency and navigation efficiency of the agent in a benchmark test; S402, evaluating the cross-system navigation energy efficiency of the UI operation agent in multiple operating systems; verifying the single-task operation energy efficiency of the UI operation agent actually completing a task or the multi-task parallel operation energy efficiency of multiple tasks; Evaluate the cross-system navigation energy efficiency of the UI operation agent in multiple operating systems; verify the single-task operation energy efficiency of the UI operation agent to actually complete a task or the multi-task parallel operation energy efficiency of multiple tasks, including: in addition to the positioning energy efficiency of the model, it is also necessary to evaluate the navigation energy efficiency of the model in different operating systems; navigation energy efficiency generally refers to the energy efficiency of the UI operation agent to actually complete a task; on the classic mobile operating system test set, compared with the energy efficiency indicators of ShowUI, etc., the energy efficiency of the UI operation agent has been greatly improved in the case of 3B and 7B models.

[0009] The present invention provides a system for improving the energy efficiency of intelligent body energy efficiency operation tasks, comprising: UI operation agent framework subsystem: builds the UI operation agent framework, automatically generates UI operation trajectories to train multimodal models, and automatically collects and processes multimodal UI operation tutorials; The UI operation data subsystem combines the UI operation agent framework with big data capture and processing to automatically collect and integrate multimodal UI operation tutorials and perform big data analysis and mining, parallelizing them into UI operation trajectory data sets, and constructing a diverse graphical interface operation data set; The UI operation agent training subsystem trains a multimodal model based on the UI operation trajectory data set and a diverse set of graphical interface operation data to form a UI operation agent. It also performs cross-system application loop training and tuning to generalize the UI operation agent across different operating systems and applications. The agent verification architecture subsystem builds the UI operation agent verification architecture, tests the UI operation agent on the UI operation evaluation test set, and simultaneously verifies the agent's positioning energy efficiency, navigation energy efficiency, and multi-task parallel operation energy efficiency in the benchmark test.

[0010] Preferably, the UI operation agent framework subsystem includes: The supervised precision adjustment subsystem trains the AI ​​intelligent large model to learn the operation trajectory through supervised precision adjustment. Based on incremental learning, it continuously updates the model parameters by capturing the real-time characteristics of the data stream. The model update loop reinforcement subsystem uses real-time features to update model parameters in real time during the model update phase, and cyclically updates and optimizes the AI ​​intelligent big model. This provides a significantly enhanced data foundation for the AI ​​intelligent big model in solving UI operation type problems. The diverse graphical interface data set also includes actions: HotKey (keyboard shortcuts), Drag (hold down the left button and drag), Input (input characters), etc.; it provides a significantly enhanced data foundation for the AI ​​intelligent model in solving UI operation type problems.

[0011] Preferably, the UI operation data subsystem includes: The UI operation search expansion subsystem selects the UI operation task query word, and uses the UI operation task query word in combination with the search interface to expand the data on the website; obtains links to websites containing UI operation tutorial articles and videos; The multi-operation graphical interface dataset subsystem, by combining the UI operation agent framework and the multimodal tutorial of big data capture and processing, builds a diverse graphical interface operation dataset covering multiple operating systems and multiple applications under the operating systems.

[0012] Preferably, the UI operation agent training subsystem includes: The supervised precision adjustment subsystem uses supervised precision adjustment of large-scale parameter AI intelligent large models according to the UI operation trajectory data set and the diversified graphical interface operation data set; large-scale parameters include: 3B and 7B large-scale parameters; 3B and 7B represent 3bilion of approximately 3 billion parameters and 7bilion of approximately 7 billion parameters respectively; large-scale parameters include: 3B and 7B large-scale parameters; 3B and 7B represent 3bilion of approximately 3 billion parameters and 7bilion of approximately 7 billion parameters respectively; The UI operation intelligent agent subsystem continuously trains and optimizes the AI ​​intelligent large model to form a UI operation intelligent agent. It performs cross-system application cyclic training and optimization on the UI operation intelligent agent to perform operation tasks and generalize applications across different operating systems.

[0013] Preferably, the agent verification architecture subsystem includes: The UI operation agent verification architecture subsystem builds the UI operation agent verification architecture, tests the UI operation agent on the UI operation evaluation test set, and simultaneously verifies the positioning energy efficiency and navigation energy efficiency of the agent in the benchmark test; Evaluate subsystems in parallel across systems and evaluate the cross-system navigation efficiency of UI operation agents in multiple operating systems; verify the single-task operation efficiency of the UI operation agent in actually completing a task or the multi-task parallel operation efficiency of multiple tasks.

[0014] Evaluate the cross-system navigation energy efficiency of the UI operation agent in multiple operating systems; verify the single-task operation energy efficiency of the UI operation agent to actually complete a task or the multi-task parallel operation energy efficiency of multiple tasks, including: in addition to the positioning energy efficiency of the model, it is also necessary to evaluate the navigation energy efficiency of the model in different operating systems; navigation energy efficiency generally refers to the energy efficiency of the UI operation agent to actually complete a task; on the classic mobile operating system test set, compared with the energy efficiency indicators of ShowUI, etc., the energy efficiency of the UI operation agent has been greatly improved in the case of 3B and 7B models.

[0015] Compared with the prior art, the present invention has at least the following beneficial effects: The present invention provides a method and system for improving the energy efficiency of intelligent body operation tasks, constructing a UI operation intelligent body framework, automatically forming a multimodal model for UI operation trajectory training, and automatically collecting and processing multimodal UI operation tutorials; combining the UI operation intelligent body framework and big data capture and processing, automatically collecting and integrating multimodal UI operation tutorials and big data analysis and mining processing, parallelizing into UI operation trajectory data sets, and constructing a diversified graphical interface operation data set; training a multimodal model based on the UI operation trajectory data set and the diversified graphical interface operation data set to form a UI operation intelligent body; cross-system application loop training and tuning the UI operation intelligent body to generalize across different operating systems and applications; constructing a UI operation intelligent body verification framework, testing the UI operation intelligent body on a UI operation evaluation test set, and synchronously verifying the positioning energy efficiency and navigation energy efficiency of the intelligent body in a benchmark test and the multi-task parallel operation energy efficiency; being able to capture and process multimodal tutorials through big data, and constructing a graphical interface (UI) operation data set covering multiple (including five) operating systems and (200) multiple applications under the operating system. The data set is named GUI-Net data set (diversified graphical interface data set). The dataset contains a total of 143K (annotated) UI operation trajectory data, so that the multimodal model can smoothly operate the UI to complete complex tasks after learning; a UI operation agent framework is proposed. The framework automatically collects and integrates multimodal UI operation tutorials and processes big data analysis and mining, and parallelizes them into a UI operation trajectory data set to train the multimodal model; the multimodal model includes a text-image multimodal AI intelligent question-answering model, a text-speech recognition multimodal AI intelligent model, or a text-speech-image-video hybrid multimodal AI intelligent model; the final trained UI operation agent can be generalized across different operating systems and applications; a UI operation agent is developed. Based on a diverse graphical interface dataset, adjustments are made on the AI ​​intelligent model. A UI operation agent verification architecture is developed based on the adjusted model. Based on the verification architecture, it can be verified that the model significantly improves the positioning and navigation efficiency of the agent in common benchmark tests, which is about 10% higher than the baseline agent; a UI operation trajectory training dataset to improve the adaptability and energy efficiency of the UI operation agent in different operating systems and applications; a system for automatically collecting and processing UI tutorials publicly available on the Internet to continuously build and improve the UI operation dataset; a method for adjusting AI intelligent large models, and an agent construction method to support real-time completion of UI operation tasks with significantly improved efficiency; the present invention has important technical significance and significant effects.

[0016] The present invention describes a method and system for improving the energy efficiency of intelligent body energy efficiency operation tasks. Other advantages, objectives and features of the present invention will be reflected in part through the following description, and in part will also be understood by technical personnel in this field through research and practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is a diagram of an embodiment of a system architecture for improving the energy efficiency of intelligent body operation tasks described in the present invention.

[0018] Figure 2 This is another example diagram of a method and system for improving the energy efficiency of intelligent body energy efficiency operation tasks described in the present invention.

[0019] Figure 3 This is a comparative example diagram of the prior art of a method and system for improving the energy efficiency of intelligent body energy efficiency operation tasks described in the present invention. DETAILED DESCRIPTION

[0020] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments so that those skilled in the art can implement the invention with reference to the description. Figure 1-Figure 3 As shown, the present invention provides a method for improving the energy efficiency of intelligent body energy efficiency operation tasks, comprising: S10, build a UI operation agent framework, automatically generate UI operation trajectories to train multimodal models, and automatically collect and process multimodal UI operation tutorials; S20, combined with the UI operation agent framework and big data capture and processing, automatically collects and integrates multimodal UI operation tutorials and processes big data analysis and mining, parallelizes them into UI operation trajectory data sets, and constructs a diverse graphical interface operation data set; S30, training a multimodal model based on the UI operation trajectory data set and the diversified graphical interface operation data set to form a UI operation agent; cross-system application loop training and tuning the UI operation agent to generalize across different operating systems and applications; S40, builds a UI operation agent verification architecture, tests the UI operation agent on the UI operation evaluation test set, and simultaneously verifies the agent's positioning energy efficiency, navigation energy efficiency, and multi-task parallel operation energy efficiency in the benchmark test.

[0021] The principle and effect of the above technical solution are as follows: the present invention provides a method for improving the energy efficiency of intelligent agent operation tasks, including: constructing a UI operation intelligent agent framework, automatically forming a multimodal model for training UI operation trajectories, and automatically collecting and processing multimodal UI operation tutorials; combining the UI operation intelligent agent framework and big data capture and processing, automatically collecting and integrating multimodal UI operation tutorials and big data analysis and mining processing, parallelizing them into UI operation trajectory data sets, and constructing a diversified graphical interface operation data set; training a multimodal model based on the UI operation trajectory data set and the diversified graphical interface operation data set to form a UI operation intelligent agent; cross-system application loop training and tuning the UI operation intelligent agent to generalize across different operating systems and applications; constructing a UI operation intelligent agent verification framework, testing the UI operation intelligent agent on the UI operation evaluation test set, and synchronously verifying the positioning energy efficiency and navigation energy efficiency of the intelligent agent in the benchmark test and the multi-task parallel operation energy efficiency; being able to capture and process multimodal tutorials through big data, and construct a graphical interface (UI) operation data set covering multiple (including five) operating systems and (200) multiple applications under the operating system. The data set is named GUI-Net data set (diversified graphical interface data set). The dataset contains a total of 143K (annotated) UI operation trajectory data, so that the multimodal model can smoothly operate the UI to complete complex tasks after learning; a UI operation agent framework is proposed. The framework automatically collects and integrates multimodal UI operation tutorials and processes big data analysis and mining, and parallelizes them into a UI operation trajectory data set to train the multimodal model; the multimodal model includes a text-image multimodal AI intelligent question-answering model, a text-speech recognition multimodal AI intelligent model, or a text-speech-image-video hybrid multimodal AI intelligent model; the final trained UI operation agent can be generalized across different operating systems and applications; a UI operation agent is developed; based on a diverse graphical interface dataset, it is adjusted on the AI ​​intelligent model. A UI operation agent verification architecture is developed based on the adjusted model. Based on the verification architecture, it can be verified that the model significantly improves the positioning and navigation efficiency of the agent in common benchmark tests, which is about 10% higher than the baseline agent; a UI operation trajectory training dataset to improve the adaptability and energy efficiency of the UI operation agent in different operating systems and applications; a system for automatically collecting and processing UI tutorials publicly available on the Internet to continuously build and improve the UI operation dataset; a method for adjusting AI intelligent large models, and an agent construction method to support real-time completion of UI operation tasks with significantly improved efficiency; the present invention has important technical significance and significant effects.

[0022] In one embodiment, S10 includes: S101, through supervised precision adjustment, trains AI intelligent large models to learn operation trajectories; based on incremental learning, it continuously updates model parameters by capturing real-time features of data streams; S102, in the model update phase, uses real-time features to update model parameters in real time, and cyclically updates and optimizes the AI ​​intelligent big model; this provides a significantly enhanced data foundation for the AI ​​intelligent big model in solving UI operation type problems; The diverse graphical interface data set also includes actions: HotKey (keyboard shortcuts), Drag (hold down the left button and drag), Input (input characters), etc.; it provides a significantly enhanced data foundation for the AI ​​intelligent model in solving UI operation type problems.

[0023] The principle and effect of the above technical solution are as follows: through supervised precision adjustment, the AI ​​intelligent big model is trained to learn the operation trajectory; based on incremental learning, the model parameters are continuously updated by capturing the real-time features of the data stream; in the model update stage, the model parameters are updated in real time using the real-time features, and the AI ​​intelligent big model is updated and optimized cyclically; this provides a significantly enhanced data foundation for the AI ​​intelligent big model in solving the UI operation type problem; The diverse graphical interface data set also includes actions: HotKey (keyboard shortcuts), Drag (hold down the left button to drag), Input (input characters), etc.; it provides a significantly enhanced data foundation for AI intelligent big models in solving UI operation type problems; such as Figure 3 As shown, some detailed distribution of diverse graphical interface datasets is shown; Figure 3 The distribution of operation traces of the five operating systems is shown in Figure a; the distribution of operation traces of the operating systems in the dataset; Figure 3 Figure b shows the task classification and the distribution of application categories in the dataset to show that the diverse graphical interface dataset covers a wide range of daily UI operations. Figure 3 In the figure c, the distribution of the number of task steps in the dataset is shown, which is the distribution of the number of operation trajectory steps. It can be seen that the longest operation step is up to 9 steps, which is enough to cover complex UI operation tasks. Figure 3 Figure d shows the distribution of action categories in the dataset, showing that the most common actions in the diverse graphical interface dataset are Click and Tap, corresponding to the click tasks of computers and mobile phones respectively. The diverse graphical interface dataset also includes actions such as HotKey (keyboard shortcuts), Drag (hold down the left button to drag), Input (input characters), etc.; this provides a significantly enhanced data foundation for AI intelligent large models in solving UI operation type problems.

[0024] In one embodiment, S20 includes: S201, select a UI operation task query word, and use the UI operation task query word in combination with a search interface to perform data expansion on the website; obtain a website link containing UI operation tutorial articles and videos; S202, by combining the UI operation agent framework and the multimodal tutorial for big data capture and processing, a diverse graphical interface operation dataset covering multiple operating systems and multiple applications under the operating systems was constructed.

[0025] The principle and effect of the above technical solution are: select UI operation task query words, and use the UI operation task query words in combination with the search interface to expand data on the website; obtain links to websites containing UI operation tutorial articles and videos; and build a diverse graphical interface operation data set covering multiple operating systems and multiple applications under the operating systems by combining the UI operation agent framework and the big data capture and processing multimodal tutorials; By combining the UI operation agent framework and big data capture and processing multimodal tutorials, a diverse graphical interface operation dataset covering multiple operating systems and multiple applications under the operating systems is constructed, including: big data capture of articles and videos on these web pages based on big data capture technology; since article-type web pages generally have a certain article structure, it is relatively easy to align and parse them into UI operation trajectories. However, for video tutorials, automatic speech recognition (Automatic Speech Recognition, ASR) and key frame detection are required. The ASR results of key frames and those near key frames will be used as a step in the UI operation trajectory; finally, the video tutorials and article data are organized into agent operation trajectory data, and the AI ​​intelligent big model is allowed to learn the operation trajectory through supervised fine-tune (Supervised Fine-tune); a large number of video article tutorials are collected through the UI operation agent framework; such as Figure 2 As shown, these tutorials are divided into 5 operating systems, covering various types of operating systems; the dataset includes 143K operation trajectories for supervised precision adjustment; a diverse graphical interface dataset (GUI-Net dataset) is constructed, so that the multimodal model can smoothly operate the UI to complete complex tasks after learning.

[0026] In one embodiment, S30 includes: S301, based on the UI operation trajectory data set and the diversified graphical interface operation data set, use supervised precision to adjust the large-scale parameter AI intelligent large model; the large-scale parameters include: 3B and 7B large-scale parameters; 3B and 7B represent 3bilion approximately 3 billion parameters and 7bilion approximately 7 billion parameters respectively; S302, continuously train and tune the AI ​​big model to form a UI operation agent, and perform cross-system application cyclic training and tuning of the UI operation agent to perform operation tasks and generalize applications across different operating systems.

[0027] The principle and effect of the above technical solution are as follows: according to the UI operation trajectory data set and the diversified graphical interface operation data set, supervised precision adjustment of large-scale parameter AI intelligent big model is used; large-scale parameters include: 3B and 7B large-scale parameters; 3B and 7B represent 3bilion of about 3 billion parameters and 7bilion of about 7 billion parameters respectively; continuous training and tuning of AI intelligent big model is used to form UI operation intelligent agent, and cross-system application cycle training and tuning of UI operation intelligent agent is carried out across different operating systems to perform operation tasks and applications for generalization; based on diversified graphical interface data sets, supervised precision adjustment of large-scale parameter AI intelligent big model is used; large-scale parameters include: 3B and 7B large-scale parameters; 3B and 7B represent respectively The 3bilion has about 3 billion parameters and the 7bilion has about 7 billion parameters. The test was conducted on an evaluation test set of common UI operations, and improvements were made in various energy efficiency indicators. First, a UI operation agent needs to be able to locate UI interface elements based on text or button labels. The positioning energy efficiency of the model was tested based on a high-resolution benchmark energy efficiency test set for positioning energy efficiency. Compared with ShowUI, the model has achieved great energy efficiency improvements on the same training set and model size. Compared with UI-TARS, with only one-fortieth of the training data, it is only slightly behind in energy efficiency indicators. The positioning and navigation energy efficiency of the agent in commonly used benchmark tests is improved by about 10% compared with the baseline agent.

[0028] In one embodiment, S40 includes: S401, constructing a UI operation agent verification framework, testing the UI operation agent on a UI operation evaluation test set, and simultaneously verifying the positioning efficiency and navigation efficiency of the agent in a benchmark test; S402, evaluating the cross-system navigation energy efficiency of the UI operation agent in multiple operating systems; verifying the single-task operation energy efficiency of the UI operation agent actually completing a task or the multi-task parallel operation energy efficiency of multiple tasks; Evaluate the cross-system navigation energy efficiency of the UI operation agent in multiple operating systems; verify the single-task operation energy efficiency of the UI operation agent to actually complete a task or the multi-task parallel operation energy efficiency of multiple tasks, including: in addition to the positioning energy efficiency of the model, it is also necessary to evaluate the navigation energy efficiency of the model in different operating systems; navigation energy efficiency generally refers to the energy efficiency of the UI operation agent to actually complete a task; on the classic mobile operating system test set, compared with the energy efficiency indicators of ShowUI, etc., the energy efficiency of the UI operation agent has been greatly improved in the case of 3B and 7B models.

[0029] The principle and effect of the above technical solution are: constructing a UI operation agent verification framework, testing the UI operation agent on the UI operation evaluation test set, and simultaneously verifying the positioning energy efficiency and navigation energy efficiency of the agent in the benchmark test; evaluating the cross-system navigation energy efficiency of the UI operation agent in multiple operating systems; verifying the single-task operation energy efficiency of the UI operation agent to actually complete a task or the multi-task parallel operation energy efficiency of multiple tasks; Evaluate the cross-system navigation energy efficiency of UI operation agents in multiple operating systems; verify the single-task operation energy efficiency of UI operation agents to actually complete a task or the multi-task parallel operation energy efficiency of multiple tasks, including: in addition to the positioning energy efficiency of the model, it is also necessary to evaluate the navigation energy efficiency of the model in different operating systems; navigation energy efficiency generally refers to the energy efficiency of UI operation agents to actually complete a task; on the classic mobile operating system test set, compared with the energy efficiency indicators of ShowUI, the energy efficiency of UI operation agents has been greatly improved in the case of 3B and 7B models; In addition to mobile phones, navigation efficiency on browsers is also very important for UI operation agents. Tests were conducted on the reference dataset for measuring the Internet access capability of large AI models in the browser operation test set. The agent was ahead of ShowUI in all aspects on the reference dataset for measuring the Internet access capability of large AI models, and achieved greater improvements on more difficult Cross-Domain tasks. Most of the test data sets are in English, and the industry lacks test sets that can test navigation tasks in a Chinese environment. To solve this problem, 102 UI operation trajectory data were annotated and tested on the UI operation agent. It was found that the energy efficiency of the UI operation agent was significantly improved compared with similar work in the industry. The test set not only reflects the evaluation of intelligence in offline conditions; the agent completes the task by performing UI operations on the real web page test set MiniWob online browser web page; the UI operation agent achieved the highest scores on both 3B and 7B; proving that the UI operation agent can complete UI operation tasks on a real web browser.

[0030] The present invention provides a system for improving the energy efficiency of intelligent body energy efficiency operation tasks, comprising: UI operation agent framework subsystem: builds the UI operation agent framework, automatically generates UI operation trajectories to train multimodal models, and automatically collects and processes multimodal UI operation tutorials; The UI operation data subsystem combines the UI operation agent framework with big data capture and processing to automatically collect and integrate multimodal UI operation tutorials and perform big data analysis and mining, parallelizing them into UI operation trajectory data sets, and constructing a diverse graphical interface operation data set; The UI operation agent training subsystem trains a multimodal model based on the UI operation trajectory data set and a diverse set of graphical interface operation data to form a UI operation agent. It also performs cross-system application loop training and tuning to generalize the UI operation agent across different operating systems and applications. The agent verification architecture subsystem builds the UI operation agent verification architecture, tests the UI operation agent on the UI operation evaluation test set, and simultaneously verifies the agent's positioning energy efficiency, navigation energy efficiency, and multi-task parallel operation energy efficiency in the benchmark test.

[0031] The principle and effect of the above technical solution are as follows: The present invention provides a system for improving the energy efficiency of intelligent body energy efficiency operation tasks, including: a UI operation intelligent body framework subsystem, which constructs a UI operation intelligent body framework, automatically forms a multimodal model for UI operation trajectory training, and automatically collects and processes multimodal UI operation tutorials; a UI operation data subsystem, which combines the UI operation intelligent body framework and big data capture and processing, automatically collects and integrates multimodal UI operation tutorials and performs big data analysis and mining, and parallelizes them into UI operation trajectory data sets, and constructs a diversified graphical interface operation data set; a UI operation intelligent body training subsystem, according to the UI operation trajectory data set and diversified graphical interface operation data sets, train multimodal models, and form UI operation agents; cross-system application loop training and tuning of UI operation agents generalize across different operating systems and applications; agent verification architecture subsystem, build UI operation agent verification architecture, test UI operation agents on UI operation evaluation test sets, and simultaneously verify the positioning energy efficiency and navigation energy efficiency of agents in benchmark tests and multi-task parallel operation energy efficiency; can capture and process multimodal tutorials through big data, and build a graphical interface (UI) operation data set covering multiple (including five) operating systems and (200) more than applications under the operating system. The data set is named GUI-Net data set (diverse graphical interface data set). The data set contains a total of 143K (annotated) UI operation trajectory data, so that the multimodal model can smoothly operate the UI to complete complex tasks after learning; a UI operation agent framework is proposed. The framework automatically collects and integrates multimodal UI operation tutorials and processes big data analysis and mining, and parallelizes them into a UI operation trajectory data set to train a multimodal model; the multimodal model includes a text-image multimodal AI intelligent question-answering model, a text-speech recognition multimodal AI intelligent model, or a text-speech-image-video hybrid multimodal AI intelligent model; the final trained UI operation agent can be generalized across different operating systems and applications; a UI operation agent is developed; based on a diverse graphical interface data set, adjustments are made on the AI ​​intelligent model. A UI operation agent verification framework is developed based on the adjusted model. Based on the verification framework, it can be verified that the model significantly improves the positioning and navigation energy efficiency of the agent in common benchmark tests, which is about 10% higher than the baseline agent; a UI operation trajectory training data set to improve the adaptability and energy efficiency of the UI operation agent in different operating systems and applications; a system for automatically collecting and processing online public UI tutorials to continuously build and improve UI operation data sets; a method for adjusting the AI ​​intelligent model, and an agent construction method to support the real-time completion of UI operation tasks with significantly improved efficiency; the present invention has important technical significance and significant effects.

[0032] In one embodiment, the UI operation agent framework subsystem includes: The supervised precision adjustment subsystem trains the AI ​​intelligent large model to learn the operation trajectory through supervised precision adjustment. Based on incremental learning, it continuously updates the model parameters by capturing the real-time characteristics of the data stream. The model update loop reinforcement subsystem uses real-time features to update model parameters in real time during the model update phase, and cyclically updates and optimizes the AI ​​intelligent big model. This provides a significantly enhanced data foundation for the AI ​​intelligent big model in solving UI operation type problems. The diverse graphical interface data set also includes actions: HotKey (keyboard shortcuts), Drag (hold down the left button and drag), Input (input characters), etc.; it provides a significantly enhanced data foundation for the AI ​​intelligent model in solving UI operation type problems.

[0033] The principle and effect of the above technical solution are as follows: UI operation agent framework subsystem, including: The supervised precision adjustment subsystem trains the AI ​​intelligent large model to learn the operation trajectory through supervised precision adjustment. Based on incremental learning, it continuously updates the model parameters by capturing the real-time characteristics of the data stream. The model update loop reinforcement subsystem uses real-time features to update model parameters in real time during the model update phase, and cyclically updates and optimizes the AI ​​intelligent big model. This provides a significantly enhanced data foundation for the AI ​​intelligent big model in solving UI operation type problems. The diverse graphical interface data set also includes actions: HotKey (keyboard shortcuts), Drag (hold down the left button to drag), Input (input characters), etc.; it provides a significantly enhanced data foundation for AI intelligent big models in solving UI operation type problems; such as Figure 3 As shown, some detailed distribution of diverse graphical interface datasets is shown; Figure 3 The distribution of operation traces of the five operating systems is shown in Figure a; the distribution of operation traces of the operating systems in the dataset; Figure 3 Figure b shows the task classification and the distribution of application categories in the dataset to show that the diverse graphical interface dataset covers a wide range of daily UI operations. Figure 3 In the figure c, the distribution of the number of task steps in the dataset is shown, which is the distribution of the number of operation trajectory steps. It can be seen that the longest operation step is up to 9 steps, which is enough to cover complex UI operation tasks. Figure 3 Figure d shows the distribution of action categories in the dataset, showing that the most common actions in the diverse graphical interface dataset are Click and Tap, corresponding to the click tasks of computers and mobile phones respectively. The diverse graphical interface dataset also includes actions such as HotKey (keyboard shortcuts), Drag (hold down the left button to drag), Input (input characters), etc.; this provides a significantly enhanced data foundation for AI intelligent large models in solving UI operation type problems.

[0034] In one embodiment, the UI operation data subsystem includes: The UI operation search expansion subsystem selects the UI operation task query word, and uses the UI operation task query word in combination with the search interface to expand the data on the website; obtains links to websites containing UI operation tutorial articles and videos; The multi-operation graphical interface dataset subsystem, by combining the UI operation agent framework and the multimodal tutorial of big data capture and processing, builds a diverse graphical interface operation dataset covering multiple operating systems and multiple applications under the operating systems.

[0035] The principle and effect of the above technical solution are as follows: UI operation data subsystem includes: The UI operation search expansion subsystem selects the UI operation task query word, and uses the UI operation task query word in combination with the search interface to expand the data on the website; obtains links to websites containing UI operation tutorial articles and videos; The multi-operation graphical interface dataset subsystem builds a diverse graphical interface operation dataset covering multiple operating systems and multiple applications under the operating systems by combining the UI operation agent framework and the multimodal tutorial on big data capture and processing; By combining the UI operation agent framework and big data capture and processing multimodal tutorials, a diverse graphical interface operation dataset covering multiple operating systems and multiple applications under the operating systems is constructed, including: big data capture of articles and videos on these web pages based on big data capture technology; since article-type web pages generally have a certain article structure, it is relatively easy to align and parse them into UI operation trajectories. However, for video tutorials, automatic speech recognition (Automatic Speech Recognition, ASR) and key frame detection are required. The ASR results of key frames and those near key frames will be used as a step in the UI operation trajectory; finally, the video tutorials and article data are organized into agent operation trajectory data, and the AI ​​intelligent big model is allowed to learn the operation trajectory through supervised fine-tune (Supervised Fine-tune); a large number of video article tutorials are collected through the UI operation agent framework; such as Figure 2 As shown, these tutorials are divided into 5 operating systems, covering various types of operating systems; the dataset includes 143K operation trajectories for supervised precision adjustment; a diverse graphical interface dataset (GUI-Net dataset) is constructed, so that the multimodal model can smoothly operate the UI to complete complex tasks after learning.

[0036] In one embodiment, the UI operates the agent training subsystem, including: The supervised precision adjustment subsystem uses supervised precision adjustment of large-scale parameter AI intelligent large models based on the UI operation trajectory data set and the diversified graphical interface operation data set; large-scale parameters include: 3B and 7B large-scale parameters; 3B and 7B represent 3bilion, about 3 billion parameters, and 7bilion, about 7 billion parameters, respectively; The UI operation intelligent agent subsystem continuously trains and optimizes the AI ​​intelligent large model to form a UI operation intelligent agent. It performs cross-system application cyclic training and optimization on the UI operation intelligent agent to perform operation tasks and generalize applications across different operating systems.

[0037] The principle and effect of the above technical solution are as follows: UI operation agent training subsystem includes: The supervised precision adjustment subsystem uses supervised precision adjustment of large-scale parameter AI intelligent large models based on the UI operation trajectory data set and the diversified graphical interface operation data set; large-scale parameters include: 3B and 7B large-scale parameters; 3B and 7B represent 3bilion, about 3 billion parameters, and 7bilion, about 7 billion parameters, respectively; The UI operation agent subsystem continuously trains and tunes the AI ​​intelligent big model to form a UI operation agent, and performs cross-system application cyclic training and tuning of the UI operation agent to perform operation tasks and generalize applications across different operating systems; based on a diverse graphical interface data set, a supervised precision adjustment of large-scale parameter AI intelligent big models is used; large-scale parameters include: 3B and 7B large-scale parameters; 3B and 7B represent 3bilion, about 3 billion parameters, and 7bilion, about 7 billion parameters, respectively; and tested on the evaluation test set of common UI operations, and achieved improvements in various energy efficiency indicators; first, a UI operation agent needs to have the energy efficiency to locate UI interface elements according to text or button labels; the positioning energy efficiency of the model is tested based on the positioning energy efficiency high-resolution benchmark energy efficiency test set; the model has achieved a great energy efficiency improvement compared to ShowUI on the same training set and model size, and compared to UI-TARS, it is only slightly behind in energy efficiency indicators when the training data is only one-fortieth; the positioning and navigation energy efficiency of the agent in common benchmark tests is improved by about 10% compared to the baseline agent.

[0038] In one embodiment, the agent verification architecture subsystem includes: The UI operation agent verification architecture subsystem builds the UI operation agent verification architecture, tests the UI operation agent on the UI operation evaluation test set, and simultaneously verifies the positioning energy efficiency and navigation energy efficiency of the agent in the benchmark test; Evaluate subsystems in parallel across systems and evaluate the cross-system navigation efficiency of UI operation agents in multiple operating systems; verify the single-task operation efficiency of the UI operation agent in actually completing a task or the multi-task parallel operation efficiency of multiple tasks.

[0039] Evaluate the cross-system navigation energy efficiency of the UI operation agent in multiple operating systems; verify the single-task operation energy efficiency of the UI operation agent to actually complete a task or the multi-task parallel operation energy efficiency of multiple tasks, including: in addition to the positioning energy efficiency of the model, it is also necessary to evaluate the navigation energy efficiency of the model in different operating systems; navigation energy efficiency generally refers to the energy efficiency of the UI operation agent to actually complete a task; on the classic mobile operating system test set, compared with the energy efficiency indicators of ShowUI, etc., the energy efficiency of the UI operation agent has been greatly improved in the case of 3B and 7B models.

[0040] The principle and effect of the above technical solution are as follows: The intelligent agent verification architecture subsystem includes: The UI operation agent verification architecture subsystem builds the UI operation agent verification architecture, tests the UI operation agent on the UI operation evaluation test set, and simultaneously verifies the positioning energy efficiency and navigation energy efficiency of the agent in the benchmark test; Evaluate subsystems in parallel across systems and evaluate the cross-system navigation efficiency of UI operation agents in multiple operating systems; verify the single-task operation efficiency of the UI operation agent in actually completing a task or the multi-task parallel operation efficiency of multiple tasks.

[0041] Evaluate the cross-system navigation energy efficiency of UI operation agents in multiple operating systems; verify the single-task operation energy efficiency of UI operation agents to actually complete a task or the multi-task parallel operation energy efficiency of multiple tasks, including: in addition to the positioning energy efficiency of the model, it is also necessary to evaluate the navigation energy efficiency of the model in different operating systems; navigation energy efficiency generally refers to the energy efficiency of UI operation agents to actually complete a task; on the classic mobile operating system test set, compared with the energy efficiency indicators of ShowUI, the energy efficiency of UI operation agents has been greatly improved in the case of 3B and 7B models; In addition to mobile phones, navigation efficiency on browsers is also very important for UI operation agents. Tests were conducted on the reference dataset for measuring the Internet access capability of large AI models in the browser operation test set. The agent was ahead of ShowUI in all aspects on the reference dataset for measuring the Internet access capability of large AI models, and achieved greater improvements on more difficult Cross-Domain tasks. Most of the test data sets are in English, and the industry lacks test sets that can test navigation tasks in a Chinese environment. To solve this problem, 102 UI operation trajectory data were annotated and tested on the UI operation agent. It was found that the energy efficiency of the UI operation agent was significantly improved compared with similar work in the industry. The test set not only reflects the evaluation of intelligence in offline conditions; the agent completes the task by performing UI operations on the real web page test set MiniWob online browser web page; the UI operation agent achieved the highest scores on both 3B and 7B; proving that the UI operation agent can complete UI operation tasks on a real web browser.

[0042] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and the implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and the illustrations shown and described herein.

Claims

1. A method for improving the energy efficiency of intelligent agent energy efficiency operation tasks, characterized in that: include: S10, build a UI operation agent framework, automatically generate UI operation trajectories to train multimodal models, and automatically collect and process multimodal UI operation tutorials; S20, combined with the UI operation agent framework and big data capture and processing, automatically collects and integrates multimodal UI operation tutorials and processes big data analysis and mining, parallelizes them into UI operation trajectory data sets, and constructs a diverse graphical interface operation data set; S30, training a multimodal model according to the UI operation trajectory data set and the diversified graphical interface operation data set to form a UI operation agent; Cross-system application loop training and tuning of UI operation agents to generalize across different operating systems and applications; S40, builds a UI operation agent verification architecture, tests the UI operation agent on the UI operation evaluation test set, and simultaneously verifies the agent's positioning energy efficiency, navigation energy efficiency, and multi-task parallel operation energy efficiency in the benchmark test.

2. A method for improving the energy efficiency of intelligent agent energy efficiency operation tasks according to claim 1, characterized in that: S10 includes: S101, through supervised precision adjustment, trains AI intelligent large models to learn operation trajectories; based on incremental learning, it continuously updates model parameters by capturing real-time features of data streams; S102, in the model updating stage, uses real-time features to update model parameters in real time, and cyclically updates and optimizes the AI ​​intelligent big model; this provides a significantly enhanced data foundation for the AI ​​intelligent big model in solving UI operation type problems.

3. A method for improving the energy efficiency of intelligent body energy efficiency operation tasks according to claim 1, characterized in that: The S20 includes: S201, select a UI operation task query word, and use the UI operation task query word in combination with a search interface to perform data expansion on the website; obtain a website link containing UI operation tutorial articles and videos; S202, by combining the UI operation agent framework and the multimodal tutorial for big data capture and processing, a diverse graphical interface operation dataset covering multiple operating systems and multiple applications under the operating systems was constructed.

4. A method for improving the energy efficiency of intelligent body energy efficiency operation tasks according to claim 1, characterized in that: S30 includes: S301, using supervised precision adjustment of large-scale parameter AI intelligent big model based on UI operation trajectory data set and diversified graphical interface operation data set; S302, continuously train and tune the AI ​​big model to form a UI operation agent, and perform cross-system application cyclic training and tuning of the UI operation agent to perform operation tasks and generalize applications across different operating systems.

5. A method for improving the energy efficiency of intelligent body energy efficiency operation tasks according to claim 1, characterized in that: S40 includes: S401, constructing a UI operation agent verification framework, testing the UI operation agent on a UI operation evaluation test set, and simultaneously verifying the positioning efficiency and navigation efficiency of the agent in a benchmark test; S402, evaluating the cross-system navigation efficiency of the UI operation agent in multiple operating systems; verifying the single-task operation efficiency of the UI operation agent in actually completing a task or the multi-task parallel operation efficiency of multiple tasks.

6. A system for improving the energy efficiency of intelligent body operation tasks, characterized in that: include: UI operation agent framework subsystem: builds the UI operation agent framework, automatically generates UI operation trajectories to train multimodal models, and automatically collects and processes multimodal UI operation tutorials; The UI operation data subsystem combines the UI operation agent framework with big data capture and processing to automatically collect and integrate multimodal UI operation tutorials and perform big data analysis and mining, parallelizing them into UI operation trajectory data sets, and constructing a diverse graphical interface operation data set; The UI operation agent training subsystem trains a multimodal model based on the UI operation trajectory data set and a diverse graphical interface operation data set to form a UI operation agent; Cross-system application loop training and tuning of UI operation agents to generalize across different operating systems and applications; The agent verification architecture subsystem builds the UI operation agent verification architecture, tests the UI operation agent on the UI operation evaluation test set, and simultaneously verifies the agent's positioning energy efficiency, navigation energy efficiency, and multi-task parallel operation energy efficiency in the benchmark test.

7. A system for improving energy efficiency of intelligent body operation tasks according to claim 6, characterized in that: The UI operation agent framework subsystem includes: The supervised precision adjustment subsystem trains the AI ​​intelligent large model to learn the operation trajectory through supervised precision adjustment. Based on incremental learning, it continuously updates the model parameters by capturing the real-time characteristics of the data stream. The model update loop reinforcement subsystem uses real-time features to update model parameters in real time during the model update phase, and cyclically updates and optimizes the AI ​​intelligent big model; it provides a significantly enhanced data foundation for the AI ​​intelligent big model in solving UI operation type problems.

8. The system for improving energy efficiency of intelligent body operation tasks according to claim 6, characterized in that: UI operation data subsystem, including: The UI operation search expansion subsystem selects UI operation task query words, and uses the UI operation task query words in combination with the search interface to expand data on the website; obtains links to websites containing UI operation tutorial articles and videos; The multi-operation graphical interface dataset subsystem, by combining the UI operation agent framework and the multimodal tutorial of big data capture and processing, builds a diverse graphical interface operation dataset covering multiple operating systems and multiple applications under the operating systems.

9. The system for improving energy efficiency of intelligent body operation tasks according to claim 6, characterized in that: UI operation agent training subsystem, including: The supervised precision adjustment subsystem uses supervised precision adjustment of large-scale parameter AI intelligent models based on the UI operation trajectory data set and the diversified graphical interface operation data set; The UI operation intelligent agent subsystem continuously trains and tunes the AI ​​intelligent large model to form a UI operation intelligent agent. It performs cross-system application cyclic training and tuning of the UI operation intelligent agent to perform operation tasks and generalize applications across different operating systems.

10. The system for improving energy efficiency of intelligent body operation tasks according to claim 6, characterized in that: The agent verification architecture subsystem includes: The UI operation agent verification architecture subsystem builds the UI operation agent verification architecture, tests the UI operation agent on the UI operation evaluation test set, and simultaneously verifies the positioning energy efficiency and navigation energy efficiency of the agent in the benchmark test; Evaluate subsystems in parallel across systems and evaluate the cross-system navigation efficiency of UI operation agents in multiple operating systems; verify the single-task operation efficiency of the UI operation agent in actually completing a task or the multi-task parallel operation efficiency of multiple tasks.

Citation Information

Cited By

  • Mobile application intelligent agent method

    CN121125830A

  • A Smart Agent Method for Mobile Applications

    CN121125830B