Model construction method, program product and storage medium for GUI intelligent body

By constructing personalized models of mobile GUI agents, including pre-training the base model, online collection and labeling of user operation trajectories, task planning training and the use of plug-in scene knowledge bases, the problems of insufficient adaptability and insufficient execution of high-frequency tasks in different usage scenarios are solved, and more efficient personalized adaptation and task execution are achieved.

CN119576470BActive Publication Date: 2025-05-13BEIJING LANZHOU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510136285.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-13
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

When existing mobile GUI agents respond to different usage scenarios, they lack the ability to adapt to the user's personalized usage environment and perform high-frequency tasks.

Method used

By pre-training based on the GUI field general data set, the base model is obtained; the user demonstration operation trajectory is collected online, and the user's demonstration operation trajectory is automatically marked, and the base model is obtained; the base model is task-planned and the training results are integrated to obtain the scene planner, and the base model and the scene planner are combined to obtain a personalized model; the plug-in scene knowledge base is obtained, supplementary scene information is provided and corresponding scene knowledge is retrieved during action planning.

Benefits of technology

It greatly improves the adaptability of GUI agents to the user's personalized usage environment and enhances the efficiency and accuracy of high-frequency task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119576470B_ABST
    Figure CN119576470B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and in particular to a model construction method, program product and storage medium for GUI intelligent bodies. The model construction method for GUI intelligent bodies provided by the present invention is to perform pre-training based on a general data set in the GUI field to obtain a base model; collect user demonstration operation trajectories online to obtain personalized task execution trajectories; automatically annotate the task execution trajectories to obtain annotated trajectories; based on the annotated trajectories, perform task planning training on the base model to obtain a scene planner, and obtain a personalized model by combining the base model and the scene planner; based on an external scene knowledge base, retrieve corresponding scene knowledge as a context supplement when performing personalized task action planning. By automatically annotating the task execution trajectory, guiding the large model to actively annotate and understand the thinking of humans in executing operations, and further planning the next step of behavior, the adaptability of the GUI intelligent body to the user's personalized use environment is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a model building method, program product and storage medium for a GUI intelligent body. Background Art

[0002] With the development of large language models and multimodal large model technologies, AI agents based on large models are becoming more flexible in executing behavior planning, memory storage, tool calling, etc. Among them, GUI (Graphical User Interface) agents have been widely used due to their high functional practicality and anthropomorphism.

[0003] However, GUI agents, especially mobile GUI agents based on multimodal large models, lack the ability to adapt to users' personalized usage environments and perform high-frequency tasks when dealing with different usage scenarios, because each user's device version, APP type, operating system and other usage environments are very different. Summary of the invention

[0004] In order to solve the problems of the existing mobile GUI intelligent body's lack of adaptability to the user's personalized usage environment and insufficient execution of high-frequency tasks, the present invention provides a model building method, program product and storage medium for GUI intelligent body.

[0005] The solution to the technical problem of the present invention is to provide a model construction method for a GUI agent, and the model construction method for a GUI agent comprises the following steps:

[0006] Pre-training is performed based on a common dataset in the GUI field to obtain a base model;

[0007] Collect user demonstration operation trajectories online to obtain personalized task execution trajectories;

[0008] Automatically labeling the task execution trajectory to obtain a labeled trajectory;

[0009] Based on the marked trajectory, the base model is trained for task planning, the training results are integrated to obtain a scenario planner, and a personalized model is obtained by combining the base model with the scenario planner;

[0010] Acquire an external scene knowledge base, provide supplementary scene information to the personalized model based on the external scene knowledge base, and retrieve corresponding scene knowledge as context supplement when acquiring action planning in natural language form;

[0011] Based on the marked trajectory, the base model is trained for task planning, and the training results are integrated to obtain a scenario planner, which specifically includes the following steps:

[0012] Based on the marked trajectory, the task execution sequence is dynamically decomposed and planned into a series of low-level action instructions to obtain a task planning sequence;

[0013] Repeatedly acquiring different labeled trajectories in multiple scenarios and dynamically decomposing the task execution sequence to obtain a task planning sequence group consisting of a plurality of the task planning sequences;

[0014] Combining the labeled trajectory with the task planning sequence group, fine-tuning the base model, and obtaining a fine-tuning layer after training;

[0015] The fine-tuning layers under various scenarios are comprehensively sorted out to obtain the scenario planner adapted to various scenarios.

[0016] Preferably, collecting user demonstration operation traces online specifically includes the following steps:

[0017] Obtain high-frequency task instructions input by the user, and demonstrate task execution in the order of high-frequency task instructions;

[0018] Through the Android background debugging tool, the interface status is captured at a high frequency during the user demonstration process to collect the user demonstration operation track.

[0019] Preferably, collecting the user demonstration operation track online to obtain the personalized task execution track specifically includes:

[0020] Based on the similarity of the interface states, the collected user demonstration operation traces are cleaned and deduplicated through the Android background debugging tool and the supporting image analysis model to obtain the task execution trace.

[0021] Preferably, the task execution trajectory is automatically labeled to obtain the labeled trajectory, which specifically includes the following steps:

[0022] Collect interface switching information, task instruction information, supplementary APP introduction information, and action history information in the task execution trajectory, and integrate them to obtain additional contextual input;

[0023] The action-thinking chain technology is used to simulate the thinking process of humans during operation, and the task execution trajectory is actively understood and annotated in combination with the additional context input to obtain an annotated trajectory.

[0024] Preferably, based on the marked trajectory, the base model is trained for task planning, and the training results are integrated to obtain a scenario planner, which specifically includes the following steps:

[0025] Based on the marked trajectory, the task execution sequence is dynamically decomposed and planned into a series of low-level action instructions to obtain a task planning sequence;

[0026] Repeatedly acquiring different labeled trajectories in multiple scenarios and dynamically decomposing the task execution sequence to obtain a task planning sequence group consisting of a plurality of the task planning sequences;

[0027] Combining the labeled trajectory with the task planning sequence group, fine-tuning the base model, and obtaining a fine-tuning layer after training;

[0028] The fine-tuning layers under various scenarios are comprehensively sorted out to obtain the scenario planner adapted to various scenarios.

[0029] Preferably, the base model is fine-tuned by a lightweight fine-tuning method, specifically including:

[0030] The base model is trained using a low-rank adaptive fine-tuning scheme to obtain the LoRA fine-tuning layer.

[0031] Preferably, obtaining the plug-in scenario knowledge base specifically includes:

[0032] According to the task execution trajectory, retrieve the function introduction and usage example information of the corresponding scenario in the Internet and user guide, and obtain a network search library after collating and unifying the form;

[0033] Through the large model, random walks are performed in the corresponding scenes to collect a series of scene data of screen switching states and switching relationships, and the random walk library is obtained after sorting.

[0034] Based on the network search library and the random walk library, the plug-in scenario knowledge base is obtained.

[0035] Preferably, obtaining the plug-in scenario knowledge base based on the network search library and the random walk library specifically includes:

[0036] The plug-in scenario knowledge base is constructed in the form of a state diagram or a narrative memory module.

[0037] In order to solve the above technical problems, the present invention also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the model building method for GUI intelligent body as described in any of the above items.

[0038] In order to solve the above technical problems, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the model building method for a GUI intelligent agent as described in any of the above items.

[0039] Compared with the prior art, the model building method, program product and storage medium for GUI agent of the present invention have the following advantages:

[0040] 1. The model construction method for GUI intelligent body of the present invention includes the following steps: pre-training based on the general data set in the GUI field to obtain the base model; collecting the user demonstration operation trajectory online to obtain the personalized task execution trajectory; automatically annotating the task execution trajectory to obtain the annotated trajectory; based on the annotated trajectory, the base model is trained for task planning, the training results are integrated to obtain the scene planner, and the personalized model is obtained by combining the base model and the scene planner; the plug-in scene knowledge base is obtained, and based on the plug-in scene knowledge base, the personalized model is provided with supplementary scene information, and the corresponding scene knowledge is retrieved as context supplement when obtaining the action planning in the form of natural language. By automatically annotating the task execution trajectory, the large model is guided to actively annotate and understand the thinking of humans in the process of executing the operation, and further plan the next step of behavior, which greatly improves the adaptability of the GUI intelligent body to the user's personalized use environment.

[0041] 2. The online collection of user demonstration operation trajectories in the present invention specifically includes the following steps: obtaining high-frequency task instructions input by the user, and demonstrating task execution in the order of high-frequency task instructions; using the Android background debugging tool, capturing the interface state at high frequency during the user demonstration process, and collecting the user demonstration operation trajectory. The collection of user demonstration operation trajectories by the Android background debugging tool provides highly targeted and highly personalized interactive data in different scenarios for large model training, which is conducive to the subsequent response training of GUI intelligent agents to high-frequency task instructions.

[0042] 3. The online collection of user demonstration operation trajectories in the present invention to obtain personalized task execution trajectories specifically includes: based on the similarity of the interface state, the collected user demonstration operation trajectories are cleaned and deduplicated through the Android background debugging tool and the supporting image analysis model to obtain the task execution trajectory. By cleaning and deduplicating the user demonstration operation trajectories through the Android background debugging tool and the supporting image analysis model, the repeated data interference in the task execution trajectory is reduced, the repeated calculation of big data is avoided, the storage space occupied by the task execution trajectory is reduced, and the data processing efficiency of the GUI intelligent body is improved.

[0043] 4. The present invention automatically labels the task execution trajectory to obtain the labeled trajectory, which specifically includes the following steps: collecting the interface switching information, task instruction information, supplementary APP introduction information and action history information in the task execution trajectory, and integrating them to obtain additional context input; using the action thinking chain technology to simulate the thinking process of humans in the operation process, and actively understand and label the task execution trajectory in combination with additional context input to obtain the labeled trajectory. Through the use of the action thinking chain technology, the large model is further guided to actively understand and label the trajectory information demonstrated by humans, restore the human thinking process, enhance the action semantic understanding of the GUI intelligent body to the user's demonstration operation trajectory, improve the generalization performance of scene adaptation, and reduce the risk of overfitting.

[0044] 5. The present invention performs task planning training on the base model based on the marked trajectory, integrates the training results and obtains a scene planner, which specifically includes the following steps: based on the marked trajectory, dynamically decompose the task execution sequence, and plan it into a series of low-order action instructions to obtain a task planning sequence; repeatedly obtain different marked trajectories in multiple scenarios and dynamically decompose the task execution sequence to obtain a task planning sequence group composed of multiple task planning sequences; combine the marked trajectory with the task planning sequence group to fine-tune the base model, and obtain a fine-tuning layer after training; comprehensively organize the fine-tuning layers in multiple scenarios to obtain a scene planner adapted to multiple scenarios. On the one hand, by dynamically decomposing the task execution sequence, the high-order task instructions are decoupled from the specific low-order action instructions, and the action execution capability of the base model is maximized, which is convenient for the function realization of the GUI intelligent body; on the other hand, by fine-tuning the base model to obtain the fine-tuning layer, it is beneficial for the GUI intelligent body to use the existing trained general base model, greatly reducing the number of training parameters required to obtain the fine-tuning layer, reducing the storage cost of the GUI intelligent body, and improving the computing efficiency.

[0045] 6. In the present invention, the base model is fine-tuned by a lightweight fine-tuning method, and a low-rank adaptive fine-tuning scheme is used to train the base model to obtain the LoRA fine-tuning layer. By adopting a low-rank adaptive fine-tuning training scheme, the scene adaptation is achieved in a lightweight manner, while avoiding the scene conflicts that occur when expanding and enriching third-party scenes, and improving the adaptability of the GUI agent in a variety of personalized usage environments.

[0046] 7. The method of obtaining the plug-in scene knowledge base of the present invention specifically includes the following steps: according to the task execution trajectory, the function introduction and usage example information of the corresponding scene in the Internet and the user guide are retrieved, and the network search library is obtained after being sorted and unified; a large model is used to perform random walks in the corresponding scene to collect a series of scene data of screen switching states and switching relationships, and a random walk library is obtained after sorting; based on the network search library and the random walk library, the plug-in scene knowledge base is obtained. By obtaining the network search library and the random walk library, the plug-in scene knowledge base synchronizes the latest external information to the personalized model in real time, alleviating the defect that the data coverage of the user's online demonstration data may be insufficient, and can automatically synchronize the latest scene usage information in real time, reducing the AI ​​illusion in the intelligent body's task execution and avoiding the AI ​​illusion caused by the single source of data source, thereby improving the retrieval accuracy of the GUI intelligent body.

[0047] 8. The invention obtains a plug-in scene knowledge base based on a network search library and a random walk library, specifically including: the plug-in scene knowledge base is in the form of a state diagram or a narrative memory module. By introducing the plug-in scene knowledge base, adding prompt words to the state diagram or narrative memory module is conducive to guiding the GUI intelligent body to think and understand in combination with state transition and task flow in planning thinking, thereby improving the planning effect of the GUI intelligent body.

[0048] 9. The present invention also provides a computer program product and a computer-readable storage medium, which have the same beneficial effects as the above-mentioned model building method for GUI intelligent body, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0050] Figure 1 It is a step flow chart of the model building method for GUI intelligent body provided by the first embodiment of the present invention.

[0051] Figure 2 It is a flowchart of step S2 of the model building method for GUI intelligent body provided by the first embodiment of the present invention.

[0052] Figure 3 It is a flow chart of step S3 of the model building method for GUI intelligent body provided by the first embodiment of the present invention.

[0053] Figure 4It is a flow chart of step S4 of the model building method for GUI intelligent body provided by the first embodiment of the present invention.

[0054] Figure 5 It is an example diagram of the model building method for GUI intelligent agent provided by the first embodiment of the present invention. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and implementation examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0056] See also Figure 1 The first embodiment of the present invention provides a model building method for a GUI agent, comprising the following steps:

[0057] S1: Pre-train based on the general dataset in the GUI field to obtain the base model.

[0058] As an optional implementation, the general dataset in the GUI field is based on large-scale pre-annotated offline datasets, such as the Android device control AitW dataset and the AndroidControl dataset. Pre-training is performed on the general dataset to construct a large amount of sequence data of GUI operations, thereby quickly and effectively improving the GUI operation capabilities of large models.

[0059] S2: Collect user demonstration operation trajectories online to obtain personalized task execution trajectories.

[0060] As an optional implementation method, the Android background debugging tool ADB (Android Debug Bridge) and mobile application development technology are used to design a user demonstration and data collection interface, and the user demonstration operation trajectory is collected online.

[0061] S3: Automatically label the task execution trajectory to obtain a labeled trajectory.

[0062] As an optional implementation, based on the large model's automatic labeling of task execution trajectories, human manual labeling, human static labeling and human-machine enhanced hybrid labeling can be added to supplement the labeling trajectories and enhance the GUI agent's ability to understand the task execution trajectories.

[0063] S4: Based on the labeled trajectory, the base model is trained for task planning, the training results are integrated to obtain the scenario planner, and the personalized model is obtained by combining the base model and the scenario planner.

[0064] As an optional implementation, task planning training can be formalized into a large model to input user goals, observe scenarios and interaction history, output a process operation of action plans in natural language form, and repeat the training.

[0065] S5: Obtain an external scene knowledge base, and provide supplementary scene information to the personalized model based on the external scene knowledge base. When obtaining action planning in natural language form, retrieve corresponding scene knowledge as context supplement.

[0066] As an optional implementation, the external scene knowledge base is a database obtained by collecting corresponding scene information when the large model is connected to the Internet; in order to avoid external data information affecting the training process of user data for GUI intelligent agents, the external scene knowledge base is only input as supplementary information after the personalized model training is completed.

[0067] It can be understood that by guiding the big model to automatically label the task execution trajectory, it is helpful for the big model to understand human thinking during the operation process and plan the next step of behavior, enhance the GUI intelligent agent's ability to think and learn about different human behaviors, and greatly improve the adaptability of the GUI intelligent agent to the user's personalized usage environment.

[0068] See also Figure 2 , further, step S2 specifically includes:

[0069] S21: Obtain high-frequency task instructions input by the user, and demonstrate task execution according to the sequence of high-frequency task instructions.

[0070] As an optional implementation, the high-frequency task instructions input by the user are divided into plain text modal input and multi-modal information input.

[0071] S22: Through the Android background debugging tool, the interface status is captured at a high frequency during the user demonstration process to collect the user demonstration operation track.

[0072] As an optional implementation, when the high-frequency task instructions in the user demonstration process are plain text modal input, the Android background debugging tool uses the visual tools of the Accessibility Tree, optical character recognition technology (OCR) and GUI element recognition network (IconNet) to capture the interface state and obtain a user demonstration operation trajectory consisting of a series of interface state screenshots.

[0073] It can be understood that collecting user demonstration operation trajectories is to collect data pairs of user personalized task-operation sequences and store them in the form of screenshot images; by collecting a large number of user demonstration operation trajectories, a large amount of highly targeted and personalized interactive data containing user tasks and corresponding operations is provided to the large model, which is conducive to training the GUI intelligent agent's response ability to high-frequency task instructions in different scenarios.

[0074] Furthermore, after step S22, the following steps are specifically included:

[0075] S23: Based on the similarity of the interface states, the collected user demonstration operation trajectories are cleaned and deduplicated through the Android background debugging tool and the supporting image analysis model to obtain the task execution trajectory.

[0076] It can be understood that cleaning and deduplication of user demonstration operation trajectories eliminates the interference of duplicate data. On the one hand, it prevents duplicate data from affecting subsequent model training, causing repeated calculations in big data and leading to overfitting; on the other hand, it reduces the data size of task execution trajectories, thereby reducing the storage space they occupy, which is conducive to storing more task execution trajectories in limited data storage space, greatly improving the data processing efficiency of GUI intelligent agents.

[0077] See also Figure 3 , further, step S3 specifically includes:

[0078] S31: collecting interface switching information, task instruction information, supplementary APP introduction information and action history information in the task execution trajectory, and integrating them to obtain additional context input;

[0079] S32: Use action-thinking chain technology to simulate the thinking process of humans during operation, and actively understand and annotate the task execution trajectory in combination with additional contextual input to obtain the annotated trajectory.

[0080] As an optional implementation, since the task execution trajectory is a group of screenshots collected for each step in the corresponding user operation sequence, the automatic annotation process of the large model is to make a brief understanding and analysis of each current screenshot interface, and combine historical actions and behavior planning to simulate the thinking process of humans when performing the operation.

[0081] Understandably, the action thinking chain technology enables GUI agents to further expand the connection between thinking behavior and the environment, restore the human thinking process, rather than relying solely on preset programs or direct task-behavior mapping to determine the annotation content of the trajectory information of the large model, and enhance the GUI agent's action semantic understanding of the user's demonstration operation trajectory, which is conducive to generalized learning of data knowledge in new scenarios; at the same time, it reduces the GUI agent's dependence on specific environments, improves the generalization performance of scene adaptation, and reduces the risk of overfitting.

[0082] See also Figure 4 , further, step S4 specifically includes:

[0083] S41: Based on the labeled trajectory, dynamically decompose the task execution sequence and plan it into a series of low-level action instructions to obtain the task planning sequence.

[0084] As an optional implementation, the low-level action instructions are pre-trained GUI base models that can be implemented into designated task operations on the interface through certain specific action plans.

[0085] For example, the task execution sequence is: collect a store in a certain application;

[0086] The task execution sequence is broken down into a series of low-level action instructions: click the app icon, click the search box, enter the store name, click search, click the first search result, and click the favorite button.

[0087] In the process of executing the task execution sequence, low-level action instructions are all action instructions that can be implemented through the base model.

[0088] S42: Repeatedly obtain different labeled trajectories in multiple scenarios and dynamically decompose the task execution sequence to obtain a task planning sequence group consisting of multiple task planning sequences.

[0089] S43: Fine-tune the base model by combining the labeled trajectory and the task planning sequence group, and obtain the fine-tuning layer after training.

[0090] As an optional implementation, each fine-tuning layer trains a specific scenario planner for the personalized scenario corresponding to each task planning sequence.

[0091] S44: Comprehensively organize the fine-tuning layers under various scenarios to obtain a scenario planner that is suitable for various scenarios.

[0092] It can be understood that dynamically decomposing the task execution sequence decouples high-level task instructions from specific low-level action instructions, that is, breaking down complex actions into multiple action groups that can be implemented by the base model, maximizing the use of the action execution capabilities of the base model, and facilitating the GUI intelligent agent to realize the functions required by users.

[0093] It can be understood that, on the one hand, the fine-tuning training process of the base model is conducive to the GUI agent to utilize the existing model data, greatly reducing the number of training parameters, saving the computing resources and time of the large model, and further improving the operating efficiency of the GUI agent; on the other hand, the use of specific data sets for small-scale data training during fine-tuning reduces the dependence on the overall task planning sequence group and its annotations, effectively avoiding the risk of overfitting of the large model during the training process, thereby improving the generalization ability and versatility of the GUI agent.

[0094] Furthermore, in step S43, a low-rank adaptive fine-tuning scheme is used to train the base model to obtain a LoRA fine-tuning layer.

[0095] It should be noted that LoRA fine-tunes the base model by introducing a low-rank matrix whose rank is much lower than the parameter dimension of the base model, and freezes the update state of the base model during the fine-tuning process. It only trains the low-rank matrix, and then merges the low-rank matrix into the base model after the training is completed, ensuring that the training process of the low-rank matrix is ​​independent and does not affect the base model, and improving the fine-tuning training speed and the efficiency of the scenario planner.

[0096] It can be understood that the LoRA fine-tuning solution can be orthogonally combined with other scenarios and other fine-tuning technologies, adapting scenarios in a lightweight manner and avoiding conflicts when expanding third-party scenarios, thereby improving the composability of GUI agents with other fine-tuners and their adaptability in a variety of personalized usage environments.

[0097] See also Figure 1 and Figure 2 Further, after step S23, the following steps are specifically included:

[0098] S24: According to the task execution trajectory, the function introduction and usage example information of the corresponding scenario in the Internet and the user guide are retrieved, and the network search library is obtained after being sorted and unified.

[0099] S25: Perform random walks in the corresponding scene through the large model to collect a series of scene data of screen switching states and switching relationships, and obtain a random walk library after sorting.

[0100] S26: Based on the network search library and the random walk library, obtain the plug-in scenario knowledge base.

[0101] It can be understood that the user guide includes a user operation manual or a user instruction manual.

[0102] Understandably, the web search library provides external real-time network information supplements to the plug-in scene knowledge base, provides timely verification information, alleviates the defect that the data coverage of the user's online demonstration data may be insufficient, and can automatically synchronize the latest scene usage information in real time, helping the personalized model to reduce possible fictitious or erroneous results; the random walk library helps the large model identify the key nodes and thinking structures in the corresponding scene, helps to reveal the intrinsic properties of the scene, strengthens the large model's in-depth understanding of the current scene information, and improves the retrieval accuracy of the GUI intelligent agent.

[0103] Furthermore, in step S26, the external scenario knowledge base is in the form of a state diagram or a narrative memory module.

[0104] It should be noted that the plug-in scenario knowledge base in the form of a state diagram models the switching logic of the scene interface in the personalized scenario in the form of a state diagram, in which each node corresponds to an interface state containing a corresponding screenshot and interface summary, and the directed edges represent the switching logic between interfaces and their corresponding operations: the plug-in scenario knowledge base in the form of a narrative memory module is presented as a brief description of the task execution process of some high-frequency tasks.

[0105] It can be understood that adding state diagrams or brief descriptions of the execution process of high-frequency tasks as supplementary prompts to the retrieval process of the personalized model is conducive to guiding the GUI agent to introduce thinking and understanding of the logic of the state transition diagram when planning thinking tasks, and conduct secondary comparison checks with commonly used execution tasks to improve the planning effect of the GUI agent.

[0106] Please further combine Figure 5 For example, when a user downloads the Google Drive application and hopes that the model can learn to "create a new folder named 'Dreams and Ambitions' in Google Drive", the GUI agent can learn the user's demonstration operation trajectory through the following process, using the action thinking chain technology from observation to action thinking, combined with fine-tuning technology to train the base model for task planning, and obtain a scenario planner adapted to the personalized scenario.

[0107] Figure 5 A in the figure is the main interface of the Google Drive application, including the notification bar at the top of the interface, the search box, the main application body, the selection bar at the bottom, and the plus icon; among them, the suggestion and notification columns are displayed below the search box, two pictures are displayed in the main application body, and the selection bar at the bottom includes four columns: home page, starred files, sharing, and files, and the home page column is currently lit. Based on the action thinking of "In order to create a new folder with a specified name, I need to click the plus button at the bottom right of the screen to create a new process", the GUI agent generates the action of "clicking the plus icon" and reflects on the result of "after clicking the plus button, specific new content options appear".

[0108] Figure 5 B in the figure is the pop-up window interface that appears after clicking the plus icon, including the darkened and non-interactive application body and the six new options displayed in the pop-up window; the six new options are creating a new folder, uploading files, adding files by scanning, creating a new Google document, creating a new Google spreadsheet, and creating a new Google slide. Based on the action thinking of "In order to create a new folder with a specified name, I need to click the New Folder button to start creating a folder", the GUI agent generates the action of "clicking the New Folder button" and reflects on the result of "after clicking the New Folder button, entering the New Folder Naming Interface".

[0109] Figure 5 C in the figure is a new folder naming interface, including a naming pop-up window with a text input box, a cancel button and a new button, and a mobile phone keyboard below the pop-up window; the current content in the text box is the fully selected "Untitled Folder". Based on the action thinking of "In order to create a new folder with a specified name, I need to click the delete button on the mobile phone keyboard to delete the current name, and then enter the target folder name", the GUI agent generates the action of "clicking the delete button on the mobile phone keyboard" and reflects on the result of "after clicking the delete button, the default content in the text box is cleared".

[0110] Figure 5 D in the figure is a new folder naming interface, including a naming pop-up window with a text input box, a cancel button and a new button, and a mobile phone keyboard below the pop-up window; the current content in the text box is blank. Based on the action thinking of "In order to create a new folder with a specified name, I need to enter the specified name in the text box", the GUI agent generates the action of "entering 'dreams and ambitions'" and reflects on the result of "after input, the specified folder name appears in the text box".

[0111] Figure 5 E in the figure is a new folder naming interface, including a naming pop-up window with a text input box, a cancel button and a new button, and a mobile phone keyboard below the pop-up window; the current content in the text box is "Dreams and Ambitions". Based on the action thinking of "In order to create a new folder with a specified name, I need to click the new button to complete the folder creation", the GUI agent generates the action of "clicking the new button" and reflects on the result of "after clicking the new button, the target folder is successfully created".

[0112] Figure 5F in the figure is the interface of the user's cloud disk content in Google Drive, including the user's cloud disk folder list, the new creation success pop-up window, and the selection bar at the bottom; the top layer of the user's cloud disk folder list is the newly created folder with the specified name "Dreams and Ambitions", and the file column in the selection bar at the bottom is lit. Based on the action thinking of "the task of creating a folder with a specified name has been completed, and I need to output a task completion signal", the GUI agent generates the action of "outputting a task completion signal, task completed".

[0113] Furthermore, according to the process described above, the large model can be actively summarized and concluded into the following summary: "In order to create a new folder named Dreams and Ambitions in Google Drive, we click the plus button in the main interface to start the new creation process, and the new content option appears; click the folder icon to enter the naming interface of the new file; since there is currently a default folder name, I click the Delete button to clear the default content in the text box, then enter the specified name and click the Confirm Creation option, and the task is completed."

[0114] After obtaining the above content, the summary is optionally stored in the plug-in scene knowledge base, or is further updated with the iteration of actual deployment, retrieval and demonstration of the large model, and an optimized summary is obtained.

[0115] The second embodiment of the present invention provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the model building method for a GUI agent as in any one of the first embodiments is implemented. The model building method for a GUI agent can have the same beneficial effects as the model building method for a GUI agent in the first embodiment, and will not be described in detail here.

[0116] The third embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, a model building method for a GUI agent as in any one of the first embodiments is implemented. It can have the same beneficial effects as the model building method for a GUI agent in the first embodiment, and will not be described in detail here.

[0117] In the embodiments provided by the present invention, it should be understood that "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.

[0118] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. Those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0119] In various embodiments of the present invention, it should be understood that the size of the serial numbers of the above-mentioned processes does not mean the necessary order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0120] The flow chart and block diagram in the accompanying drawings of the present invention illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which is determined based on the functions involved. It should be particularly noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs a specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0121] Compared with the prior art, the model building method, program product and storage medium for GUI agent of the present invention have the following advantages:

[0122] 1. The model construction method for GUI intelligent body of the present invention includes the following steps: pre-training based on the general data set in the GUI field to obtain the base model; collecting the user demonstration operation trajectory online to obtain the personalized task execution trajectory; automatically annotating the task execution trajectory to obtain the annotated trajectory; based on the annotated trajectory, the base model is trained for task planning, the training results are integrated to obtain the scene planner, and the personalized model is obtained by combining the base model and the scene planner; the plug-in scene knowledge base is obtained, and based on the plug-in scene knowledge base, the personalized model is provided with supplementary scene information, and the corresponding scene knowledge is retrieved as context supplement when obtaining the action planning in the form of natural language. By automatically annotating the task execution trajectory, the large model is guided to actively annotate and understand the thinking of humans in the process of executing the operation, and further plan the next step of behavior, which greatly improves the adaptability of the GUI intelligent body to the user's personalized use environment.

[0123] 2. The online collection of user demonstration operation trajectories in the present invention specifically includes the following steps: obtaining high-frequency task instructions input by the user, and demonstrating task execution in the order of high-frequency task instructions; using the Android background debugging tool, capturing the interface state at high frequency during the user demonstration process, and collecting the user demonstration operation trajectory. The collection of user demonstration operation trajectories by the Android background debugging tool provides highly targeted and highly personalized interactive data in different scenarios for large model training, which is conducive to the subsequent response training of GUI intelligent agents to high-frequency task instructions.

[0124] 3. The online collection of user demonstration operation trajectories in the present invention to obtain personalized task execution trajectories specifically includes: based on the similarity of the interface state, the collected user demonstration operation trajectories are cleaned and deduplicated through the Android background debugging tool and the supporting image analysis model to obtain the task execution trajectory. By cleaning and deduplicating the user demonstration operation trajectories through the Android background debugging tool and the supporting image analysis model, the repeated data interference in the task execution trajectory is reduced, the repeated calculation of big data is avoided, the storage space occupied by the task execution trajectory is reduced, and the data processing efficiency of the GUI intelligent body is improved.

[0125] 4. The present invention automatically labels the task execution trajectory to obtain the labeled trajectory, which specifically includes the following steps: collecting the interface switching information, task instruction information, supplementary APP introduction information and action history information in the task execution trajectory, and integrating them to obtain additional context input; using the action thinking chain technology to simulate the thinking process of humans in the operation process, and actively understand and label the task execution trajectory in combination with additional context input to obtain the labeled trajectory. Through the use of the action thinking chain technology, the large model is further guided to actively understand and label the trajectory information demonstrated by humans, restore the human thinking process, enhance the action semantic understanding of the GUI intelligent body to the user's demonstration operation trajectory, improve the generalization performance of scene adaptation, and reduce the risk of overfitting.

[0126] 5. The present invention performs task planning training on the base model based on the marked trajectory, integrates the training results and obtains a scene planner, which specifically includes the following steps: based on the marked trajectory, dynamically decompose the task execution sequence, and plan it into a series of low-order action instructions to obtain a task planning sequence; repeatedly obtain different marked trajectories in multiple scenarios and dynamically decompose the task execution sequence to obtain a task planning sequence group composed of multiple task planning sequences; combine the marked trajectory with the task planning sequence group to fine-tune the base model, and obtain a fine-tuning layer after training; comprehensively organize the fine-tuning layers in multiple scenarios to obtain a scene planner adapted to multiple scenarios. On the one hand, by dynamically decomposing the task execution sequence, the high-order task instructions are decoupled from the specific low-order action instructions, and the action execution capability of the base model is maximized, which is convenient for the function realization of the GUI intelligent body; on the other hand, by fine-tuning the base model to obtain the fine-tuning layer, it is beneficial for the GUI intelligent body to use the existing trained general base model, greatly reducing the number of training parameters required to obtain the fine-tuning layer, reducing the storage cost of the GUI intelligent body, and improving the computing efficiency.

[0127] 6. In the present invention, the base model is fine-tuned by a lightweight fine-tuning method, and a low-rank adaptive fine-tuning scheme is used to train the base model to obtain the LoRA fine-tuning layer. By adopting a low-rank adaptive fine-tuning training scheme, the scene adaptation is achieved in a lightweight manner, while avoiding the scene conflicts that occur when expanding and enriching third-party scenes, and improving the adaptability of the GUI agent in a variety of personalized usage environments.

[0128] 7. The method of obtaining the plug-in scene knowledge base of the present invention specifically includes the following steps: according to the task execution trajectory, the function introduction and usage example information of the corresponding scene in the Internet and the user guide are retrieved, and the network search library is obtained after being sorted and unified; a large model is used to perform random walks in the corresponding scene to collect a series of scene data of screen switching states and switching relationships, and a random walk library is obtained after sorting; based on the network search library and the random walk library, the plug-in scene knowledge base is obtained. By obtaining the network search library and the random walk library, the plug-in scene knowledge base synchronizes the latest external information to the personalized model in real time, alleviating the defect that the data coverage of the user's online demonstration data may be insufficient, and can automatically synchronize the latest scene usage information in real time, reducing the AI ​​illusion in the intelligent body's task execution and avoiding the AI ​​illusion caused by the single source of data source, thereby improving the retrieval accuracy of the GUI intelligent body.

[0129] 8. The invention obtains a plug-in scene knowledge base based on a network search library and a random walk library, specifically including: the plug-in scene knowledge base is in the form of a state diagram or a narrative memory module. By introducing the plug-in scene knowledge base, adding prompt words to the state diagram or narrative memory module is conducive to guiding the GUI intelligent body to think and understand in combination with state transition and task flow in planning thinking, thereby improving the planning effect of the GUI intelligent body.

[0130] 9. The present invention also provides a computer program product and a computer-readable storage medium, which have the same beneficial effects as the above-mentioned model building method for GUI intelligent body, and will not be elaborated here.

[0131] The above is a detailed introduction to a model building method, program product and storage medium for a GUI intelligent body disclosed in an embodiment of the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention, and any modifications, equivalent substitutions and improvements made within the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A model construction method for a GUI agent, characterized in that: The following steps are involved: Pre-training is performed based on a common dataset in the GUI field to obtain a base model; Collect user demonstration operation trajectories online to obtain personalized task execution trajectories; Automatically labeling the task execution trajectory to obtain a labeled trajectory; Based on the marked trajectory, the base model is trained for task planning, the training results are integrated to obtain a scenario planner, and a personalized model is obtained by combining the base model with the scenario planner; Acquire an external scene knowledge base, provide supplementary scene information to the personalized model based on the external scene knowledge base, and retrieve corresponding scene knowledge as context supplement when acquiring action planning in natural language form; Based on the marked trajectory, the base model is trained for task planning, and the training results are integrated to obtain a scenario planner, which specifically includes the following steps: Based on the marked trajectory, the task execution sequence is dynamically decomposed and planned into a series of low-level action instructions to obtain a task planning sequence; Repeatedly acquiring different labeled trajectories in multiple scenarios and dynamically decomposing the task execution sequence to obtain a task planning sequence group consisting of a plurality of the task planning sequences; Combining the labeled trajectory with the task planning sequence group, fine-tuning the base model, and obtaining a fine-tuning layer after training; The fine-tuning layers under various scenarios are comprehensively sorted out to obtain the scenario planner adapted to various scenarios.

2. The model building method for a GUI agent according to claim 1, characterized in that: The online collection of user demonstration operation traces specifically includes the following steps: Obtain high-frequency task instructions input by the user, and demonstrate task execution in the order of high-frequency task instructions; Through the Android background debugging tool, the interface status is captured at a high frequency during the user demonstration process to collect the user demonstration operation track.

3. The model building method for a GUI agent as claimed in claim 2, characterized in that: Obtaining a personalized task execution trajectory specifically includes: Based on the similarity of the interface states, the collected user demonstration operation traces are cleaned and deduplicated through the Android background debugging tool and the supporting image analysis model to obtain the task execution trace.

4. The model building method for a GUI agent according to claim 1, characterized in that: Automatically labeling the task execution trajectory to obtain a labeled trajectory specifically includes the following steps: Collect interface switching information, task instruction information, supplementary APP introduction information, and action history information in the task execution trajectory, and integrate them to obtain additional contextual input; The action-thinking chain technology is used to simulate the thinking process of humans during operation, and the task execution trajectory is actively understood and annotated in combination with the additional context input to obtain an annotated trajectory.

5. The model building method for a GUI agent according to claim 1, characterized in that: The base model is fine-tuned by a lightweight fine-tuning method, specifically including: The base model is trained using a low-rank adaptive fine-tuning scheme to obtain the LoRA fine-tuning layer.

6. The model building method for a GUI agent according to claim 1, characterized in that: Obtaining the plug-in scenario knowledge base specifically includes: According to the task execution trajectory, retrieve the function introduction and usage example information of the corresponding scenario in the Internet and user guide, and obtain a network search library after collating and unifying the form; Through the large model, random walks are performed in the corresponding scenes to collect a series of scene data of screen switching states and switching relationships, and the random walk library is obtained after sorting. Based on the network search library and the random walk library, the plug-in scenario knowledge base is obtained.

7. The model building method for GUI agent according to claim 6, characterized in that: The method of obtaining the plug-in scenario knowledge base based on the network search library and the random walk library specifically includes: The plug-in scenario knowledge base is constructed in the form of a state diagram or a narrative memory module.

8. A computer program product, characterized in that: The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the model building method for a GUI agent as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the model building method for a GUI agent as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Machine writing method and device based on Gaussian mixture model and dynamic motion primitives

    CN116721464A

  • Intelligent task sequence planning method based on language vision large model and knowledge graph

    CN117874258A