Automatic application program control method based on visual language large model agent

By adopting the multi-VLM agent method with real-time strategy planning as the main and global strategy planning as the supplement in the automatic control method of application based on visual language big model, and combining external GUI-Grounding technology to identify the GUI interface, the problem of dead loops and interface changes in the automatic control of complex software systems is solved, and the exploration and robustness of automatic control is improved.

CN120215768AActive Publication Date: 2025-06-27FUZHOU INSTITUE OF TECH

Patent Information

Application Number
CN202510720016.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-06-27
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

The existing application automatic control method based on visual language big model is prone to falling into a dead loop due to improper global strategy planning, and traditional methods are difficult to cope with the problem of frequent changes in the software interaction interface.

Method used

The multi-VLM agent method is adopted with real-time policy planning as the main and global policy planning as the auxiliary. The screenshot of the current application GUI interface is identified through the external GUI-Grounding method, and it is used as the input of VLM Agent's real-time policy planning, which enhances VLM's image understanding of the GUI interface and avoids dead loops.

Benefits of technology

It improves the exploratory and robustness of VLM Agent, enhances the automatic control ability of complex software systems, avoids the problem of dead loop caused by improper global policy planning, and is suitable for unknown applications or user automatic control tasks with high operation difficulty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215768A_ABST
    Figure CN120215768A_ABST
Patent Text Reader

Abstract

The invention relates to an automatic application program control method based on a visual language large model agent, and belongs to the technical field of information. According to the method, a user automatic control task is jointly completed in a cooperative scheduling mode of a plurality of VLM agents, and an agent decision-making method which takes instant strategy planning as a main mode and takes global strategy planning as an auxiliary mode is adopted, so that the defects of the global strategy planning are overcome, and the generalization ability and universality of the method are improved. In order to excavate the potential of the VLM Agent in solving the automatic control problem of the application program, a general rule element extraction mode is adopted to replace a mainstream GUI-Grounding method so as to improve the UI control identification accuracy as much as possible. Besides, according to the method, multi-modal messages generated when the VLM Agent executes the automatic control task are spliced by using an image splicing technology, so that the proportion of multi-round long dialogue type image information in cue words is reduced, the operation speed of the method is improved, and the problem that the shared historical context is too long is relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information technology, and particularly relates to an automatic control method for application programs based on an intelligent agent of a vision-language large model. Background Art

[0002] In the field of human-computer interaction, graphical user interfaces (GUIs) are widely used in electronic devices such as smartphones and computers, allowing users to freely control these interactive devices through input devices such as mice and keyboards, and complete a series of complex user tasks on various application software. However, due to the flexible and changeable user requirements and the increasing requirements of users for software functions, the interaction interfaces of application software are becoming more and more complex and bloated. Users often need to spend more time learning and memorizing the usage methods of application software first, which makes the user-software interface interaction experience worse. It becomes difficult and troublesome for users to complete some user interaction tasks with more operation steps. In addition, due to the increase in UI control elements and the complexity of the layout in the GUI, the difficulty of writing program automation test control scripts also increases accordingly. At the same time, traditional automation test methods usually rely on hard coding and are difficult to cope with frequently changing software interaction interfaces. For some large software, once there is a version update and iteration, the automation test control scripts need to be rewritten.

[0003] Due to the popularity of vision-language large models (VLMs) in recent years, more and more research scholars have used VLM technology to solve such human-computer interaction problems at the software level, such as the CogAgent method proposed by Hong et al. and the UFO method proposed by Zhang et al. These methods utilize the characteristic that VLMs have strong understanding capabilities for text and images, and automate the control process of the steps of manually operating UI control elements, thereby simplifying the interaction process between users and complex software. VLMs trained based on zero-shot learning have the generalization ability to handle unknown scenarios, can better understand the task purposes of users, and give corresponding language text or image outputs. This enables VLMs to still maintain a good response effect when facing large software with rapid version updates and iterations, and assist users in completing a series of complex human-computer interaction tasks.

[0004] Although current advanced VLMs, such as GPT-4V(o), perform poorly in scenario tasks with high-dimensional state and action spaces, lack complex image logical reasoning capabilities, and have limited context awareness, for problems in automatic control software applications, when users interact with the GUI, they can complete tasks through discrete actions such as mouse clicks and keyboard inputs. The intermediate Tokens generated during this process are few, which can make good use of the context awareness of the VLM. At the same time, most tasks in user-GUI interaction do not involve highly logical reasoning steps. Therefore, in relatively simple user interaction tasks, the VLM can better replace users to operate applications to achieve automatic control of applications.

[0005] One of the key steps in implementing automatic control of applications by VLM is the coordinate positioning of UI controls, that is, the 2D-Grounding task. The accuracy of UI control recognition directly determines the performance and effect of the VLM in automatic control applications. Currently, methods for automatic control of applications based on VLM are mainly divided into two categories: The first category is methods based on GUI-Grounding, such as the ScreenAI method proposed by Baechler et al. and the CogAgent method proposed by Hong et al. While the VLM understands the content of the GUI image, it locates the coordinates of the UI controls that appear in the image and outputs the UI controls required to complete the application control task specified by the user according to the user's prompt words; The second category of methods combines an external GUI-Grounding model or tool and embeds it into the VLM, such as the UFO method proposed by Zhang et al. and the OS-ATLAS method proposed by Wu et al. The UI control recognition task in the GUI image is decoupled from the inside of the VLM and handed over to an external GUI-Grounding method. In order to improve the performance of the method and the compatibility of the VLM model, the present invention adopts a method of combining an external GUI-Grounding to identify UI controls, and uses the visual perception characteristics of the VLM that are sensitive to image marking to enhance the decision-making accuracy of the VLM.

[0006] When designing the intelligent agent strategy in the existing VLM-based application automatic control method, most of them adopt global strategy planning. Before controlling the specified application, first, a VLM agent plans all possible subsequent UI operations according to the specified operation tasks proposed by the user. This method is suitable for application scenarios with fewer operation steps or simple software GUIs. However, when facing some large and complex software systems with rapid version iteration and updates, since this method depends on the prior knowledge of GUI operations input during the training of the VLM model, it is relatively easy to give incorrect UI operation strategy planning in the initial stage of the method, resulting in a dead loop when the subsequent VLM automatically controls the GUI operations. Therefore, the present invention adopts a multi-VLM agent method with immediate strategy planning as the main and global strategy planning as the auxiliary. When the VLM agent controls the application, it will continuously correct the global strategy according to the screenshot of the current application GUI interface, and immediately judge the next UI operation, similar to the behavior strategy when reinforcement learning explores the environment. Summary of the Invention

[0007] The purpose of the present invention is to provide an automatic control method for applications based on visual language large model agents, aiming to overcome the limitations of the global strategy planning of VLM agents (VLM Agent). It adopts a VLM multi-agent method CCMAgent (Computer Control Multi-Agent) with immediate strategy planning as the main and global strategy planning as the auxiliary, and utilizes the visual perception characteristics of VLM being sensitive to image marking. Through the external GUI-Grounding method, the UI controls of the screenshot of the current application GUI interface are identified and used as the input for the immediate strategy planning of the VLM Agent, enhancing the VLM's image understanding of the GUI interface screenshot, enabling the VLM to immediately judge the next UI operation step according to the GUI image content, and correcting the global strategy planning, increasing the exploration of the VLM Agent, improving the robustness of the method of the present invention, and avoiding the dead loop of automatic control caused by improper global strategy planning.

[0008] To achieve the above object, the present invention provides the following technical solutions: An application program automatic control method based on a vision-language large model agent, which designs three vision-language large model agents VLM Agent with different roles, namely an application program agent Application Agent, a user interface agent UI Agent, and a user task inspection agent Check Agent. After the user inputs a task description prompt or voice, the Application Agent is first responsible for parsing the user input and splitting the user operation task into a series of executable user interface UI control operations, that is, global policy planning. Then, the Application Agent obtains the corresponding window handle from the environment variable or configuration file according to the specific name of the application program extracted from the global policy planning, that is, starts the specified application program, and transfers the VLM execution leadership to the UI Agent. The UI Agent uses the user interface tool set UI Tools designed by the present invention to capture screenshots of the graphical user interface GUI of the current application window, and combines the external graphical user interface localization GUI-Grounding method to identify all operable UI operation controls appearing in the current GUI interface. In order to enhance the VLM's understanding of the GUI image, after the UI Agent identifies the UI controls on the GUI interface, it will label detection frames and unique control identification numbers ID for all the identified UI controls, so as to utilize the visual perception characteristics of the VLM sensitive to image annotation. At the same time, the UI Agent will give an immediate policy planning based on the global policy planning and the GUI interface screenshot after labeling the detection frames, select the UI control to be operated in the current step, and correct the global policy planning. Finally, the Check Agent will judge whether the user's current task has been completed according to the GUI interface screenshot after operating the specified UI control. If not, it will continue to transfer the VLM execution leadership to the UI Agent for the next UI control operation. If it is completed, the application program agent Application Agent will output a terminator, end the current automatic control task, and notify the user.

[0009] Further, the method includes the following steps: Step S1, construct a vision-language large model tool VLM Tools; Step S2, construct a vision-language large model team VLM Team and a vision-language large model agent VLM Agent; Step S3, evaluate the application program automatic control method; Further, the specific implementation of step S1 is as follows: Step S11, construct a multimodal message; Step S12, define a global policy planning tool; Step S13: Define the application window handle tool; Step S14: Define the UI recognition and image annotation tool; Step S15: Define the image stitching tool; Step S16: Define the interactive UI control tool; Step S17: Define the instant strategy planning tool; Step S18: Define the tool for verifying user operation tasks.

[0010] Furthermore, the VLM Tools set of the method is a key component of VLM Agent. The present invention splits and encapsulates the core content and logic of the application automatic control method into different visual language large model tools VLM Tool. Among them, VLM Tool is divided into Vision Task and Non-Vision Task in the present invention. VLMAgent selects and calls the appropriate VLM Tool according to the user prompt Prompt and the current automatic control task execution context. If the VLM Tool belongs to Vision Task, then VLM Agent constructs a multimodal message list , and takes the image information generated during the call process, such as screenshots of the application GUI interface, etc., and the text information together as the request content of the multimodal message, where , and then denote the VLM model used by VLM Agent as , and its parameters as , then the output of VLM, that is, the multimodal message passed by VLM Agent to the user, can be expressed as: (1); Since in multi-round long conversations, the token consumption of the system for each round of conversation when calling VLM increases linearly, and when the input content contains multimodal information such as pictures, due to limitations such as network bandwidth, the execution speed of VLM Agent is slow, and at the same time, the total number of tokens consumed by VLM is very high, which is not conducive to the Memory component storing the context of VLM Agent, making it difficult for VLM Agents to share memory and unable to persist in multi-round long conversations. To alleviate this phenomenon and reduce the total consumption of Prompt and Token, the present invention uses image stitching technology to stitch the image content output by VLM in each round of conversation and the image content generated during the intermediate process of VLM Agent (2); Among them, the stitching image stitching function sequentially stitches all the input image contents and outputs the image stitching result , and then, together with all the text contents generated by the VLM and the VLM Agent serve as the storage unit elements of the Memory component, enabling the decisions made by each VLM Agent to be open to each other for access, sharing historical contexts with each other, and at the same time ensuring that the number of images stored in the Memory component by the system is always 1 each time.

[0011] A higher UI control recognition accuracy rate can improve the visual sensitivity of the VLM to image marking, enabling the VLM to make correct immediate decisions better for the current user task operation stage. To explore the potential of the VLM in the field of application program automatic control, the present invention adopts a general rule element extraction method to identify UI control recognition, sacrificing the cross-platform nature of the CCMAgent method proposed by the present invention to improve the UI control recognition accuracy rate as much as possible, and transferring the error bottleneck from GUI-Grounding to the understanding of image content by the VLM and the design of the VLM Agent. However, when the number of UI controls appearing in a certain GUI interface of an application program is large, it is necessary to utilize the excellent ability of the VLM to understand text content to filter out irrelevant UI controls and only retain the main UI controls relevant to the current immediate decision. In addition, after identifying the UI controls, it is found that the operable areas of some UI controls vary greatly, such as multi-line text input boxes and ordinary buttons. To improve the VLM's image understanding ability of the GUI interface, the present invention not only marks the upper right corner of each interactive UI control with a pure digital ID, but also uses a red detection frame to enclose all the operable areas of these UI controls to highlight them, enabling the VLM to better understand the meaning and interaction area of each UI control, as shown in the appendix Figure 1 shown. The UFO method based on the general rule element extraction technology, which is similar to the idea of the present invention, only marks a pure digital ID in the upper right corner of the interactive UI control.

[0012] Before the VLM Agent interacts with the UI control, it is also necessary to determine the exact coordinate position of the UI interaction in the pixel coordinate system, which can be calculated through the coordinate positions of the four vertices on the boundary of the UI control: (3); where is the abscissa of the upper left vertex of the UI control boundary in the pixel coordinate system, is the abscissa of the upper right vertex of the UI control boundary in the pixel coordinate system, is the ordinate of the upper left vertex of the UI control boundary in the pixel coordinate system, is the ordinate of the lower left corner vertex of the UI control boundary in the pixel coordinate system.

[0013] When a UI control interaction operation is completed, in order to check whether the current user operation task is completed and better correct the global policy planning , the present invention designs a VLM Tool specifically for checking task completion. The VLM Agent will use this Tool to capture screenshots of the specified GUI interface of the specified application, and capture the GUI interface screenshots and the overall task description to encapsulate them into multimodal messages , and let the VLM Agent judge whether the task is completed. If not, it will re-identify the application UI and perform the next interaction operation, while correcting the global policy planning , and give the next immediate policy planning to the VLM Team , as shown in Equation (4): (4); where the global policy planning and the immediate policy planning are both encapsulated into multimodal messages, indicating the process of the VLM updating and correcting the global policy planning and the immediate policy planning . After the update is completed, the VLM Agent stores the policy information in the Memory component, so that all VLM Agents can share the corrected global policy planning and the immediate policy planning , in order to improve the VLM's image understanding of the current GUI interface. If it is completed, the application agent Application Agent outputs a terminator, ends this automatic control task, and notifies the user.

[0014] Furthermore, the specific implementation of step S2 is as follows: Step S21: Initialize the model interaction client Model Client; Step S22: Build a voice interaction component; Step S23: Define the scheduling strategy of the visual language large model team VLM Team; Step S24: Build the memory Memory component and the retrieval augmented generation RAG component; Step S25: Build the application agent Application Agent; Step S26: Build the user interface agent UI Agent; Step S27: Build the user task check agent Check Agent.

[0015] Most existing Agent frameworks are relatively cumbersome and not specifically designed for building Agents. For example, LangChain is difficult to be compatible with the scenarios of Multi-Agent and multiple VLM Tools, and it is also difficult to expand different latest commercial VLM APIs.

[0016] Furthermore, based on the AutoGen framework, a secondary encapsulation and development are carried out, and a simplified API is provided externally. Users only need to modify the system prompt words to complete the update of the entire CCMAgent system. It can also complete the scenarios of Multi-Agent and multiple VLM Tools at the same time, and can horizontally be compatible with a variety of the latest commercial VLM API interfaces, such as GPT-4V(o) and Claude-3.7-sonnet. In addition, in order to optimize the interaction experience when users operate this system, the present invention uses the STT tool to convert the way for users from inputting text to operating the VLM Agent through voice device conversations.

[0017] The present invention has designed a total of three VLM Agents for the problem of automatic control of application programs, namely Application Agent, UI Agent, and Check Agent, which is a typical Multi-Agent scenario. In order to enable these three VLM Agents to cooperate and schedule their work and share historical contexts, the present invention uses the VLM Team technology to share the Memory component among the three VLM Agents, and at the same time hands over the operation permission of the main thread of the method system to the VLM Team for management. Therefore, it is necessary to define a suitable VLM Team scheduling strategy to allocate the main thread to the corresponding VLM Agent at appropriate steps. Since the scheduling and use of various VLM Tools defined in step S1 of the present invention are based on semantics, it is necessary to determine the next VLM Tool to be called according to the multimodal output messages of the VLM Agent Therefore, the present invention defines the scheduling strategy of the VLM Team as that the VLM Agent currently in control of the main thread operation permission can choose to hand over the main thread to other VLM Agents or still keep it for itself, making the VLM Team more flexible and free when dealing with user control tasks, that is, the VLM Team selects a suitable VLM Agent according to the semantic information of the global policy planning or the immediate policy planning in the historical shared context.

[0018] The set of Visual Language Model (VLM) tools that the Application Agent can operate on, the VLM Tools set, consists of a global policy planning tool, an image stitching tool, and an application window handle tool, mainly for processing and generating global policy planning and constructing it into a multimodal message which is stored in the Memory component, and performing basic operations on the application, such as obtaining application environment variables, window handles, and starting the application; The set of Visual Language Model (VLM) tools that the UI Agent can operate on, the VLM Tools set, consists of a UI recognition and image annotation tool, an interactive UI control tool, and an image stitching tool, mainly for processing screenshots of the application GUI interface and UI-related steps, such as UI recognition, GUI detection box and control ID annotation, controlling UI elements, etc.; The set of Visual Language Model (VLM) tools that the Check Agent can operate on, the VLM Tools set, consists of an immediate policy planning tool, a tool for verifying user operation tasks, and an image stitching tool, mainly for processing and generating immediate policy planning which is also constructed into a multimodal message and stored in the Memory component, and sharing the policy with the Application Agent and the Check Agent.

[0019] Furthermore, the specific implementation of step S3 is as follows: Step S31: Perform unified operation on the VNC remote connection object; Step S32: Mark the optimal operation step sequence of the WindowsBench dataset; Step S33: Calculate the completion rate of the automatic control task; Step S34: Calculate the computer control score CC-Score; Step S35: Statistically calculate the average number of operation steps; Step S36: Statistically calculate the average time spent per step; Step S37: Statistically calculate the number of tokens consumed by the VLM.

[0020] Furthermore, in order to more comprehensively evaluate various application automatic control methods based on the VLM Agent, the present invention uses a total of 5 evaluation metrics, including the user automatic control task completion rate (Complete Rate), the computer control score CC-Score (Computer Control Score), the average number of steps, the average time spent per step, and the token situation consumed by the VLM. Among them, the user automatic control task completion rate takes into account the importance of the step sequence and the task completion degree, and calculates the ratio of the operation steps useful for task completion in the predicted operation sequence to the optimal operation steps: (5); Among them, the prediction operation sequence executed by the VLM Agent is , including operation steps; the optimal operation step sequence is , including operation steps. represents the set of operation sequences effective for completing the automatic control task. Although a certain prediction step in the prediction operation sequence is the same as a certain step in the optimal operation sequence, but if this prediction step is of no help in completing the task, then record . On the contrary, if the execution of the prediction step improves the progress of the user's automatic control task completion, then record . Calculating the user's automatic control task completion rate will increase the weight ratio of some prediction operation sequences that eventually complete the task despite having errors in the intermediate steps, which can reflect the exploratory, trial-and-error, and reflective abilities of the VLM Agent.

[0021] The CC-Score, on the other hand, considers the importance of operation step matching and step order. Taking the optimal operation sequence as the standard, it calculates the matching degree between the optimal operation sequence and the prediction operation sequence. That is, it calculates the number of steps that are the same when the indexes are the same in the two operation sequences, and finally calculates the ratio of the number of step matches to the total number of steps in the optimal operation sequence : (6); In terms of the dataset, the present invention uses the user task descriptions of Word, Excel, and three mainstream browsers in the WindowsBench benchmark as the method evaluation dataset, that is, the WindowsBench subset. This benchmark is applicable to most VLM Agent automatic control methods. Since the present invention uses the optimal operation sequence when calculating both the user's automatic control task completion rate and the CC-Score, and WindowsBench does not provide the optimal operation sequence but only provides user task descriptions, the present invention labels the optimal operation sequence for each user task in this dataset.

[0022] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention designs three VLM Agents, and uses the VLM Team technology to enable the collaborative scheduling of work among the VLM Agents. It adopts a decision-making method for VLM Agents that gives priority to immediate strategy planning and supplements it with global strategy planning, improving the exploration ability of VLM Agents and the generalization ability of the method, and is applicable to unknown application programs or user automatic control tasks with high operation difficulty. (2) The Multi-Agent framework independently built by the present invention is a secondary encapsulation based on the AutoGen framework, with a simple API. It can horizontally expand a variety of the latest commercial VLM API interfaces, such as GPT-4V(o) and Claude-3.7-sonnet, and can be compatible with both the Multi-Agent and multi-VLM Tool scenarios at the same time. (3) The recognition of UI controls in the GUI interface of the application program adopts a method based on the extraction of general rule elements, sacrificing the cross-platform nature of the present invention to improve the UI recognition accuracy as much as possible, exploring the potential of VLM Agents in the problem of application program automatic control. At the same time, the present invention uses image stitching technology to reduce the proportion of image information in the Prompt in multi-round long conversations, so as to improve the running speed of the present invention and alleviate the problem of too long shared historical context in Memory.

[0023] (4) The present invention uses a subset of WindowsBench to comprehensively evaluate and compare five indicators such as task completion rate and CC-Score between the GUI-Grounding-based method and the general rule element extraction-based method. Brief Description of the Drawings

[0024] Figure 1 It is a flowchart of the automatic control method of the CCMAgent application program designed by the present invention.

[0025] Figure 2 It is an interaction framework diagram of the VLM Team component implemented by the present invention.

[0026] Figure 3 It is a diagram showing the proportion of UI operations in the optimal operation sequence labeled for the WindowsBench subset by the present invention. Detailed Embodiment

[0027] Next, in combination with the drawings, the technical solutions of the present invention will be specifically described.

[0028] The present invention provides an application program automatic control method CCMAgent based on a vision-language large model agent. It uses VLM Team technology to collaboratively schedule 3 independently designed VLM Agents and 7 VLM Tools, and adopts an Agent decision-making method with immediate strategy planning as the main and global strategy planning as the auxiliary to improve the generalization ability and versatility of the method, providing ideas for automatically controlling unknown software and complex control tasks. In order to improve the recognition of UI controls and the annotation accuracy of their interactive area detection frames as much as possible, the present invention replaces GUI-Grounding based on the extraction method of general rule elements, explores the potential of VLM Agents in the problem of application program automatic control, and uses image stitching technology to stitch the multi-modal messages generated during the intermediate process of VLM Agents, reducing the proportion of image information in the Prompt in multi-round long conversations, improving the running speed of the CCMAgent method, and at the same time alleviating the problem of too long shared historical context in the Memory component. Finally, the present invention comprehensively compares and evaluates the application program automatic control method by calculating 5 indicators using a subset of WindowsBench.

[0029] The following is the implementation process of the specific embodiments of the present invention.

[0030] The flowchart and component interaction framework diagram of an application program automatic control method CCMAgent based on a vision-language large model agent proposed by the present invention are shown in Figure 1 and Figure 2 . The present invention includes the following steps: Step S1: Construct a vision-language large model tool set VLM Tools set; Step S2: Construct a vision-language large model team VLM Team and a vision-language large model agent VLM Agent; Step S3: Evaluate the application program automatic control method; Since the method based on the extraction of general rule elements cannot identify UI controls across platforms, in order to facilitate the subsequent method evaluation steps, the present invention uses Windows 10 VNC with a graphical interface as an automatic control embodiment. In addition, the VLM deployment and inference process of various automatic control methods are completed on an L20 GPU with a video memory capacity of 48GB and a memory capacity of 120GB, while the commercial VLM API uses the OpenAI platform.

[0031] (1) Construction of the VLM Tools set in Step 1 The core steps of the application automatic control method of the present invention are split into a set of VLM Tools that can be executed by the VLM Agent. It is one of the key components in the VLM Agent and is divided into Vision Task and Non-Vision Task. For the VLM Agent that is performing an operation task, it can select one or more appropriate VLM Tools to execute according to the historical context semantics and image information. The principle is JSON serialization, and the execution result of the VLMTool will also be used as part of the historical context stored in the Memory component and shared with other VLM Agents in the same VLM Team. Therefore, the definition and construction of the VLM Tools set will directly determine the accuracy and performance of the application automatic control method based on the VLMAgent.

[0032] Step 1.1: Construct multimodal messages: When the VLM Agent executes a VLM Tool related to the Vision Task, since the VLM Tools defined in the present invention will generate image information during the call process , such as screenshots of the application GUI interface, GUI screenshots after annotating detection boxes, etc. Therefore, the VLM Agent constructs a list of multimodal messages , and takes all the image information and text information together as the request content of the multimodal message . Among them , and then denote the VLM model used by the VLM Agent as , and its parameters are . Then the output of the VLM, that is, the multimodal message passed by the VLMAgent's answer to the user, can be expressed as: (1); Step 1.2: Define the global policy planning tool: The global policy planning tool is generally executed in the first step of the user automatic control task and is called by the ApplicationAgent. The global policy planning tool will receive the Prompt input by the user and split the automatic control task included in the user Prompt into each sub-step that can be directly executed through the UI control through VLM semantic understanding, and obtain the global policy planning . After splitting the core sub-steps, the global policy planning will also be temporarily saved in the Memory component so that the UI Agent and the Check Agent can share this policy planning for subsequent steps.

[0033] Step 1.3: Define the application window handle tool: The application window handle tool includes two main functions. One is to parse the application name contained in the user input Prompt and the global policy plan, and find the corresponding environment variable or startup path. The other is to obtain the graphical interface window handles of all currently active applications, which is used to switch the currently topmost application to the specified application window. The application window handle tool is scheduled and executed by the Application Agent. In order to test the robustness of the method, there are multiple interfering desktop windows during this evaluation experiment. At the same time, the window information will be stored in a relational database. In this invention, MySQL is used instead of a vector database for the RAG component to retrieve and enhance the VLM Agent as part of the shared historical context.

[0034] Step 1.4: Define the UI recognition and image annotation tool: The UI recognition and graphical annotation Tool is scheduled and executed by the UI Agent. The main function of this VLM Tool is to take a screenshot of the GUI interface of the current application, then recognize the UI control operations on this GUI screenshot, and mark all operable UI controls with a red detection box and a pure digital ID of the control. After the UI recognition and image annotation tool recognizes all UI controls, it will also persist all UI control information, such as name, coordinates, control ID, etc., to MySQL for the RAG component to retrieve and enhance the VLM Agent.

[0035] ​A higher UI control recognition accuracy rate can improve the visual sensitivity of the VLM to image marking, enabling the VLM to make more accurate immediate decisions for the current user task operation stage. To explore the potential of the VLM in the field of application automatic control, the present invention adopts a general rule element extraction method to identify UI controls, sacrificing the cross-platform nature of the CCMAgent method proposed by the present invention to improve the recognition accuracy rate of UI controls as much as possible, shifting the error bottleneck from GUI-Grounding to the understanding of image content by the VLM and the design of the VLM Agent. However, when there are a large number of UI controls in a certain GUI interface of an application, it is necessary to use the excellent ability of the VLM to understand text content to filter out irrelevant UI controls and only retain the main UI controls relevant to the current immediate decision. In addition, after identifying the UI controls, it is found that the operable areas of some UI controls vary greatly, such as multi-line text input boxes and ordinary buttons. To improve the VLM's image understanding ability of the GUI interface, the present invention not only marks each interactive UI control with a pure digital ID in the upper right corner, but also uses a red detection frame to enclose the operable areas of these UI controls to highlight them, enabling the VLM to better understand the meaning and interaction area of each UI control, as shown in Figure 1 shown. The UFO method proposed by Zhang et al., which is similar to the idea of the present invention based on the general rule element extraction technology, only marks a pure digital ID in the upper right corner of the interactive UI controls.

[0036] Step 1.5: Define an image stitching tool: Since image information will be generated during the process of the VLM Agent outputting answers and executing the VLM Tool with visual tasks, the Application Agent, UI Agent, and Check Agent can all schedule the image stitching tool to perform stitching operations on the multi-modal message list At the same time, during multi-round long conversations, the total number of tokens consumed by the system each round when calling the VLM increases linearly. And when the input content contains multi-modal information such as pictures, due to limitations such as network bandwidth, the execution speed of the VLM Agent is slow, and the total number of tokens consumed by the VLM is very high, which is not conducive to the memory component (Memory) storing the context of the VLM Agent, making it difficult for the VLM Agents to share memory and unable to sustain multi-round long conversations. To alleviate this phenomenon and reduce the total consumption of Prompts and Tokens, the present invention uses image stitching technology to stitch the image content output by the VLM in each round of conversation and the image content generated during the intermediate process of the VLM Agent in the order of message generation: (2); Among them, the stitching image stitching function stitches all the input image contents in sequence and outputs the image stitching result , the image stitching result and all the text contents generated by VLM and VLM Agent together serve as the storage unit element of the Memory component, enabling the decisions made by each VLM Agent to be open to each other for access, sharing historical context with each other, and at the same time ensuring that the number of images stored in the Memory component by the system each time is always 1.

[0037] Step 1.6: Define the interactive UI control tool: Before the VLM Agent interacts with the UI control, it first needs to know the coordinates of the UI control. The coordinates of the UI control are obtained from the UI recognition and image annotation tool and have been persisted in the MySQL database, and can be directly retrieved through the RAG component by the control ID. The interactive UI control tool is scheduled and used by the UI Agent. After the UI Agent obtains the coordinates of the specified UI control through this Tool, it also needs to calculate the interactive coordinates of the UI control: (3); Among them is the abscissa of the upper left vertex of the UI control boundary in the pixel coordinate system, is the abscissa of the upper right vertex of the UI control boundary in the pixel coordinate system, is the ordinate of the upper left vertex of the UI control boundary in the pixel coordinate system, is the ordinate of the lower left vertex of the UI control boundary in the pixel coordinate system.

[0038] Step 1.7: Define the immediate policy planning tool: The immediate policy planning tool is called by the Check Agent. After each UI control operation is executed, the present invention specifies through the semantic meaning of the system prompt that the Check Agent needs to schedule the immediate policy planning tool and the tool for verifying the user operation task once respectively to determine the immediate policy planning for the next operation , and at the same time according to the immediate policy planning determine the VLM Agent for the next operation to be executed and temporarily store it in the shared historical context of the Memory component for correcting the global policy planning , that is, initialize the error of the policy planning.

[0039] Step 1.8: Define the tool for verifying the user operation task: As can be seen from step 17, the inspection of the user operation task tool is also scheduled and executed by the Check Agent. When a UI control interaction operation is completed, in order to check whether the current user operation task is completed and better correct the global policy planning , the present invention designs a VLM Tool specifically for checking task completion. The VLM Agent will use this Tool to take screenshots of the specified GUI interface of the specified application, and take screenshots of the GUI interface and the overall task description and encapsulate them into multimodal messages , so that the VLM Agent can judge whether the task is completed. If not, it will re-identify the application UI and perform the next interaction operation, and at the same time correct the global policy planning , and give the next immediate policy planning to the VLM Team , as shown in formula (4): (4); Among them, the global policy planning and the immediate policy planning are both encapsulated into multimodal messages, indicating the process of the VLM updating and correcting the global policy planning and the immediate policy planning . After the update, the VLM Agent stores the policy information in the Memory component, so that all VLM Agents can share the corrected global policy planning and the immediate policy planning , in order to improve the VLM's image understanding of the current GUI interface

[0040] (2) Building the VLM Team and Agent in step 2 The VLM Team is a scheduling mechanism technology that can coordinate multiple VLM or LLM Agents and can define various scheduling strategies. The scheduling strategy of the VLM Team of the present invention is defined as that the VLM Agent currently performing the operation task can decide which VLM Agent to hand over the initiative of the next automatic control operation to, which can be itself to continue to execute cyclically, or any other VLM Agent. In addition, most of the existing Agent frameworks are relatively cumbersome and not specifically designed for building Agents. For example, LangChain is difficult to be compatible with the Multi-Agent and multi-VLM Tool scenarios, and it is difficult to expand different latest commercial VLM APIs. Therefore, the present invention is developed by secondary encapsulation based on AutoGen, and provides a simplified API externally. Users only need to modify the system prompt words to complete the update of the entire CCMAgent system, and can also complete the Multi-Agent and multi-VLM Tool scenarios, and can be horizontally compatible with a variety of latest commercial VLM API interfaces, such as GPT-4V(o) and Claude-3.7-sonnet.

[0041] Step 2.1: Initialize the Model Client:

[0042] The Model Client is one of the basic components of the VLM Team and is also an agent component connecting to external VLMs, such as GPT-4V(o) or Claude-3.7-sonnet on the OpenAI platform. Before running the CCMAgent method, it is first necessary to initialize the Model Client component. Complex parameters, such as api_key, base_url, model_type, etc., are simplified through Yml configuration or default parameters, and at the same time, it can be compatible with advanced commercial VLM API interfaces, enabling it to have the ability to format the output using Tools, which is convenient for users to use.

[0043] Step 2.2: Build the voice interaction component: To facilitate blind interaction and simplify user interaction, and eliminate the need for users to manually enter Prompts, the present invention uses STT tools to implement voice interaction functions, converting the way for users to operate the VLM Agent from entering text to conversing through voice devices. In addition, when the VLM Agent outputs an answer to the user, the text content output by the VLM Agent can also be converted into voice output through TTS tools.

[0044] Step 2.3: Define the VLM Team scheduling strategy: The present invention has designed a total of three VLM Agents for the problem of automatic control of application programs, namely Application Agent, UI Agent, and Check Agent, which is a typical Multi-Agent scenario. In order to enable these three VLM Agents to cooperate and schedule their work and share historical context, the present invention uses the VLM Team technology to share the Memory component among the three VLM Agents, and at the same time hands over the main thread operation permission of the method system to the management of the VLM Team. Therefore, it is necessary to define a suitable VLM Team scheduling strategy to allocate the main thread to the corresponding VLM Agent at the appropriate step. Since the scheduling and use of various VLM Tools defined in step S1 of the present invention are based on semantics, it is necessary to determine the next VLM Tool to be called according to the multi-modal output message of the VLM Agent to determine the next VLM Tool to be called. Therefore, the present invention defines the scheduling strategy of the VLM Team as allowing the VLM Agent that currently holds the main thread operation permission to choose whether to hand over the main thread to other VLM Agents or still keep it for itself, making the VLM Team more flexible and free when processing user control tasks, that is, the VLM Team selects the appropriate VLM Agent according to the semantic information of the global strategy planning or the immediate strategy planning in the historical shared context

[0045] Step 2.4: Construct the Memory and RAG components: Both the Memory and RAG components are one of the key components of the VLM Team. The Memory component is responsible for storing all the multi-modal messages generated by all VLM Agents , where is the set of image information after image stitching, and is the set of all text information and broadcasts and shares it to the system prompt words of all VLM Agents; the role of the RAG component is to retrieve and enhance the VLM Agent, that is, to build a knowledge base or search engine to provide external knowledge for the VLM Agent. In the present invention, the RAG is implemented by storing all the identified UI control information, including control names, control IDs, coordinates, etc., and all application window handle information in the relational database MySQL. Therefore, the RAG component can be used when the VLM Agent calls any VLM Tool to enhance the function of the VLM Tool

[0046] Step 2.5, Step 26, Step 27: Construct the Application Agent, UI Agent, and Check Agent: The present invention has designed a total of three VLM Agents for the problem of automatic control of application programs, namely Application Agent, UI Agent, and Check Agent, which is a typical Multi-Agent scenario. In order to enable these three VLM Agents to cooperate and schedule their work and share historical context, the present invention uses the VLM Team technology to share the Memory component among the three VLM Agents, and at the same time hands over the main thread operation permission of the method system to the VLM Team for management. Therefore, it is necessary to define a suitable VLM Team scheduling strategy to allocate the main thread to the corresponding VLM Agent at the appropriate step. Since the scheduling and use of various VLM Tools defined in step S1 of the present invention are based on semantics, it is necessary to determine the next VLM Tool to be called according to the multimodal output message of the VLM Agent Therefore, the present invention defines the scheduling strategy of the VLM Team as allowing the VLM Agent that currently holds the main thread operation permission to choose whether to hand over the main thread to other VLM Agents or still keep it for itself, making the VLM Team more flexible and free when processing user control tasks, that is, the VLM Team selects the appropriate VLM Agent according to the semantic information of the global policy planning or the immediate policy planning in the historical shared context.

[0047] The set of VLM Tools that the Application Agent can operate on includes the global policy planning tool, the image stitching tool, and the application window handle tool, which are mainly used to generate the global policy planning and construct it into a multimodal message and store it in the Memory component, as well as perform basic operations on the application program, such as obtaining the application program environment variables, window handles, and starting the application program; the set of VLM Tools that the UI Agent can operate on includes the UI recognition and image annotation tool, the interactive UI control tool, and the image stitching tool, which are mainly used to capture screenshots of the application program GUI interface and perform UI-related steps, such as UI recognition, GUI detection box and control ID annotation, and controlling UI elements; the set of VLM Tools that the Check Agent can operate on includes the immediate policy planning tool, the tool for verifying user operation tasks, and the image stitching tool, which are mainly used to generate the immediate policy planning and also construct it into a multimodal message and store it in the Memory component, and share the policy with the Application Agent and the Check Agent.

[0048] (3)Evaluation of the Application Program Automatic Control Method in Step 3 To comprehensively compare and evaluate various advanced application program automatic control methods based on VLM, the present invention uses a subset of the WindowsBench dataset as the evaluation dataset, which contains user task descriptions of Word, Excel, and three mainstream browsers (Google, Edge, Firefox), and uses 5 evaluation metrics to comprehensively analyze these automatic control methods.

[0049] Step 3.1: Unified operation object for VNC remote connection Since the automatic control target example used in the automatic control method evaluation experiment of the present invention is Windows 10 VNC with a graphical interface, before conducting the evaluation experiment, it is also necessary to set up a VNC client and a server. Set up a VNC server on the target example, and set up a VNC client on the L20 GPU for the VLM Agent to operate the VNCGUI interface. Note that Docker VNC virtual Windows 10 can be used, but in this experiment, a physical machine is used as the target example.

[0050] Step 3.2: Mark the optimal operation step sequence of the WindowsBench dataset In terms of the dataset, the present invention uses the user task descriptions of Word, Excel, and three mainstream browsers in the WindowsBench benchmark proposed by Zhang et al. as the method evaluation dataset, that is, the WindowsBench subset. This benchmark is applicable to most VLM Agent automatic control methods. Since the present invention uses the optimal operation sequence when calculating the user automatic control task completion rate and CC-Score , and WindowsBench does not provide the optimal operation sequence, but only provides user task descriptions, so the present invention marks the optimal operation sequence for each user task of this dataset. The proportion of the operation types of UI controls in the optimal operation sequence is as shown in the appendix Figure 3 As shown, the following provides a Prompt example for operating Excel and the optimal operation sequence marked by the present invention for it: Prompt: Open Excel application and input "100.0" into the cell at the row 3, column 2. Finally please bold formatting to the cell. 1.Open Excel. 2.Single Click on the cell at row 3, column 2. 3.Enter "100.0". 4.Single Click the "Bold" button. Step 3.3: Calculate the completion degree of the automatic control task To more comprehensively evaluate various automatic control methods for VLM Agent-based applications, the present invention uses a total of 5 evaluation metrics, including the user's automatic control task completion rate (Complete Rate), CC-Score (Computer Control Score), average number of steps, average time spent per step, and the token situation consumed by VLM. Among them, the user's automatic control task completion rate takes into account the importance of step order and task completion, and calculates the ratio of the operation steps useful for task completion in the predicted operation sequence to the optimal operation steps: (5); where the predicted operation sequence executed by the VLM Agent is , including operation steps; the optimal operation step sequence is , including operation steps. represents the set of operation sequences effective for completing the automatic control task. Although a certain predicted step in the predicted operation sequence is the same as a certain step in the optimal operation sequence, if this predicted step is not helpful for completing the task at all, then record , on the contrary, if executing the predicted step improves the progress of the user's automatic control task completion, then record .

[0051] Calculating the user's automatic control task completion rate will increase the weight ratio of some predicted operation sequences that still complete the task despite errors in the intermediate steps, and can reflect the exploratory, trial-and-error, and reflection abilities of the VLM Agent.

[0052] Step 3.4: Calculate the computer control score CC-Score: CC-Score takes into account the importance of operation step matching and step order, and takes the optimal operation sequence as the standard to calculate the matching degree between the optimal operation sequence and the predicted operation sequence. That is, calculate the number of steps that are the same when the indexes in the two operation sequences are the same, and finally calculate the ratio of the number of step matches to the total number of steps of the optimal operation sequence: (6); where, represents the predicted operation sequence of this task; represents the optimal operation sequence of this task; is used to judge when the operation sequence index is , the predicted operation Is it the same as the optimal operation If it is the same, its value is set to 1; if it is different, it is set to 0; In terms of the dataset, the present invention uses the user task descriptions of Word, Excel, and three mainstream browsers in the WindowsBench benchmark proposed by Zhang et al. as the method evaluation dataset, namely the WindowsBench subset. This benchmark is applicable to most VLM Agent automatic control methods. Since the present invention uses the optimal operation sequence when calculating the user automatic control task completion rate and CC-Score and WindowsBench does not provide the optimal operation sequence, but only provides user task descriptions, the present invention labels the optimal operation sequence for each user task in this dataset.

[0053] The present invention uses a total of two types of VLM-based application automatic control methods in this comparative evaluation experiment. Table 1 below shows the comparison results of the application automatic control method effects. Table 1 Comparison Results of Application Automatic Control Method Effects

[0054] They are ScreenAgent proposed by Niu et al. and CogAgent proposed by Hong et al. based on VLM-Grouding, and UFO method proposed by Zhang et al. and CCMAgent+Claude-3.7-sonnet and CCMAgent+GPT-4V(o) method proposed by the present invention based on general rule element extraction. In Table 1, Agent Amount represents the number of VLM Agents used in this method; VLM represents the VLM used in this method; UI Recognition represents the UI recognition and positioning method used in this method, Object Grounding is a method based on GUI-Grounding, and ElementExtraction is a method based on general rule element extraction; Step represents the total average number of steps required to execute a single complete user task in the WindowsBench subset; Complete Rate represents the user task completion degree, as shown in Equation (5); CC-Score is used to calculate the matching degree between the optimal operation sequence and the predicted operation sequence, as shown in Equation (6); Time represents the average time required for the VLM Agent to execute a complete step, such as a mouse click; Prompt represents the average number of prompt words required to execute a single complete user task in the WindowsBench subset, which can indirectly reflect the Token consumption.

[0055] As can be seen from Table 1, both the Complete Rate and CC-Score based on the mainstream GUI-Grounding method are lower than those based on the general rule element extraction method. The reason may be that the GUI-Grounding method directly identifies and locates the coordinates of UI controls in the GUI interface screenshot through the VLM Agent, and its accuracy is not as good as that of the general rule element extraction method. In addition, the set of Agents and Tools designed for each method must be different. Moreover, since ScreenAgent and CogAgent use pre-trained VLMs instead of commercial VLM API services, these reasons combined lead to poor performance. At the same time, it can be found that the total number of steps taken by the methods based on UFO and the present invention is more than that of CogAgent and ScreenAgent, and both are greater than 10 steps. This may be because CogAgent and ScreenAgent both use only global policy planning, resulting in fixed execution steps and less exploration of the VLM Agent.

[0056] In addition, ScreenAgent takes the longest time per step on average, and the gap with other methods is relatively large. This may be because its model occupies a large amount of video memory, and the time for the VLM to output answers is very slow. CogAgent takes the shortest time, probably because the average total number of steps for it to execute all tasks is the least. The two methods of CCMAgent proposed in the present invention take a slightly longer time than UFO, perhaps because the number of VLM Agents in the method of the present invention is more than that of UFO, and the shared historical context is more than that of UFO. However, it can also be found that the number of Prompts consumed by the method of the present invention is reduced by 52% - 63% compared with UFO, indicating that the image stitching technology adopted in the present invention has an obvious effect on reducing the proportion of images in Prompts. The Complete Rate and CC-Score are respectively reduced by 4.9% and 2.4% compared with UFO, and are increased by 19.2% and 7.6% compared with CogAgent. Considering comprehensively, if only the accuracy of completing the user's automatic control task is considered, then the UFO method is selected. From an economic perspective, the CCMAgent+Claude-3.7-sonnet proposed in the present invention can be selected. Although CCMAgent+Claude-3.7-sonnet consumes fewer Prompts, its accuracy is not as good as that of CCMAgent+GPT-4V(o). Therefore, if both economy and accuracy are considered, the CCMAgent+GPT-4V(o) proposed in the present invention can be selected.

[0057] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. An automatic control method for application programs based on a visual language large model intelligent agent, characterized in that, The framework of the method designs three types of visual language model agents (VLM Agents) with different roles, namely the Application Agent, the User Interface Agent (UI Agent), and the User Task Check Agent; After the user inputs a task description prompt or voice, the Application Agent is first responsible for parsing the user input and splitting the user operation task into a series of executable UI control operations of the user interface, i.e., global policy planning; After that, the Application Agent obtains the corresponding window handle from the environment variables or configuration files according to the specific name of the application extracted from the global policy planning, i.e., starts the specified application, and transfers the VLM execution leadership to the User Interface Agent (UI Agent); the UI Agent uses the designed user interface tool set (UI Tools) to take a screenshot of the GUI interface of the current application window, and combines the external graphical user interface localization GUI-Grounding method to identify all operable UI control elements appearing in the current GUI interface; in order to enhance the VLM's understanding of the GUI image, the UI Agent will label the detection box and the unique control identification number (ID) for all the identified UI control elements after the UI control element recognition of the GUI interface, using the visual perception characteristics of the VLM that is sensitive to image annotation; at the same time, the UI Agent gives an immediate policy planning based on the global policy planning and the screenshot of the GUI interface after the detection box annotation, selects the UI control element to be operated in the current step, and corrects the global policy planning; Finally, the User Task Check Agent will judge whether the user's current task has been completed according to the screenshot of the GUI interface after operating the specified UI control element. If not, it will continue to transfer the VLM execution leadership to the UI Agent for the next UI control element operation; if completed, the Application Agent outputs a terminator, ends the current automatic control task, and notifies the user.

2. The automatic control method of an application program based on a vision-language large model agent according to claim 1, wherein The method includes the following steps: Step S1: Construct a visual language model tool set (VLM Tools); Step S2: Construct a visual language model team (VLM Team) and visual language model agents (VLM Agents); Step S3: Evaluate the application automatic control method.

3. The automatic control method of an application program based on a vision-language large model agent according to claim 2, characterized in that, The specific implementation of Step S1 is as follows: Step S11: Construct multimodal messages; Step S12: Define global policy planning tools; Step S13: Define application window handle tools; Step S14: Define UI recognition and image annotation tools; Step S15: Define image stitching tools; Step S16: Define interactive UI control tools; Step S17: Define immediate policy planning tools; Step S18: Define tools for verifying user operation tasks.

4. The automatic control method of an application program based on a vision-language large model agent according to claim 3, characterized in that, The visual language large model tool set VLM Tools of the described method is a key component of the visual language large model agent VLMAgent. The core content and logic of the application automatic control method are split and encapsulated into different visual language large model tools VLM Tool. Among them, the visual language large model tool VLM Tool is divided into vision tasks Vision Task and non-vision tasks Non-Vision Task. The visual language large model agent VLM Agent selects and invokes the appropriate visual language large model tool VLM Tool according to the user prompt Prompt and the current automatic control task execution context. If the visual language large model tool VLM Tool belongs to the vision task Vision Task, the visual language large model agent VLM Agent constructs a multi-modal message list , and the image information generated during the call , together with the text information , are used as the request content of the multi-modal message. Among them , and then record the VLM model used by the visual language large model agent VLMAgent as , and its parameters are . Then the output of VLM, that is, the multi-modal message answered by the visual language large model agent VLMAgent and passed to the user, can be expressed as: (1); Since during multi-turn long conversations, the token consumption of the VLM called by the system in each turn of the conversation grows linearly, and when the input content contains multimodal information such as images, due to limitations such as network bandwidth, the execution speed of the Visual Language Model Agent (VLM Agent) is slow. At the same time, the total number of tokens consumed by the VLM is very high, which is not conducive to the Memory component storing the context of the Visual Language Model Agent (VLM Agent), making it difficult for the Visual Language Model Agents (VLM Agents) to share memory and unable to sustain multi-turn long conversations. To alleviate this phenomenon and reduce the total consumption of prompts and tokens, using image stitching technology, the image content output by the VLM in each turn of the conversation and the image content generated during the intermediate process of the Visual Language Model Agent (VLM Agent) are stitched together in the order of message generation: (2); Among them, the stitching image stitching function sequentially stitches all the input image contents and outputs the image stitching result , and all the text contents generated by the VLM and the visual language large model agent VLM Agent together serve as the storage unit elements of the Memory component, enabling the decisions made by each visual language large model agent VLM Agent to be open and accessible to each other, sharing historical contexts with each other, and at the same time ensuring that the number of images stored in the Memory component by the system is always 1 each time; A higher accuracy rate of UI control recognition can improve the visual sensitivity of VLM to image tagging, enabling VLM to make more accurate and immediate decisions for the current user task operation stage. To explore the potential of VLM in the field of application automatic control, a general rule element extraction method is adopted to identify UI controls, sacrificing the cross-platform nature of the proposed multi-agent CCMAgent method for application automatic control to maximize the accuracy rate of UI control recognition. This shifts the error bottleneck from the external graphical user interface localization GUI-Grounding to the understanding of image content by VLM and the design of the visual language large model agent VLMAgent. When there are a large number of UI controls in a certain GUI interface of an application, the excellent ability of VLM to understand text content is utilized to filter out irrelevant UI controls, retaining only the main UI controls relevant to the current immediate decision. After identifying the UI controls, it is found that the operable areas of some UI controls vary significantly. To improve VLM's image understanding ability of the GUI interface, each interactive UI control is marked with a pure digital ID in the upper right corner, and the operable areas of these UI controls are highlighted by wrapping them with a red detection box, enabling VLM to better understand the meaning and interaction area of each UI control. In contrast, the UFO method based on the general rule element extraction technology only marks a pure digital ID in the upper right corner of the interactive UI controls. Before the visual language large model agent VLM Agent interacts with UI controls, it is also necessary to determine the exact coordinate position of UI interaction in the pixel coordinate system, which can be calculated based on the coordinate positions of the four vertices on the boundary of the UI control. (3); wherein is the abscissa of the upper left vertex of the boundary of the UI control in the pixel coordinate system, is the abscissa of the upper right vertex of the boundary of the UI control in the pixel coordinate system, is the ordinate of the upper left vertex of the boundary of the UI control in the pixel coordinate system, is the ordinate of the lower left vertex of the boundary of the UI control in the pixel coordinate system; When a UI control interaction operation is completed, in order to verify whether the current user operation task is completed and better correct the global policy planning , a visual language large model tool VLM Tool dedicated to verifying task completion is designed. The visual language large model agent VLM Agent will use this Tool to take screenshots of the specified GUI interface of the specified application, and take screenshots of the GUI interface and the overall task description and encapsulate them into multimodal messages , so that the visual language large model agent VLMAgent can judge whether the task is completed. If not, it will re-identify the application UI and perform the next interaction operation, and at the same time correct the global policy planning , and give the next immediate policy planning to the visual language large model team VLM Team , as shown in Equation (4): (4); Among them, the global policy planning and the immediate policy planning are both encapsulated as multimodal messages, indicating the process of the VLM updating and correcting the global policy planning and the immediate policy planning ; after the update, the Visual Language Model Agent (VLM Agent) stores the policy information in the Memory component, enabling all VLM Agents to share the corrected global policy planning and the immediate policy planning , in order to improve the VLM's understanding of the images on the current GUI interface.

5. The automatic control method of an application program based on a vision-language large model agent according to claim 2, wherein Step S2 is specifically implemented as follows: Step S21: Initialize the model interaction client Model Client; Step S22: Build the voice interaction component; Step S23: Define the scheduling strategy of the visual language large model team VLM Team; Step S24: Build the memory Memory component and the retrieval-augmented generation RAG component; Step S25: Build the application agent Application Agent; Step S26: Build the user interface agent UI Agent; Step S27: Build the user task check agent Check Agent.

6. The automatic control method of an application program based on a vision-language large model agent according to claim 5, wherein, Based on the AutoGen framework, secondary encapsulation and development are carried out, providing a simplified API externally. Users only need to modify the system prompt words to complete the update of the entire CCMAgent system, and can also complete scenarios of multi-agent Multi-Agent and multi-visual language large model tools VLM Tool simultaneously, and can be horizontally compatible with a variety of the latest commercial VLM API interfaces. To optimize the interaction experience when users operate this system, the STT tool is used to convert the user's input method from text to dialogue through a voice device to operate the visual language large model agent VLM Agent. A total of three Visual Language Model Agents (VLM Agents) are designed for the automatic control of applications, namely the Application Agent, the User Interface Agent (UI Agent), and the User Task Check Agent. This is a typical Multi-Agent scenario. To enable these three VLM Agents to work collaboratively and share historical context, the VLM Team technology is used to share the Memory component of the three VLM Agents. At the same time, the main thread operation permissions of the method system are handed over to the VLM Team for management. It is necessary to define an appropriate scheduling strategy for the VLM Team to allocate the main thread to the corresponding VLM Agent at the appropriate step. Since the scheduling and use of various Visual Language Model Tools (VLM Tools) defined in step S1 are based on semantics, it is necessary to determine the next VLM Tool to be called according to the multimodal output message of the VLM Agent. to determine the next VLM Tool to be called. The scheduling strategy of the VLM Team is defined as allowing the VLM Agent currently in control of the main thread to choose whether to hand over the main thread to another VLM Agent or keep it for itself, making the VLM Team more flexible and free when handling user control tasks. That is, the VLM Team selects the appropriate VLM Agent according to the global policy planning or immediate policy planning semantic information in the historical shared context. The collection of Visual Language Model Tools (VLM Tools) that the Application Agent can operate on consists of a global policy planning tool, an image stitching tool, and an application window handle tool, which mainly handle the generation of global policy planning , and construct it into a multimodal message Store it in the Memory component and perform basic application operations; The collection of Visual Language Model Tools (VLM Tools) that the User Interface Agent (UI Agent) can operate on consists of a UI recognition and image annotation tool, an interactive UI control tool, and an image stitching tool, which mainly handle screenshots of the application GUI interface and UI-related steps; The set of Visual Language Model Tools (VLM Tools) available for the User Task Checking Intelligent Agent (Check Agent) consists of an immediate policy planning tool, a user operation task checking tool, and an image stitching tool, mainly for processing and generating immediate policy planning. It is also constructed into a multimodal message. And it is stored in the Memory component. The policy is shared with the Application Agent and the User Task Checking Intelligent Agent (Check Agent).

7. The automatic control method for an application program based on a vision-language large model agent according to claim 2, characterized in that, Step S3 is specifically implemented as follows: Step S31: Use VNC to remotely connect to the unified operation object; Step S32: Mark the optimal operation step sequence of the WindowsBench dataset; Step S33: Calculate the completion degree of the automatic control task Step S34: Calculate the computer control score CC-Score; Step S35: Statistically analyze the average number of operation steps; Step S36: Statistically analyze the average time spent per step; Step S37: Statistically analyze the number of tokens Token consumed by the VLM.

8. The automatic control method of an application program based on a vision-language large model agent according to claim 7, characterized in that, In order to more comprehensively evaluate various application automatic control methods of vision-language large model agents VLM Agent, a total of 5 evaluation metrics are used, including the user automatic control task completion rate Complete Rate, the computer control score CC-Score, the average number of steps, the average time spent per step, and the token Token situation consumed by the VLM; among them, the user automatic control task completion rate takes into account the importance of the step sequence and the task completion degree, and calculates the ratio of the operation steps useful for task completion in the predicted operation sequence to the optimal operation steps: (5); The prediction operation sequence executed by the Visual Language Model Agent (VLM Agent) is , which contains operation steps; the optimal operation step sequence is , which contains operation steps; represents the set of operation sequences effective for completing the automatic control task; although a certain prediction step in the prediction operation sequence is the same as a certain step in the optimal operation sequence, if this prediction step is not helpful for completing the task at all, then mark ; on the contrary, if executing the prediction step improves the progress of the user's automatic control task completion, then mark ; calculating the user's automatic control task completion rate will increase the weight ratio of some prediction operation sequences that still complete the task despite errors in the intermediate steps, which can reflect the exploratory, trial-and-error, and reflection abilities of the Visual Language Model Agent (VLM Agent); The computer control score CC-Score takes into account the matching of operation steps and the importance of the step sequence. Based on the optimal operation sequence, it calculates the matching degree between the optimal operation sequence and the predicted operation sequence; that is, it calculates the number of steps that are the same when the indices are the same in the two operation sequences, and finally calculates the ratio of the number of step matches to the total number of steps in the optimal operation sequence ratio: (6); Among them, represents the predicted operation sequence of the task; represents the optimal operation sequence of the task; is used to judge that when the operation sequence index is , the predicted operation and the optimal operation are the same. If they are the same, its value is set to 1. If they are different, it is set to 0; In terms of the dataset, the user task descriptions of Word, Excel, and three mainstream browsers in the WindowsBench benchmark are used as the dataset for method evaluation, namely the WindowsBench subset; this benchmark is applicable to most visual language large model agents (VLM Agents) automatic control methods; since the optimal operation sequence is used when calculating the user automatic control task completion rate and the computer control score (CC-Score), and WindowsBench does not provide the optimal operation sequence but only the user task description, so the optimal operation sequence is labeled for each user task in this dataset.

Citation Information

Patent Citations

  • Intelligent question and answer method and device, computer equipment and program product

    CN118820436A

  • Multi-agent cooperation method and system

    CN118917632A

  • Large language model automatic penetration testing method based on multiple agents

    CN119025878A

  • Multi-round dialogue system, method, device, medium and program product of body-equipped agent

    CN119378541A

  • Method and system for image categorization using a visual language model

    US20250094482A1

Cited By

  • Multi-agent system token consumption optimization method, system and equipment

    CN120951616A

  • A multi-agent system token consumption optimization method, system and device

    CN120951616B

  • Intelligent agent development system and method based on global variable and any hooking type

    CN121501258A