Operation automation system based on screen intelligent agent and implementation method
Through the operation automation system based on screen agents, using a large language model to process screen content and task instructions, the existing RPA system has solved the problems of high maintenance costs, long implementation cycles and limited applicable scenarios, and achieved lower maintenance costs, shorter implementation cycles and wider applicable scenarios.
Patent Information
- Application Number
- CN202510147016.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-06
AI Technical Summary
In the problems of high maintenance costs, long implementation cycles and limited applicable scenarios, existing RPA systems are difficult to effectively automate complex unstructured data and tasks that require judgment or reasoning.
Using a work automation system based on screen agents, through the websocket protocol communication between the client and the server, the client receives user task instructions and obtains screen shots, the server recognizes screen content and generates screen operation instructions through a large language model, and the client executes instructions to complete tasks.
Reduces system maintenance costs and implementation cycles, expands the applicability of the system in complex business scenarios, and enables automated tasks without major process changes.
Smart Images

Figure CN120104207A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more specifically to a job automation system based on screen intelligent body and an implementation method thereof. Background Art
[0002] In the daily operation of enterprises, a lot of repetitive and routine work is involved, and enterprises pay a lot of manpower costs in this regard. In order to cope with this situation, enterprises have gradually adopted RPA systems to realize the automation of work processes. However, the existing RPA systems have the following shortcomings in actual applications:
[0003] High maintenance cost: When the system or operation process changes, the RPA system needs to be reconfigured, which increases the maintenance labor cost;
[0004] Long implementation cycle: The screen content cannot be understood, and the system relies entirely on fixed rules configured manually. The process streamlining of the business system may lead to a long implementation cycle.
[0005] Limited applicable scenarios: RPA can only work in scenarios with clear rules and structured data. It has limited effectiveness for complex unstructured data and tasks that require judgment or reasoning.
[0006] Therefore, how to provide a job automation system based on screen intelligence and an implementation method is an issue that technical personnel in this field urgently need to solve. Summary of the invention
[0007] In view of this, the present invention provides a job automation system based on screen intelligence and an implementation method.
[0008] In order to achieve the above object, the present invention adopts the following technical solution:
[0009] A screen agent-based operation automation system includes: a client and a server, the client and the server communicate via a websocket protocol;
[0010] The client is used to receive user task instructions, obtain computer screen screenshots, and execute screen operation instructions on the computer screen;
[0011] The server is used to identify the screen content corresponding to the computer screen shot, input the screen content and the task instruction into the large language model, and output the screen operation instruction.
[0012] Preferably, the client comprises:
[0013] User instruction receiving unit: used to receive the user's task instruction and send it to the server;
[0014] Screen content capture unit: used to obtain computer screen screenshots and send them to the server;
[0015] Screen operation execution unit: used to perform corresponding operations on the computer screen after receiving the screen operation instruction from the server.
[0016] Preferably, the server includes:
[0017] Icon recognition unit: used to recognize the icon type on the computer screenshot and the specific position coordinates of the icon type on the screen through the target detection model;
[0018] A text recognition unit: used to recognize the text content on the computer screen shot and the specific position coordinates of the text on the screen through an OCR model;
[0019] Large language model unit: used to input the identified icon type, the specific position coordinates of the icon type on the screen, the text content, the specific position coordinates of the text on the screen and the task instruction into the large language model after being assembled with prompt words, and output the screen operation instruction.
[0020] Preferably, the task instruction is in text form.
[0021] Preferably, the target detection model is a yolo target detection model.
[0022] A method for realizing job automation based on a screen agent, comprising:
[0023] The client receives the user's task instruction and obtains a computer screen screenshot, and sends the task instruction and the computer screen screenshot to the server;
[0024] The server identifies the screen content corresponding to the computer screen shot, inputs the screen content and the task instruction into the large language model, outputs the screen operation instruction, and sends the screen operation instruction to the client;
[0025] The client executes the screen operation instruction on the computer screen.
[0026] Preferably, the client receives the user's task instruction and obtains a computer screen screenshot, and sends the task instruction and the computer screen screenshot to the server. The specific process is:
[0027] The user instruction receiving unit receives the user's task instruction and sends it to the server;
[0028] The screen content capture unit obtains a screenshot of the computer screen and sends it to the server;
[0029] The client executes the screen operation instruction on the computer screen, and the specific process is as follows:
[0030] After receiving the screen operation instruction from the server, the screen operation execution unit performs corresponding operations on the computer screen.
[0031] Preferably, the server identifies the screen content corresponding to the computer screen shot, inputs the screen content and the task instruction into the large language model, and outputs the screen operation instruction. The specific process is:
[0032] The icon recognition unit recognizes the icon type on the computer screen shot and the specific position coordinates of the icon type on the screen through the target detection model;
[0033] The text recognition unit recognizes the text content on the computer screen shot and the specific position coordinates of the text on the screen through an OCR model;
[0034] The large language model unit inputs the identified icon type, the specific position coordinates of the icon type on the screen, the text content, the specific position coordinates of the text on the screen and the task instruction into the large language model after assembling the prompt words, and outputs the screen operation instruction.
[0035] Preferably, the task instruction is in text form.
[0036] Preferably, the target detection model is a yolo target detection model.
[0037] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a screen-based intelligent agent-based operation automation system and implementation method, which has the following advantages:
[0038] 1) Low maintenance cost:
[0039] Through the icon recognition and text recognition units, the system can effectively understand the icons and text information on the screen and their corresponding coordinates. Without major process changes and revisions to the existing system, the system does not require new configuration, which greatly reduces the system's personnel maintenance costs.
[0040] 2) Short implementation period:
[0041] Since the system itself has a certain ability to understand screen information and user intentions, when configuring a job, it does not need to be accurate to the pixel level for each step like a traditional RPA system. It only needs to configure the general process, which greatly reduces the labor cost of system initialization configuration.
[0042] 3) Applicable to more business scenarios:
[0043] The large language model can infer the next operation based on the current information provided by the icon recognition and text recognition units and the original process configuration, rather than only being able to work under a completely fixed process. Therefore, compared with traditional RPA systems, it can work in more business scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0045] Figure 1 A schematic diagram of the structure of a screen-based intelligent operation automation system provided by the present invention.
[0046] Figure 2 A flow chart of a method for implementing job automation based on screen intelligence provided by the present invention. DETAILED DESCRIPTION
[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0048] The embodiment of the present invention discloses a screen agent-based job automation system, which can effectively understand screen information, coordinates and user intentions, and perform operations on a computer to complete user tasks. Figure 1 As shown, it includes a client and a server, and the client and the server communicate through the websocket protocol;
[0049] The client is used to receive the user's task instructions, obtain a computer screen screenshot, and execute screen operation instructions on the computer screen, wherein the computer screen screenshot is the computer screen screenshot corresponding to the moment when the user's instruction is received;
[0050] The server is used to identify the screen content corresponding to the computer screenshot, input the screen content and task instructions into the large language model, and output screen operation instructions.
[0051] In this embodiment, the client includes:
[0052] User instruction receiving unit: used to receive the user's task instruction and send it to the server. The task instruction is in text form;
[0053] Screen content capture unit: used to obtain computer screen screenshots and send them to the server;
[0054] Screen operation execution unit: used to perform corresponding operations on the computer screen after receiving screen operation instructions from the server.
[0055] In this embodiment, the server includes:
[0056] Icon recognition unit: used to identify the icon type on the computer screenshot and the specific position coordinates of the icon type on the screen through the target detection model; the target detection model adopts the yolo target detection model. The general yolo target detection model does not have the ability to recognize PC icons, especially icons in specific industry fields. The present invention collects computer screenshots in the target industry field, and marks the icons that need to be identified in the screenshots, and then trains the yolo model through the marked data. The trained yolo model can be used to realize icon recognition.
[0057] Currently, we only consider the processing of normal icons and zoom icons, and do not consider the case of icon rotation. Because the processing scenario is computer screen operation, icon rotation only occurs in very rare cases, so this scenario is not considered. The recognition of zoom icons is achieved through the collection and annotation of training data.
[0058] Text recognition unit: used to recognize the text content on the computer screen shot and the specific position coordinates of the text on the screen through the OCR model; the OCR model can use PaddleOCR. Because the text on the computer screen is basically clear print, there is no need to train the OCR model again.
[0059] Large language model unit: used to assemble the recognized icon type, the specific position coordinates of the icon type on the screen, the text content, the specific position coordinates of the text on the screen, and the task instructions through prompt words, input them into the large language model, and output screen operation instructions.
[0060] After comprehensively considering model performance, computing power requirements, and system costs, the large language model of qwen2.5-32B was adopted. In order to improve the decision-making accuracy of this language model in specific industry fields, it needs to be fine-tuned.
[0061] The fine-tuning data is collected as follows:
[0062] Collect the tasks that need to be performed on the computer screen of a specific industry (the tasks performed are screen operations, including mouse clicks and keyboard input) and their corresponding user task instructions;
[0063] The icon recognition results, text recognition results and user task instructions are assembled with prompt words and input into the large language model;
[0064] If the decision information output by the large language model is incorrect, such as the returned decision information cannot be converted into an operation instruction on the screen; or the screen operation fails to execute; or the user feedback: the user intention is not achieved after the screen operation, then the current icon recognition result, text recognition result, user intention and correct decision result are recorded as one of the large language model training data.
[0065] When the above operations are completed on all the task sets in the industry, the collected data is the training data for model fine-tuning. This data is used to fine-tune the selected large language model, and the obtained large language model can make accurate decisions.
[0066] Strategies for specific situations:
[0067] 1. User command ambiguity: Use a large-parameter large language model (such as qwen2.5-72B) or a closed-source large language model service to exhaustively generalize the user commands collected above. The generalized user commands, icon recognition results, text recognition results, and ideal outputs together constitute part of the model training data. The ideal output is the screen operation command that can meet the intention under the current computer screen display content.
[0068] 2. Multi-step operation tasks: Convert multi-step operation tasks into multi-round dialogues, and combine the training data of multi-step operation tasks in the form of multi-round dialogue fine-tuning data. After the large language model is trained with the above data, it can effectively handle multi-step task scenarios.
[0069] The present invention has the following properties:
[0070] 1) Compatibility processing: Because the intelligent agent of the present invention uses screen recognition to understand the target system and also uses screen operation to operate the target system, and does not involve API calls, there is no system compatibility issue.
[0071] 2) System performance:
[0072] System performance is related to user experience, product efficiency and operating costs, and is one of the decisive factors for product success or failure. The performance of the entire system is determined by the following points, and is also optimized for the following points:
[0073] Screen understanding part: Yolo model is used for icon recognition, and PaddleOCR tool is used for text recognition. When the Yolo model is running, it only needs less than 200MB of video memory and completes icon recognition in about 10ms; PaddleOCR tool also only needs about 500MB of video memory and completes text recognition in tens of milliseconds. Therefore, ordinary low-end graphics cards can support the efficient operation of the screen understanding part.
[0074] Large language model: The current model is the qwen2.5-32B large language model. After quantization, the model can be used for inference on graphics cards with less than 24G video memory. The inference result format is: [{"box_id":$box_id,"action_type":$action_type,"action_content":$action_content}], where box_id: When the screen is recognized for content, the recognition result will be presented in the form of many boxes, and box represents a rectangular bounding box. Each box has an id and coordinate information. action_type: Screen operation type, including: left mouse click, right mouse click, and input box text entry. action_content: Indicates the specific content corresponding to the screen operation type. When action_type is input box text entry, action_content is the specific information entered. The number of tokens that need to be output by the large language model is very small, and can usually be controlled within 200 tokens. On the 4090 graphics card, qwen2.5-32B can usually achieve an output speed of 100 tokens + / s. Therefore, for the large language model part, the recognition speed can be controlled within 2 seconds.
[0075] Comprehensive optimization: When the prediction accuracy of the intelligent agent is stable, the language part can be switched from qwen2.5-32B to qwen2.5-14B. This allows the system to run smoothly on a single 4090 or even 4080 GPU, and complete screen understanding and operation within 3 seconds.
[0076] The embodiment of the present invention discloses a method for realizing job automation based on screen intelligent agent, such as Figure 2 As shown, including:
[0077] The client receives the user's task instructions and obtains a computer screen screenshot, and sends the task instructions and computer screen screenshot to the server;
[0078] The server recognizes the screen content corresponding to the computer screenshot, inputs the screen content and task instructions into the large language model, outputs the screen operation instructions, and sends the screen operation instructions to the client;
[0079] The client executes the screen operation instructions on the computer screen.
[0080] In this embodiment, the client receives the user's task instruction and obtains a computer screen screenshot, and sends the task instruction and the computer screen screenshot to the server. The specific process is as follows:
[0081] The user instruction receiving unit receives the user's task instruction and sends it to the server, and the task instruction is in text form;
[0082] The screen content capture unit obtains a screenshot of the computer screen and sends it to the server;
[0083] The client executes screen operation instructions on the computer screen. The specific process is as follows:
[0084] After receiving the screen operation instruction from the server, the screen operation execution unit performs the corresponding operation on the computer screen.
[0085] In this embodiment, the server identifies the screen content corresponding to the computer screenshot, inputs the screen content and task instructions into the large language model, and outputs the screen operation instructions. The specific process is as follows:
[0086] The icon recognition unit identifies the icon type on the computer screen shot and the specific position coordinates of the icon type on the screen through the target detection model; the target detection model adopts the yolo target detection model. The general yolo target detection model does not have the ability to recognize PC icons, especially icons in specific industry fields. The present invention collects computer screenshots in the target industry field, and marks the icons that need to be identified in the screenshots, and then trains the yolo model through the marked data. The trained yolo model can be used to realize icon recognition.
[0087] Currently, we only consider the processing of normal icons and zoom icons, and do not consider the case of icon rotation. Because the processing scenario is computer screen operation, icon rotation only occurs in very rare cases, so this scenario is not considered. The recognition of zoom icons is achieved through the collection and annotation of training data.
[0088] The text recognition unit recognizes the text content on the computer screen shot and the specific position coordinates of the text on the screen through the OCR model; the OCR model can use PaddleOCR. Because the text on the computer screen is basically clear print, there is no need to train the OCR model again.
[0089] The large language model unit inputs the recognized icon type, the specific position coordinates of the icon type on the screen, the text content, the specific position coordinates of the text on the screen, and the task instruction into the large language model after assembling the prompt words, and outputs the screen operation instruction.
[0090] Other specific details are consistent with the system part and will not be repeated here.
[0091] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0092] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A screen agent-based operation automation system, characterized in that: include: Client and server, the client and server communicate through the websocket protocol; The client is used to receive user task instructions, obtain computer screen screenshots, and execute screen operation instructions on the computer screen; The server is used to identify the screen content corresponding to the computer screen shot, input the screen content and the task instruction into the large language model, and output the screen operation instruction.
2. The screen agent-based operation automation system according to claim 1, characterized in that: The client comprises: User instruction receiving unit: used to receive the user's task instruction and send it to the server; Screen content capture unit: used to obtain computer screen screenshots and send them to the server; Screen operation execution unit: used to perform corresponding operations on the computer screen after receiving the screen operation instruction from the server.
3. The screen agent-based operation automation system according to claim 1, characterized in that: The server includes: Icon recognition unit: used to recognize the icon type on the computer screenshot and the specific position coordinates of the icon type on the screen through the target detection model; A text recognition unit: used to recognize the text content on the computer screen shot and the specific position coordinates of the text on the screen through an OCR model; Large language model unit: used to input the identified icon type, the specific position coordinates of the icon type on the screen, the text content, the specific position coordinates of the text on the screen and the task instruction into the large language model after being assembled with prompt words, and output the screen operation instruction.
4. The screen agent-based operation automation system according to claim 2, characterized in that: The task instruction is in text form.
5. The screen agent-based operation automation system according to claim 3, characterized in that: The target detection model is a yolo target detection model.
6. A method for realizing job automation based on screen intelligent agent, characterized in that: include: The client receives the user's task instruction and obtains a computer screen screenshot, and sends the task instruction and the computer screen screenshot to the server; The server identifies the screen content corresponding to the computer screen shot, inputs the screen content and the task instruction into the large language model, outputs the screen operation instruction, and sends the screen operation instruction to the client; The client executes the screen operation instruction on the computer screen.
7. The method for realizing job automation based on screen agent according to claim 6, characterized in that: The client receives the user's task instruction and obtains a computer screen screenshot, and sends the task instruction and the computer screen screenshot to the server. The specific process is as follows: The user instruction receiving unit receives the user's task instruction and sends it to the server; The screen content capture unit obtains a screenshot of the computer screen and sends it to the server; The client executes the screen operation instruction on the computer screen, and the specific process is as follows: After receiving the screen operation instruction from the server, the screen operation execution unit performs corresponding operations on the computer screen.
8. The method for realizing job automation based on screen agent according to claim 6, characterized in that: The server identifies the screen content corresponding to the computer screenshot, inputs the screen content and the task instruction into the large language model, and outputs the screen operation instruction. The specific process is as follows: The icon recognition unit recognizes the icon type on the computer screen shot and the specific position coordinates of the icon type on the screen through the target detection model; The text recognition unit recognizes the text content on the computer screen shot and the specific position coordinates of the text on the screen through an OCR model; The large language model unit inputs the identified icon type, the specific position coordinates of the icon type on the screen, the text content, the specific position coordinates of the text on the screen and the task instruction into the large language model after assembling the prompt words, and outputs the screen operation instruction.
9. The method for realizing job automation based on screen agent according to claim 7, characterized in that: The task instruction is in text form.
10. The method for realizing job automation based on screen agent according to claim 8, characterized in that: The target detection model is a yolo target detection model.