Plug-and-play desktop interaction agent module and interaction method thereof
By parsing GUI views and voice commands through an embedded AI computing platform and a large language model, the flexibility and adaptability issues of the interactive intelligent agent are solved, plug-and-play desktop interaction is achieved, and the efficiency and robustness of human-computer interaction are improved.
Patent Information
- Application Number
- CN202510958693.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-10-17
AI Technical Summary
In the existing technology, interactive intelligent agents have poor flexibility and adaptability in understanding the intention of instructions, and their computing resources rely on the host computer, making it difficult to achieve plug-and-play and universal interaction across applications.
It uses an embedded AI computing platform, combined with a large language model and deep learning algorithms, to achieve autonomous parsing and execution of GUI views and voice commands through a USB interface and keyboard and mouse protocol, build structured text representations, and reduce dependence on host computing resources.
It achieves flexible parsing of complex instructions and cross-application adaptability, reduces dependence on the host computer's computing power, provides a plug-and-play convenient upgrade solution, and improves the efficiency and robustness of human-computer interaction.
Smart Images

Figure CN120803322A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of human-computer interaction, and particularly relates to a plug-and-play desktop interaction agent module and an interaction method thereof. BACKGROUND
[0002] The wide popularity of personal computers has made desktop software an indispensable tool in people's work and life. Such software provides a functionally rich graphical user interface, and users can interact through a keyboard and mouse. However, this interaction method has the problems of tedious operation and low efficiency. With the development of artificial intelligence technology, specific operations in the interaction process can be completed by an intelligent agent, which brings great convenience and efficiency improvement to human-computer interaction.
[0003] In the prior art, the intention understanding of the interaction intelligent agent for instructions is based on a specific recognition mode, key words and parameters in a sentence are extracted, and matching is performed in an intention template set to query an execution scheme. The flexibility is poor when facing complex and diverse instruction expression methods. Similarly, the operation of the desktop system GUI adopts a fixed behavior mode, and the operation flow needs to be constructed in advance based on the GUI content. This process requires human participation and has poor adaptability. In addition, the algorithms used by these technologies are usually directly run on a computer system, which needs to occupy the computing resources of the host, and is not conducive to the deployment of the intelligent agent.
[0004] Patent application No. CN116243826A discloses a GUI interface interaction method based on voice instructions, which includes: designing an instruction template according to the function area of the GUI interface; using intention understanding and slot recognition to analyze the instruction type and key parameters; matching the instruction template based on the instruction type and key parameters; and responding to the corresponding function area of the GUI according to the template. In this patent, the instructions are one-to-one mapped with the function areas of the GUI, which is difficult to extend to instructions with higher levels of abstraction. The function areas of the GUI need to be divided in advance, which lacks adaptability to different application programs. The algorithm is directly run on the computer system, which has strict requirements on the computing performance of the host computer. These problems make the technical solution of this patent application have certain limitations in actual deployment, and a general interaction intelligent agent has not been realized. SUMMARY
[0005] In order to overcome the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide a plug-and-play desktop interactive agent module and its interaction method, through the powerful semantic understanding ability of the large language model, the analysis of complex and diverse instruction expression mode is realized, without relying on fixed intent template matching, the flexibility of instruction understanding is significantly improved; by constructing a GUI view text representation with an HTML syntax structure, the GUI interface is converted into a structured text that can be processed by the large language model, avoiding pre-classifying functional areas, so that the system can adapt to the GUI layout of different application programs, realizing the general interaction ability across applications; using an embedded AI computing platform to run deep learning algorithms and large language models, the computing task is transferred from the desktop computer system to the agent module itself, completely eliminating the dependence on the host computing resources, and greatly reducing the deployment threshold; through prompt word engineering, the large language model is guided to perform complex potential logical reasoning, supporting automatic mapping from abstract instructions to specific GUI operations, expanding the processing capacity of the system for high-level instructions; a USB interface is designed to bridge the video capture card and the keyboard and mouse control protocol, realizing non-invasive connection with the computer system, without modifying the host system to complete the deployment, providing a plug-and-play convenient intelligent upgrading solution for the existing desktop environment.
[0006] Through the above technical innovation, the present application not only solves the problems of rigid instruction understanding, poor adaptability, and serious algorithm dependence in the prior art, but also realizes a truly general, flexible and easy-to-deploy desktop interactive agent, providing users with an efficient and natural human-computer interaction experience.
[0007] In order to achieve the above-mentioned purpose, the technical solution adopted by the present application is:
[0008] A plug-and-play desktop interactive agent module, comprising a plurality of USB interfaces and an embedded AI computing platform;
[0009] The embedded AI computing platform bridges the video capture card through the USB interface and then connects the HDMI display output interface of the computer, for capturing the GUI view;
[0010] The embedded AI computing platform connects the microphone through the USB, for acquiring the user's voice instruction and converting it into instruction text;
[0011] The embedded AI computing platform uses deep learning algorithms to analyze the GUI view, constructs a structured GUI view text representation, and inputs the instruction text and the GUI view text representation into the large language model to predict the instruction execution scheme, sends operation events to the computer system through the USB interface and the keyboard and mouse control protocol, and realizes the autonomous execution of user voice instructions and desktop interaction.
[0012] A desktop interaction method, comprising the following steps:
[0013] Step 1, the embedded AI computing platform acquires the GUI view through the HDMI display output interface of the computer via the USB interface, acquires the user's voice instruction through the microphone, and converts it into instruction text;
[0014] Step 2, the embedded AI computing platform uses a deep learning algorithm to analyze the GUI view collected in step 1, and constructs a structured GUI view text representation;
[0015] Step 3, the embedded AI computing platform inputs the instruction text obtained in step 1 and the GUI view text representation constructed in step 2 into a large language model, and predicts an instruction execution scheme represented by GUI control operation;
[0016] Step 4, the embedded AI computing platform sends operation events to the computer system based on the instruction execution scheme predicted in step 3, using the USB interface and the mouse control protocol, to realize autonomous execution of interaction.
[0017] In step 1, the embedded AI computing platform acquires the RGB image of the GUI view through the HDMI display output interface of the computer; the microphone uses a smart microphone that can convert voice to text, and the user's voice instruction is collected through the smart microphone and converted into instruction text; the RGB image and the instruction text are transmitted to the embedded AI computing platform through the USB interface.
[0018] In step 2, the RGB image of the GUI view is processed by a deep neural network model to construct a text representation with HTML syntax structure, which facilitates subsequent processing by a large language model.
[0019] In step 3, the large language model uses the text representation constructed in step 2 and the instruction text collected in step 1 to predict an execution scheme, with the operations in the execution scheme being in units of GUI controls to reduce the difficulty of prediction and achieve better results.
[0020] In step 4, according to the execution scheme predicted in step 3, a corresponding mouse operation event stream is constructed, and operation events are sent to the computer system through the USB interface and the mouse protocol to complete specific operations.
[0021] In steps 2 and 3, the deep learning algorithm and the large language model run on the embedded AI computing platform carried by the agent module.
[0022] Compared with the prior art, the present application has the following advantages:
[0023] 1. The present application uses artificial intelligence technology to analyze the interaction intent based on a large language model, realizes the execution flow prediction of complex instructions and GUI views, and improves the efficiency of human-computer interaction and the robustness of the agent.
[0024] 2.The method of the present application is based on the embedding of an intelligent agent with an embedded AI computing platform, which acquires GUI views and user instructions through an HDMI capture card and a microphone, and utilizes a USB interface and a keyboard / mouse protocol to achieve autonomous execution of interactions, with all components serving as external modules that can be plugged in and used immediately, thereby reducing the reliance on the computing power of a desktop computer system and having stronger practical application value.
[0025] In summary, the present application analyzes interactive intentions based on a large language model, realizes the prediction of the execution flow of complex instructions and GUI views, improves the efficiency of human-computer interaction and the robustness of intelligent agents, and utilizes a USB interface and a keyboard / mouse protocol to achieve autonomous execution of interactions, with all components serving as external modules that can be plugged in and used immediately, thereby reducing the reliance on the computing power of a desktop computer system and having stronger practical application value. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 The workflow diagram of the desktop interaction method of the present application.
[0027] Figure 2 The interactive hardware design schematic of the desktop interaction intelligent agent module of the present application.
[0028] Figure 3 One example of the GUI view text representation of the present application; wherein, Figure 3 (a) is a GUI view example, Figure 3 (b) is the text representation result of the GUI view example.
[0029] Figure 4 One example of the prompt word text of the present application.
[0030] Figure 5 (a) is an example before interaction that can be applied by the present application, Figure 5 (b) is the result after interaction of the example. DETAILED DESCRIPTION
[0031] The technical solutions of the present application will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0032] As shown in Figure 2 a plug-and-play desktop interaction intelligent agent module, comprising a plurality of USB interfaces and an embedded AI computing platform;
[0033] The embedded AI computing platform bridges a video capture card through a USB interface and then connects an HDMI display output interface of a computer, for collecting a GUI view;
[0034] The embedded AI computing platform connects a microphone through a USB, for obtaining a voice instruction of a user and converting the voice instruction into an instruction text;
[0035] The embedded AI computing platform analyzes the GUI view by using a deep learning algorithm, constructs a structured GUI view text representation, and inputs the instruction text and the GUI view text representation into a large language model to predict an instruction execution scheme, sends operation events to a computer system by using a USB interface and a keyboard and mouse control protocol, and realizes autonomous execution of desktop interaction of the user voice instruction.
[0036] As shown in FIG. Figure 1 A desktop interaction method includes the following steps:
[0037] Step 1: The embedded AI computing platform obtains a GUI view by connecting an HDMI display output interface of a computer through a USB interface, obtains a voice instruction of a user by using a microphone, and converts the voice instruction into an instruction text;
[0038] Step 2: The embedded AI computing platform analyzes the GUI view collected in step 1 by using a deep learning algorithm, and constructs a structured GUI view text representation;
[0039] Step 3: The embedded AI computing platform inputs the instruction text obtained in step 1 and the GUI view text representation constructed in step 2 into a large language model, and predicts an instruction execution scheme represented by a GUI control operation;
[0040] Step 4: The embedded AI computing platform sends operation events to a computer system by using a USB interface and a keyboard and mouse control protocol based on the instruction execution scheme predicted in step 3, and realizes autonomous execution of interaction.
[0041] In step 1, the embedded AI computing platform obtains the RGB image of the GUI view through the HDMI display output interface of the computer, and the video capture card also outputs the GUI view to the display for normal display. The microphone is an intelligent microphone that can convert voice into text, and the user's voice instruction is collected through the intelligent microphone and converted into instruction text. The RGB image and the instruction text are transmitted to the embedded AI computing platform through the USB interface. The embedded AI computing platform is built based on the high-performance SoC (System on Chip) of ARM architecture, integrates the NVIDIA Jetson series edge AI module, is equipped with a dedicated neural network processor (NPU) and a GPU acceleration unit, and is used for efficient operation of subsequent algorithm calculation. The intelligent microphone receives the user's voice instruction and converts it into instruction text, which is also sent to the embedded AI computing platform through the USB interface. The microphone does not need to be connected with the desktop computer.
[0042] In step 2, the RGB image of the GUI view is constructed into a text representation with HTML syntax structure through the Pix2Struct deep neural network model, as shown in the following formula: Figure 3 Pix2Struct is a cross-modal generation model based on the Transformer architecture, which realizes the direct conversion from GUI image to structured text by mapping visual features to text space. Each control in the GUI view, such as button, input box, text content, etc., has a corresponding HTML tag description in the text representation, and the tag attributes include control type, content value and coordinate information. These coordinates will be used to locate the operation position during the interaction execution in step 4. The present application adopts Pix2Struct model as the basis of end-to-end cross-modal text generation model, which uses Vision Transformer (ViT) as the visual encoder, combines with the autoregressive language model decoder, and is trained through large-scale multi-modal data to form a strong GUI description ability.
[0043] In step 3, through the prompt word engineering, the Chinese-LLaMA-Alpaca large language model fine-tuned by a large amount of Chinese corpus uses the GUI text representation constructed in step 2 and the instruction text collected in step 1 to predict the execution scheme. The operation in the execution scheme is in units of GUI controls to reduce the difficulty of prediction and achieve better results. The prompt word adopts Chain of Thought (CoT) prompt technology, as shown in the following formula: Figure 4Compared with directly giving the final answer, the thinking chain prompt word guides the model to think in logical steps, and the specific reasoning process is: (1) The model first analyzes the HTML structure in the GUI text representation, identifies all interactive elements and their position coordinates; (2) The keywords in the user instruction are semantically matched with the interface elements to locate the most relevant candidate elements; (3) Check if the candidate element type matches the type of the executable target instruction operation; (4) If the candidate element type does not match, the model searches for related controls in the adjacent area based on its position coordinates; (5) According to the finally determined operable element and its attributes, the corresponding operation is predicted. By explicitly adding the above reasoning steps in the example in the foregoing, the Chinese-LLaMA-Alpaca model can simulate this thinking process and output a more logical execution plan. The application further optimizes the prompt word template, and through multiple rounds of dialogue examples, it shows the reasoning path of different types of instructions, significantly improving the model's understanding ability of complex GUI interaction.
[0044] Step 4 constructs the corresponding keyboard and mouse operation event stream according to the execution plan predicted in step 3. Through the USB interface and the keyboard and mouse protocol, operation events are sent to the computer system to complete the specific operation. During execution, according to the operation type predicted in step 3 and the control type of the GUI text representation generated in step 2, mouse or keyboard operation events are selected for transmission. In addition, the control coordinates need to be determined to determine the operation position.
[0045] In steps 2 and 3, Figure 2 The artificial intelligence algorithm model runs on the embedded AI computing platform carried by the agent module to reduce the dependence on the computing power of the desktop computer system, which is conducive to the large-scale landing application of the interactive agent on the existing computer system.
[0046] The specific embodiments of the application will be further specifically described below by specific embodiments and in conjunction with the accompanying drawings:
[0047] As an example of an application scenario, specifically, Figure 2 As shown in the figure, the agent monitors the GUI view of the computer system in real time through the HDMI capture card, and transmits it to the embedded AI computing platform through the USB interface. When the user does not issue an instruction, the text representation of the GUI does not need to be constructed in order to improve the running efficiency of the system. In Figure 5 the example scenario shown in (a), the user issues the specific instruction "set the procurement order confirmation person to Lin Yuxia", which is processed by the intelligent microphone and transmitted to the AI computing platform in the form of text.
[0048] Referring to Figure 3 (a), after the AI computing platform obtains the instruction and the GUI view, it first constructs the text representation of the GUI, and the application adopts the method as Figure 3The HTML syntax shown in (b) in the figure describes the GUI view, and such syntax can greatly enhance the understandability of the GUI view; and then the upper input of the large language model is constructed in the style shown in Figure 4 The task description text and the example text of the prompt word part are fixed templates, and the GUI text representation to be predicted and the user instruction are added to the end of the prompt word to form the upper input; the large language model predicts the corresponding operation to be performed according to the upper input. In the present application, the large language model only predicts the operation to be performed at the next step each time, and when the instruction needs to be completed through multiple rounds of interaction, the algorithm iterates this process until the large model outputs the task end marker such as "{type="Over"}", which greatly improves the accuracy of the interaction.
[0049] Finally, the program parses the operation text predicted by the large language model, extracts the operation type, target control and parameters from the operation text, and selects the key mouse event stream to be sent according to these information. Specifically, in Figure 5 In the scenario shown in (a), the instruction "set the purchase order confirm person to Lin Yuxia" needs to fill in "Lin Yuxia" in the input box of "purchase order confirm person" in the GUI view. The agent first sends the mouse move and mouse click events to the desktop computer system to select the input box; then, the keyboard input events are sent in turn to input the three characters "Lin", "Yuxia" and "Xia". The interaction result of this example is shown in Figure 5 (b), in which the content of the "purchase order confirm person" input box is filled in with the three characters "Lin Yuxia".
[0050] It should be noted that the prediction mode and interaction mode of each operation itself have error correction function, when an operation fails, the agent obtains the wrong GUI view, and "Over" operation is not inferred at this time. Specifically, in Figure 5 the example scenario of (b), if the operation fails, that is, the input box is not correctly filled in with the three characters "Lin Yuxia", the agent will reattempt to perform the correct interaction until the "Over" operation is obtained.
[0051] It can be seen that, compared with the prior art, the present application predicts the interaction mode based on the large language model, without the need to divide the functional area according to different GUI views, and has a more powerful instruction-operation logical mapping capability, rather than a simple one-to-one mapping, and has good adaptability and flexibility; in addition, the interaction agent used in the present application is independent of the desktop computer system, and only needs to be connected to the computer through HDMI and USB interfaces to work, which reduces the requirement for the computing performance of the host computer and is easy to deploy in practice. Through the present application, the human-computer interaction process does not need the frequent participation of the user, and the interaction efficiency is greatly improved.
[0052] It will be obvious to a person skilled in the art that the application is not limited to the details of the foregoing exemplary embodiments and can be implemented in other concrete forms without departing from the spirit or essential characteristics of the application. The embodiments are therefore to be considered in all respects as illustrative and not restrictive, the scope of the application being indicated by the appended claims rather than by the foregoing description, and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein. No reference signs in the claims should be considered as limiting the scope of the claims to the identity of the reference signs therein.
[0053] Furthermore, it should be understood that although the description is made on the basis of the embodiments, not every embodiment contains only one independent technical solution, and the description of the specification is only for the sake of clarity, and those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that those skilled in the art can understand.
Claims
1. A plug-and-play desktop interactive agent module, comprising multiple USB interfaces, characterized in that: It also includes an embedded AI computing platform; The embedded AI computing platform is connected to the video capture card via a USB interface and then to the HDMI display output interface of the computer to capture the GUI view; The embedded AI computing platform is connected to a microphone via USB to obtain the user's voice commands and convert them into command text; The embedded AI computing platform uses a deep learning algorithm to parse the GUI view, construct a structured GUI view text representation, and input the instruction text and GUI view text representation into a large language model to predict the instruction execution plan. Through the USB interface and keyboard and mouse control protocol, it sends operation events to the computer system to realize the autonomous execution of user voice commands and desktop interaction.
2. A desktop interaction method, characterized in that: The steps include: Step 1: The embedded AI computing platform connects to the computer's HDMI display output interface via a USB interface to obtain the GUI view, receives the user's voice commands through the microphone, and converts them into command text; Step 2: The embedded AI computing platform uses a deep learning algorithm to parse the GUI view collected in step 1 and construct a structured GUI view text representation; Step 3: The embedded AI computing platform inputs the instruction text obtained in step 1 and the GUI view text representation constructed in step 2 into the large language model to predict the instruction execution plan represented by the GUI control operation; In step 4, the embedded AI computing platform uses the USB interface and keyboard and mouse control protocol to send operation events to the computer system based on the instruction execution plan predicted in step 3, thereby realizing autonomous execution of the interaction.
3. The desktop interaction method according to claim 2, wherein: In step 1, the embedded AI computing platform obtains the RGB image of the GUI view through the HDMI display output interface of the computer; the microphone uses an intelligent microphone that can convert voice into text, and the user's voice commands are collected by the intelligent microphone and converted into command text; the RGB image and command text are transmitted to the embedded AI computing platform through the USB interface.
4. The desktop interaction method according to claim 2, wherein: In step 2, the RGB image of the GUI view is used to construct a text representation with HTML grammatical structure through a deep neural network model, which is convenient for subsequent large language model processing.
5. The desktop interaction method according to claim 2, characterized in that: In step 3, the large language model uses the text representation constructed in step 2 and the instruction text collected in step 1 to predict the execution plan through the prompt word engineering. The operations in the execution plan are based on GUI controls to reduce the prediction difficulty and achieve better results.
6. The desktop interaction method according to claim 2, characterized in that: In step 4, according to the execution plan predicted in step 3, a corresponding keyboard and mouse operation event stream is constructed, and the operation event is sent to the computer system through the USB interface and the keyboard and mouse protocol to complete the specific operation.
7. The desktop interaction method according to claim 2, wherein: In steps 2 and 3, the deep learning algorithm and the large language model are run on the embedded AI computing platform equipped with the intelligent agent module as described in claim 1.
Citation Information
Patent Citations
UI interface design and man-machine interaction method based on voice instruction
CN116243826A