Closed-loop reasoning method, device and equipment based on multi-modal large model
By using closed-loop reasoning of a multimodal large language model, combined with screen visual information and target tasks, atomic operation instructions are generated, solving the problem of interaction difficulties in complex tasks of existing AI systems and achieving efficient control at the operating system level.
Patent Information
- Application Number
- CN202511512531.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2025-11-21
AI Technical Summary
Existing AI systems are unable to understand complex intentions when handling complex tasks, cannot manipulate any visible elements on the interface, have a clunky interactive experience, are difficult to expand functionality, and cannot handle undefined tasks.
A multimodal large language model is used for closed-loop reasoning. By acquiring screen visual information and target task, and combining operation memory and multimodal large language model, a single tool call command is generated to execute atomic operations until the target task is completed.
It reduces the control issues of numerous upper-layer applications to atomic operation combination calls at the operating system level, improving the flexibility and accuracy of task execution, adapting to complex interface changes, and increasing task execution efficiency.
Smart Images

Figure CN120996209A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence, and in particular to a closed-loop reasoning method, apparatus, and device based on a multimodal large model. Background Technology
[0002] MLLM (Multimodal Large Language Model), especially LVLM (Large Vision Language Model), achieves a deeper understanding of the real world by integrating visual and linguistic modalities. Using MLLM as the "brain" and combining it with external tools for autonomous planning and execution, an intelligent agent architecture provides a new paradigm for solving complex tasks that require interaction with the environment.
[0003] However, existing technologies attempt to enable AI (Artificial Intelligence) to "understand" specific functions of upper-layer applications. Due to the strict limitation of interaction logic, they cannot handle any undefined tasks, understand complex intentions, or manipulate any visible elements on the interface. The interaction experience is stiff, functional expansion is difficult, and the processing is mechanical. Summary of the Invention
[0004] This disclosure provides a method, apparatus, and device for closed-loop inference based on a multimodal large model, which solves the technical problems of existing AI processing of mechanical data.
[0005] According to a first aspect of this disclosure, a closed-loop reasoning method based on a multimodal large-scale model is provided. The method includes: acquiring screen visual information and a target task; performing a search operation in a pre-generated or real-time generated operation memory based on the target task; if a corresponding historical task operation sequence is found, executing the operation according to the historical task operation sequence; if no corresponding historical task operation sequence is found, inputting the screen visual information and the target task instruction into a multimodal large-scale language model to obtain a single tool invocation instruction, the single tool invocation instruction containing operation information on screen coordinates or interface elements; executing the task corresponding to the type of the single tool invocation instruction based on the screen visual information, updating the screen visual information and the single tool invocation instruction, until the execution result is consistent with the execution result corresponding to the target task, thereby completing the closed-loop reasoning.
[0006] In addition to the above aspects and any possible implementation, a further implementation is provided, in which screen visual information and target task instructions are input into a multimodal large language model to obtain a single tool invocation instruction, including: if the screen visual information is initial screen visual information, fusing the initial screen visual information and the target task to obtain first fused information, and inputting the first fused information into the multimodal large language model to obtain a first tool invocation instruction; wherein, the first tool invocation instruction includes initial screen coordinates or operation information of initial interface elements.
[0007] In some possible implementations, the closed-loop inference process includes: according to the type of the first tool call instruction, executing the first task corresponding to the type of the first tool call instruction based on the initial screen visual information to obtain a first execution result; if the first execution result is that the first task is not completed, updating the screen visual information to obtain second screen visual information; combining the second screen visual information, the target task, and the first historical record to obtain second fusion information, and inputting the second fusion information into the multimodal large language model to obtain a second tool call instruction; according to the type of the second tool call instruction, executing the second task corresponding to the type of the second tool call instruction based on the second screen visual information to obtain a second execution result; if the second execution result is that the second task is not completed, updating the screen visual information to obtain third screen visual information; combining the third screen visual information, the target task, and the second historical record to obtain third fusion information, and inputting the third fusion information into the multimodal large language model to obtain a third tool call instruction; according to the type of the third tool call instruction, executing the third task corresponding to the type of the third tool call instruction based on the third screen visual information until the final execution result is consistent with the execution result corresponding to the target task.
[0008] In some possible implementations, the generation of the operation memory includes: recording the order of task operations in the closed-loop reasoning process, which is denoted as a task operation sequence; and recording the number of times the task operation sequence is executed, which yields the number of times the task operation sequence is executed; if the number of times the task operation sequence is executed is greater than a preset number of times, then the task operation sequence that is greater than the preset number of times is denoted as a historical task operation sequence; and storing the historical task operation sequence in an operation memory that is created in advance or in real time.
[0009] In some possible implementations, closed-loop reasoning also includes: identifying user intent based on the operation information of interface elements to obtain user intent identification results; initiating an execution query based on the user intent identification results; capturing screen visual information corresponding to the execution query results to obtain fourth screen visual information; inputting the fourth screen visual information into a multimodal large language model to obtain a fourth tool invocation instruction; and executing the fourth task corresponding to the type of the fourth tool invocation instruction based on the fourth screen visual information, until the final execution result is consistent with the execution result corresponding to the target task.
[0010] In some possible implementations, the underlying control tools corresponding to a single tool call command include atomic interface operation tools, which at least include mouse click tools, keyboard input tools, and screen scrolling tools.
[0011] In some possible implementations, screen visual information includes at least one of the following: a complete screenshot of the client screen, a web page document object model tree structure, a native application's user control accessibility object model, or a combination of the above.
[0012] According to a second aspect of this disclosure, a closed-loop inference device based on a multimodal large model is provided. The device includes: an acquisition module for acquiring screen visual information and a target task; The lookup module is used to perform a lookup operation in the pre-generated or real-time operation memory based on the target task. If a corresponding historical task operation sequence is found, it will be executed according to the historical task operation sequence. The reasoning module is used to input screen visual information and target task instructions into the multimodal large language model if no corresponding historical task operation sequence is found, to obtain a single tool call instruction, which contains operation information on screen coordinates or interface elements; and to execute the task corresponding to the type of single tool call instruction based on the screen visual information, according to the type of single tool call instruction, and update the screen visual information and single tool call instruction until the execution result is consistent with the execution result corresponding to the target task, so as to complete the closed-loop reasoning.
[0013] According to a third aspect of this disclosure, an electronic device is provided. The electronic device includes a memory and a processor, wherein a computer program is stored in the memory, and the processor executes the program to implement the method described above.
[0014] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the methods according to the first and / or second aspects of this disclosure.
[0015] The beneficial effects of this application are as follows: By inputting screen visual information and target task instructions into a multimodal large language model, a single tool invocation instruction is obtained. This single tool invocation instruction contains operation information on screen coordinates or interface elements. The problem of controlling numerous upper-layer applications by an intelligent agent is reduced to a problem of combining and invoking a few atomic operations at the operating system level. Then, according to the type of the single tool invocation instruction, the task corresponding to that type is executed based on the screen visual information. The screen visual information and the single tool invocation instruction are updated until the execution result matches the execution result corresponding to the target task, thus completing closed-loop reasoning. This employs a closed-loop feedback mechanism, where each decision is based on real and up-to-date visual feedback, and dynamic adjustments and corrections are made to ensure that the processing result better meets user needs.
[0016] Furthermore, the system performs a search operation in a pre-generated or real-time operation memory based on the target task. If a corresponding historical task operation sequence is found, it is executed according to the historical task operation sequence, thereby recording high-frequency operation sequences and improving task execution efficiency.
[0017] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0018] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the scope of this disclosure. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 An exemplary system architecture diagram is shown in which embodiments of the present disclosure can be implemented; Figure 2 An exemplary alternative system architecture diagram is shown in which embodiments of the present disclosure can be implemented; Figure 3 A flowchart of a closed-loop inference method based on a multimodal large model according to an embodiment of the present disclosure is shown; Figure 4 A block diagram of a closed-loop inference device based on a multimodal large model according to an embodiment of the present disclosure is shown; Figure 5 A block diagram of an exemplary electronic device capable of implementing embodiments of the present disclosure is shown. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0020] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0021] In the evolution of large language models to multimodal large language models and intelligent agent architectures in the field of artificial intelligence, MLLM, especially large visual language models, has achieved a deeper understanding of the real world by integrating visual and language modalities. Using MLLM as the "brain" and combining it with external tools for autonomous planning and execution, an intelligent agent architecture provides a new paradigm for solving complex tasks that require interaction with the environment. However, how to enable this "brain" to interact efficiently and universally with the massive, non-cooperative GUIs (Graphical User Interfaces) we use daily is the core challenge currently facing the practical application of this technology.
[0022] Existing technologies are based on traditional interactive systems that rely on intent recognition and slot filling, as well as early tool-enhanced large language models that rely on API (Application Programming Interface) calls.
[0023] 1. Intent Recognition and Slot Filling Scheme: This is the core technology of current mainstream intelligent assistants. Developers need to predefine all the intents (such as PlayVideo) and slots (such as videoName) that the system can support, and train the model for matching. Its inherent drawback is that the interaction logic is strictly limited, it cannot handle any undefined tasks, cannot understand complex intents, and cannot manipulate any visible elements on the interface, resulting in a rigid interaction experience and difficulty in expanding functionality.
[0024] 2. API-based Agent Approach: This approach teaches the LLM to call a series of pre-packaged software APIs (e.g., calling the `search_flight(destination, date)` function) to perform tasks. It is more flexible than the previous approach, but its capabilities are limited by the quantity and quality of available APIs. It still cannot operate applications without APIs (e.g., a local calculator, an outdated desktop application), nor can it handle fine-grained operations that rely on visual positioning and are difficult to describe via APIs (e.g., "click the red close button in the upper right corner").
[0025] The common flaw in the aforementioned solutions is that they all attempt to enable AI to "understand" specific functions of upper-layer applications, rather than fundamentally endowing AI with the ability to "use" general computer interfaces. This invention proposes a novel technical solution. Its core idea is to use a multimodal large language model as the intelligent decision-making center, and the ability to simulate low-level input devices such as a mouse and keyboard as its universal "hand," driven by a continuous "observation-thinking-action" closed-loop process. MLLM does not plan the entire task path all at once, but rather, like a human, "takes one step at a time," deciding only the next atomic operation that should be executed based on the visual understanding of the current screen at each point in time, thereby gradually completing the user's final task.
[0026] The following describes an exemplary closed-loop inference system architecture based on a multimodal large model, as illustrated in the embodiments of this application. Figure 1 , Figure 1 An exemplary system architecture diagram implementing embodiments of the present disclosure is shown, including a client device 110, a voice module 120, and a multimodal reasoning and decision-making module 130. The client device 110 carries the application for the user interface and core tool scheduler, used to capture screen visual information and execute atomic operation instructions issued by MLLM. The voice module 120 is used for speech-to-text and text-to-speech, providing users with a natural voice interaction entry point. The multimodal reasoning and decision-making module 130 deploys a multimodal large language model, serving as the "brain" of the system, used to receive visual information and task objectives uploaded by the client, perform closed-loop reasoning, and decide on the next action instruction.
[0027] The following describes an exemplary architecture of a closed-loop inference system 200 based on a multimodal large model, as illustrated in an embodiment of this application. Figure 2The system 200 includes a client device 201 and an inference and decision-making module 202. The client device 201 includes a visual information acquisition unit 2010 and a tool scheduling unit 2020. The inference and decision-making module 202 inputs screen visual information and target task instructions into a multimodal large language model to obtain a single tool invocation instruction. The visual information acquisition unit 2010 cyclically captures screen visual information and sends it to the inference and decision-making unit. The tool scheduling unit 2020 receives the single tool invocation instruction from the inference and decision-making unit and invokes the corresponding underlying control tool to execute the instruction.
[0028] The client device can be a computer, mobile phone, or other similar device. It can also be AR / VR glasses, a smart cockpit, an industrial control panel, or a robot equipped with a camera. Its core technical features remain unchanged: it must internally include a visual context information acquisition module and a tool scheduler for performing atomic operations.
[0029] As can be seen from the above architecture diagram, the execution flow uses the multimodal large language model as the intelligent decision-making center, and the capabilities of simulating low-level input devices such as mice and keyboards as its general "hands," driven by a continuous "observation-thinking-action" closed-loop process. The multimodal large language model does not plan the entire task path all at once, but rather, at each point in time, based on its visual understanding of the current screen, it only decides on the next atomic operation that should be executed, thereby gradually completing the user's final task.
[0030] The following describes a closed-loop inference method based on a multimodal large model according to embodiments of this disclosure. Figure 3 A flowchart of a method 300 for closed-loop inference based on a multimodal large model according to an embodiment of the present disclosure is shown. Method 300 can be performed in... Figure 1 or Figure 2 The system architecture diagram based on a multimodal large model is executed.
[0031] Step S310: Obtain screen visual information and target task.
[0032] In some embodiments, screen visual information includes at least one of the following: a complete screenshot of the client screen, a web page document object model tree structure, a native application's user control accessibility object model, or a combination of the above.
[0033] As an example, the target task can fuse the user's original final task objective, such as "Check the weather and open related news," with the visual context and possible historical interaction records (such as previous conversations and operations) to obtain fused multimodal information. Specifically, the target speech can be acquired through a speech module, converted into text, and the target task can be identified through text recognition.
[0034] As an example, the screen visual information and the target task are obtained, and the screen visual information and the target task are fused. The client captures a complete screenshot of the current screen as the current visual context.
[0035] As an example, the fused multimodal information is input into the multimodal inference decision module. Specifically, this involves guiding the MLLM through inference using a carefully designed prompt template. This template indicates that the MLLM's role is that of a general computer operator and describes the available set of atomic tools (such as click(coordinates), type_text(text), scroll(direction), etc.). The MLLM's task is not to plan all steps, but rather to determine and generate only the most reasonable atomic operation based on the current state. The output is a single, structured tool invocation instruction in JSON format, such as {"tool_name": "click", "parameters": {"x": 850, "y": 600}}. Furthermore, the model can also decide whether the task is completed or not, such as {"tool_name": "task_complete", "parameters": {"reason": "News has been opened"}}.
[0036] In addition to capturing screen visual information, clients can also obtain visual context in more efficient ways. For example, for web pages, a simplified DOM (Document Object Model) tree structure can be extracted and sent; for native applications, the accessibility object model of their UI controls can be extracted. This structured data can be combined with image data to provide richer positioning information for MLLM.
[0037] Step S320: Search the operation memory in the pre-generated or real-time generated operation memory according to the target task. If the corresponding historical task operation sequence is found, execute it according to the historical task operation sequence.
[0038] During task execution, the operation sequence of the same task is usually similar. If tasks with similar operation sequences can be memorized and stored, task execution efficiency can be improved while ensuring execution accuracy.
[0039] Therefore, in some embodiments, the generation of the operation memory includes: recording the order of task operations in the closed-loop reasoning process, denoted as a task operation sequence; and recording the number of times the task operation sequence is executed, to obtain the number of times the task operation sequence is executed; if the number of times the task operation sequence is executed is greater than a preset number of executions, then the task operation sequence that is greater than the preset number of executions is recorded as a historical task operation sequence; and the historical task operation sequence is stored in an operation memory created in advance or in real time.
[0040] In step S330, if no corresponding historical task operation sequence is found, the screen visual information and the target task instruction are input into the multimodal large language model to obtain a single tool invocation instruction. The single tool invocation instruction contains operation information on screen coordinates or interface elements.
[0041] Step S340: According to the type of single tool call instruction, execute the task corresponding to the type of single tool call instruction based on the screen visual information, update the screen visual information and the single tool call instruction, until the execution result is consistent with the execution result corresponding to the target task, so as to complete the closed-loop reasoning.
[0042] In some embodiments, the underlying control tool corresponding to a single tool call instruction includes atomic interface operation tools, which include at least a mouse click tool, a keyboard input tool, and a screen scrolling tool.
[0043] In some embodiments, the type of a single tool invocation instruction includes multiple operations such as "UI control" (e.g., simulating mouse and keyboard) or "knowledge query" (e.g., calling a database). Specifically, when the server returns a "single tool invocation instruction," the process enters the "action" phase. The client's tool scheduler parses the instruction and determines and distributes it according to its type. The system can execute multiple operations such as "UI control" (e.g., simulating mouse and keyboard) or "knowledge query" (e.g., calling a database), demonstrating its scalability. After executing the operation, immediate user feedback (e.g., voice announcement) is provided, and the system waits for the interface status to update.
[0044] By invoking commands through a single tool, the problem of controlling countless upper-layer applications by an intelligent agent is reduced to a problem of combining a few atomic operations (click, input, scroll) at the operating system level. The model does not need to learn any application's dedicated API; instead, it learns a general "screen-based, keyboard-and-mouse-based" capability. This gives it the potential to control any visible and controllable graphical interface, fundamentally solving the problem of controlling third-party applications and applications without APIs.
[0045] In some embodiments, inputting screen visual information and target task instructions into a multimodal large language model to obtain a single tool invocation instruction includes: if the screen visual information is initial screen visual information, fusing the initial screen visual information and the target task to obtain first fused information, and inputting the first fused information into the multimodal large language model to obtain a first tool invocation instruction; wherein the first tool invocation instruction includes initial screen coordinates or operation information of initial interface elements.
[0046] In some embodiments, the closed-loop inference process includes: executing a first task corresponding to the type of the first tool invocation instruction based on initial screen visual information according to the type of the first tool invocation instruction, and obtaining a first execution result; if the first execution result is that the first task is not completed, updating the screen visual information to obtain second screen visual information; combining the second screen visual information, the target task, and the first historical record to obtain second fusion information, and inputting the second fusion information into a multimodal large language model to obtain a second tool invocation instruction; executing a second task corresponding to the type of the second tool invocation instruction based on the second screen visual information according to the type of the second tool invocation instruction, and obtaining a second execution result; if the second execution result is that the second task is not completed, updating the screen visual information to obtain third screen visual information; fusing the third screen visual information, the target task, and the second historical record to obtain third fusion information, and inputting the third fusion information into a multimodal large language model to obtain a third tool invocation instruction; executing a third task corresponding to the type of the third tool invocation instruction based on the third screen visual information according to the type of the third tool invocation instruction, until the final execution result is consistent with the execution result corresponding to the target task. As can be seen, the problem of controlling countless upper-layer applications by an intelligent agent is reduced to a problem of combining and calling a few atomic operations (click, input, scroll) at the operating system level. The model does not need to learn any application's dedicated API; it learns a general "screen-based, keyboard-and-mouse-based" capability. This gives it the potential to control any visible and controllable graphical interface, fundamentally solving the problem of controlling third-party applications and applications without APIs.
[0047] As can be seen, the closed-loop feedback mechanism bases each decision on real and up-to-date visual feedback. When unexpected pop-ups, loading delays, or interface changes occur during execution, the model can "see" these changes in the next loop and dynamically adjust and correct them, much like a human operator, greatly improving the success rate of complex tasks. As an example, recognizing user instructions, the target task is "Book me an economy class ticket to Shanghai tomorrow, choosing a window seat; remind me if the price exceeds 1000 yuan." The closed-loop process includes: Round 1: The client captures initial screen visual information: for example, the initial screen visual information is the homepage of an airline's official website (without entering the booking page), and there is no historical operation record. The initial screen visual information and the target task are fused to obtain the first fused information, namely the user's final task goal ("Book tomorrow's Shanghai economy class, window seat, reminder for over 1000 yuan") + the current screen screenshot (homepage interface) + empty historical records, forming a multimodal context. The multimodal big model receives the above context and performs inference through a preset prompt template. The "Flight Booking" entry button on the homepage (coordinates x=300, y=200) determines that "entering the booking page" is the most necessary first step. The client tool scheduler parses the instruction, simulates a mouse click at coordinates (300, 200), and triggers a page jump to the "Flight Booking Page". After the page jumps, the system waits 1 second to confirm the interface refresh, determines that the task is not completed (departure location, date, etc. are not entered), and enters the next round of the loop.
[0048] Round 2: The client captures visual information from the second screen: a flight booking page containing input boxes for "Departure City," "Arrival City," and "Departure Date" (currently empty). The second screen visual information is fused with the target task to obtain the second fused information: Target Task + Second Screen Visual Information (Booking Page) + First Historical Record (Step 1: Clicked the "Flight Booking" button). The model analyzes the current interface: "Departure City" must be entered first; the input box coordinates are x=150, y=300. The inference conclusion is: "The next step requires activating the departure city input box before the departure location can be entered." The client simulates a click, and the "Departure City" input box is activated. If the task is not completed (departure location not entered), proceed to the next round.
[0049] Round 3: The client captures visual information from the third screen, which shows that the "Departure City" input box is active, while other input boxes remain empty. The resulting fusion information includes: target task + third-screen visual information + second historical record (clicked booking entry, activated departure city box). Model analysis: The input box is active, requiring the departure city to be entered first (assuming the user departs from the current city, assuming "Beijing"). The client simulates keyboard input for "Beijing," the input box displays "Beijing," and the system automatically suggests matching "Beijing Daxing Airport" and "Beijing Capital Airport." If the task is not completed (no departure city entered), the next round of the loop begins; if the task is completed, the loop ends.
[0050] Fourth loop: The client captures visual information from the fourth screen. This visual information is as follows: after entering "Beijing," a pop-up ad ("Member Registration Discount," including a "Close" button, coordinates x=500, y=400) suddenly appears, obscuring the "Arrival City" input box. The above information is then merged. The fourth set of fused information is obtained: target task + fourth screen visual information (including pop-up window) + history (departure city "Beijing" has been entered). The multimodal large language model recognizes that the pop-up window is obscuring the operation target (arrival city input box) and determines that "the pop-up window must be closed before continuing". The atomic tool is invoked to execute the operation; the client simulates a click, the pop-up window closes, and the screen returns to the booking page (the "arrival city" input box is visible). The system checks whether the task is completed; if not, it enters the next loop until the task execution is complete.
[0051] As can be seen, the aforementioned multimodal context is transmitted to the multimodal reasoning and decision-making service module, which represents the system's "brain." The core task of this module is to perform single-step reasoning: instead of planning all steps at once, it decides the most appropriate atomic operation to execute next based solely on the current input. This "step-by-step" strategy gives the system extremely high flexibility and adaptability.
[0052] During task execution, the mouse movement trajectory and input speed may reflect the user's operational intent. Therefore, the user's intent can be identified based on the mouse movement trajectory and input speed, thereby proactively initiating clarification inquiries to the user and improving the naturalness and accuracy of the interaction.
[0053] In some embodiments, closed-loop reasoning further includes: identifying user intent based on operation information of interface elements to obtain user intent identification results; initiating an execution query based on the user intent identification results; capturing screen visual information corresponding to the execution query results to obtain fourth screen visual information; inputting the fourth screen visual information into a multimodal large language model to obtain a fourth tool invocation instruction; and executing a fourth task corresponding to the type of the fourth tool invocation instruction based on the fourth screen visual information, until the final execution result is consistent with the execution result corresponding to the target task. By introducing a user intent identification and confirmation mechanism, it is possible to handle ambiguous situations or situations requiring further user confirmation, thereby enhancing the flexibility and adaptability of the system.
[0054] Therefore, this application adopts a closed-loop feedback mechanism, where each decision is based on real and up-to-date visual feedback. When unexpected pop-ups, loading delays, or interface changes are encountered during execution, the model can "see" this change in the next loop and make dynamic adjustments and corrections, just like a human operator, greatly improving the success rate of complex tasks.
[0055] According to embodiments of this disclosure, the following technical effects are achieved: By inputting screen visual information and target task instructions into a multimodal large language model, a single tool invocation instruction is obtained, which contains operation information on screen coordinates or interface elements; the problem of controlling numerous upper-layer applications by an intelligent agent is reduced to a problem of combining and invoking a few atomic operations at the operating system level. Then, according to the type of the single tool invocation instruction, the task corresponding to the type of the single tool invocation instruction is executed based on the screen visual information, and the screen visual information and the single tool invocation instruction are updated until the execution result is consistent with the execution result corresponding to the target task, thus completing closed-loop reasoning. That is, a closed-loop feedback mechanism is adopted, where each decision is based on real and up-to-date visual feedback, and dynamic adjustments and corrections are made accordingly, making the processing result more in line with user needs.
[0056] Furthermore, the system performs a search operation in a pre-generated or real-time operation memory based on the target task. If a corresponding historical task operation sequence is found, it is executed according to the historical task operation sequence, thereby recording high-frequency operation sequences and improving task execution efficiency.
[0057] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this disclosure.
[0058] The above is an introduction to the method embodiments. The following describes the present disclosure further through device embodiments.
[0059] Figure 4 A block diagram of a closed-loop inference device 400 based on a multimodal large model according to an embodiment of the present disclosure is shown. Figure 4As shown, the closed-loop inference device 400 based on a multimodal large model includes: an acquisition module 410, a search module 420, and an inference module 430. The acquisition module acquires screen visual information and the target task; the search module searches in a pre-generated or real-time operation memory based on the target task, and if a corresponding historical task operation sequence is found, it executes according to that sequence; the inference module, if no corresponding historical task operation sequence is found, inputs the screen visual information and the target task instruction into the multimodal large language model to obtain a single tool call instruction, which includes operation information on screen coordinates or interface elements; and, according to the type of the single tool call instruction, executes the task corresponding to that type based on the screen visual information, updating the screen visual information and the single tool call instruction until the execution result matches the execution result corresponding to the target task, thus completing the closed-loop inference. By searching in the pre-generated or real-time operation memory based on the target task, and executing according to the historical task operation sequence if a corresponding historical task operation sequence is found, high-frequency operation sequences are recorded, thereby improving task execution efficiency. By inputting screen visual information and target task instructions into a multimodal large language model, a single tool invocation instruction is obtained. This single tool invocation instruction contains operation information on screen coordinates or interface elements. The problem of controlling numerous upper-layer applications by an intelligent agent is reduced to a problem of combining and invoking a few atomic operations at the operating system level. Then, according to the type of single tool invocation instruction, the task corresponding to that type is executed based on the screen visual information. The screen visual information and the single tool invocation instruction are updated until the execution result matches the execution result corresponding to the target task, thus completing closed-loop reasoning. This employs a closed-loop feedback mechanism, where each decision is based on real and up-to-date visual feedback, dynamically adjusted and corrected to ensure the processing result better meets user needs.
[0060] In some embodiments, the inference module is configured to, if the screen visual information is initial screen visual information, fuse the initial screen visual information and the target task to obtain first fused information, and input the first fused information into a multimodal large language model to obtain a first tool invocation instruction; wherein, the first tool invocation instruction includes initial screen coordinates or operation information of initial interface elements.
[0061] In some embodiments, the inference module is configured to: execute a first task corresponding to the type of the first tool invocation instruction based on the initial screen visual information, according to the type of the first tool invocation instruction, to obtain a first execution result; if the first execution result is that the first task is not completed, update the screen visual information to obtain second screen visual information; combine the second screen visual information, the target task, and the first historical record to obtain second fusion information, and input the second fusion information into a multimodal large language model to obtain a second tool invocation instruction; execute a second task corresponding to the type of the second tool invocation instruction based on the second screen visual information, according to the type of the second tool invocation instruction, to obtain a second execution result; if the second execution result is that the second task is not completed, update the screen visual information to obtain third screen visual information; fuse the third screen visual information, the target task, and the second historical record to obtain third fusion information, and input the third fusion information into a multimodal large language model to obtain a third tool invocation instruction; execute a third task corresponding to the type of the third tool invocation instruction based on the third screen visual information, according to the type of the third tool invocation instruction, until the final execution result is consistent with the execution result corresponding to the target task.
[0062] In some embodiments, the closed-loop inference device based on a multimodal large model further includes a memory generation module 440, which is used to record the order of task operations during the closed-loop inference process, denoted as a task operation sequence; and to record the number of times the task operation sequence is executed, to obtain the number of times the task operation sequence is executed; if the number of times the task operation sequence is executed is greater than a preset number of executions, then the task operation sequence that is greater than the preset number of executions is recorded as a historical task operation sequence; and the historical task operation sequence is stored in an operation memory created in advance or in real time.
[0063] In some embodiments, the inference module is configured to identify user intent based on the operation information of the interface elements to obtain user intent identification results; initiate an execution query based on the user intent identification results; capture screen visual information corresponding to the execution query results to obtain fourth screen visual information; input the fourth screen visual information into a multimodal large language model to obtain a fourth tool invocation instruction; and execute a fourth task corresponding to the type of the fourth tool invocation instruction based on the fourth screen visual information, until the final execution result is consistent with the execution result corresponding to the target task.
[0064] In some embodiments, the underlying control tool corresponding to the single tool invocation instruction includes atomic interface operation tools, which at least include mouse click tools, keyboard input tools, and screen scrolling tools.
[0065] In some embodiments, the screen visual information includes at least one of the following: a complete screenshot of the client screen, a web page document object model tree structure, and a native application's user control accessibility object model, or a combination of the above.
[0066] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0067] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0068] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0069] Figure 5 A schematic block diagram of an electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0070] Electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in ROM 502 or a computer program loaded into RAM 503 from storage unit 508. RAM 503 can also store various programs and data required for the operation of electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. I / O interface 505 is also connected to bus 504.
[0071] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0072] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the method based on closed-loop inference of a multimodal large model. For example, in some embodiments, the method based on closed-loop inference of a multimodal large model can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the method based on closed-loop inference of a multimodal large model described above can be performed. Alternatively, in other embodiments, computing unit 501 may be configured by any other suitable means (e.g., by means of firmware) to perform closed-loop inference based on a multimodal large model.
[0073] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0074] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0075] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0076] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including voice input, speech input, or tactile input).
[0077] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0078] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0079] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0080] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A closed-loop inference method based on a multimodal large model, characterized in that, include: Acquire visual information from the screen and the target task; According to the target task, a search operation is performed in the pre-generated or real-time generated operation memory. If a corresponding historical task operation sequence is found, the operation is executed according to the historical task operation sequence. If no corresponding historical task operation sequence is found, the screen visual information and target task instructions are input into the multimodal large language model to obtain a single tool call instruction, which includes operation information on screen coordinates or interface elements. According to the type of the single tool invocation instruction, the task corresponding to the type of the single tool invocation instruction is executed based on the screen visual information, and the screen visual information and the single tool invocation instruction are updated until the execution result is consistent with the execution result corresponding to the target task, so as to complete the closed-loop reasoning.
2. The closed-loop inference method based on a multimodal large model according to claim 1, characterized in that, The screen visual information and target task instructions are input into a multimodal large language model to obtain a single tool invocation instruction, including: If the screen visual information is the initial screen visual information, the initial screen visual information and the target task are fused to obtain the first fused information. The first fused information is then input into the multimodal large language model to obtain the first tool call instruction. The first tool invocation instruction includes initial screen coordinates or operation information of initial interface elements.
3. The closed-loop inference method based on a multimodal large model according to claim 2, characterized in that, The closed-loop reasoning process includes: According to the type of the first tool invocation instruction, based on the initial screen visual information, execute the first task corresponding to the type of the first tool invocation instruction to obtain the first execution result; If the first execution result indicates that the first task has not been completed, then update the screen visual information to obtain the second screen visual information; The second screen visual information, the target task, and the first historical record are used to obtain the second fusion information. The second fusion information is then input into the multimodal large language model to obtain the second tool call instruction. According to the type of the second tool invocation instruction, the second task corresponding to the type of the second tool invocation instruction is executed based on the second screen visual information to obtain the second execution result; If the second execution result indicates that the second task has not been completed, then update the screen visual information to obtain the third screen visual information; The third screen visual information, the target task, and the second historical record are fused to obtain the third fused information. The third fused information is then input into the multimodal large language model to obtain the third tool invocation instruction. According to the type of the third tool invocation instruction, the third task corresponding to the type of the third tool invocation instruction is executed based on the third screen visual information until the final execution result is consistent with the execution result corresponding to the target task.
4. The closed-loop inference method based on a multimodal large model according to claim 1, characterized in that, The generation of the operational memory includes: Record the sequence of task operations during the closed-loop reasoning process, denoted as the task operation sequence; and... Record the number of times the task operation sequence is executed to obtain the number of times the task operation sequence is executed; If the number of times the task operation sequence is executed is greater than the preset number of times, then the task operation sequence that is executed more than the preset number of times is recorded as a historical task operation sequence; The historical task operation sequence is stored in an operation memory bank created in advance or in real time.
5. The closed-loop inference method based on a multimodal large model according to any one of claims 1-4, characterized in that, The closed-loop reasoning also includes: The user intent is identified based on the operation information of the interface elements, and the user intent identification result is obtained. Based on the user intent recognition result, initiate an execution query; Capture the screen visual information corresponding to the query result to obtain the fourth screen visual information; The visual information from the fourth screen is input into the multimodal large language model to obtain the fourth tool invocation command; According to the type of the fourth tool invocation instruction, the fourth task corresponding to the type of the fourth tool invocation instruction is executed based on the visual information of the fourth screen until the final execution result is consistent with the execution result corresponding to the target task.
6. The closed-loop inference method based on a multimodal large model according to any one of claims 1-4, characterized in that, The underlying control tools corresponding to the single tool call command include atomic interface operation tools, which include at least mouse click tools, keyboard input tools, and screen scrolling tools.
7. The closed-loop inference method based on a multimodal large model according to any one of claims 1-4, characterized in that, The screen visual information includes at least one of the following: a complete screenshot of the client screen, a web page document object model tree structure, and a native application's user control accessibility object model, or a combination of the above.
8. A closed-loop inference device based on a multimodal large model, characterized in that, include: The acquisition module is used to acquire screen visual information and target tasks; The search module is used to perform a search operation in a pre-generated or real-time generated operation memory based on the target task. If a corresponding historical task operation sequence is found, the operation is executed according to the historical task operation sequence. The reasoning module is used to, if no corresponding historical task operation sequence is found, input the screen visual information and the target task instruction into the multimodal large language model to obtain a single tool call instruction, which includes operation information on screen coordinates or interface elements; and, according to the type of the single tool call instruction, execute the task corresponding to the type of the single tool call instruction based on the screen visual information, update the screen visual information and the single tool call instruction, until the execution result is consistent with the execution result corresponding to the target task, so as to complete the closed-loop reasoning.
9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the closed-loop inference method based on a multimodal large model as described in any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, in, The computer instructions are used to cause the computer to execute the closed-loop inference method based on a multimodal large model as described in any one of claims 1-7.
Citation Information
Patent Citations
Task instruction execution method and device, storage medium and electronic device
CN119088899A
Digital human interaction control method, system and equipment based on screen recognition
CN120066282A
Intelligent task guiding method, device and system based on digital assistant
CN120653334A
Auxiliary operation method, electronic equipment, storage medium and computer program product
CN120670078A