Task understanding and executing system based on multi-mode intelligent agent
Through the multimodal intelligent system, the task semantic analysis and plug-in execution under multimodal input is realized, which solves the problem of insufficient semantic understanding of intelligent assistants in complex tasks, improves the intelligence and adaptability of the system, and is suitable for office automation, intelligent assistants and other scenarios.
Patent Information
- Application Number
- CN202510464608.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-04-11
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
AI Technical Summary
The existing intelligent assistants lack the ability to understand semantics, disassembly and execute verification of complex tasks, and lack dynamic decision-making and self-learning capabilities, resulting in insufficient usability and intelligence level in real and complex environments.
The multimodal agent system is adopted, through multimodal input such as voice, text, and files, and uses a large language model to perform task semantic analysis and intention recognition, combining knowledge memory and self-optimization to realize task disassembly, plug-in execution, and result verification, and supports dynamic registration and update of plug-ins, and has the ability to learn independently.
It improves the availability and intelligence level of the agent in complex environments, supports multimodal natural interaction, has the ability to understand complex semantics and self-learning, is highly scalable, and adapts to the needs of multi-industry operations.
Smart Images

Figure CN120337977A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence, human-computer interaction, and computer automation control, and particularly to a task understanding and execution system based on a multimodal agent, which is applicable to scenarios such as office automation, intelligent assistants, and cross-platform control. Background Art
[0002] With the development of natural language processing (NLP), speech recognition, image recognition, and multimodal large models, artificial intelligence has made significant progress in task understanding and automatic execution. Existing intelligent assistants are mostly limited to simple voice command control and lack the ability to understand the semantics of complex tasks, decompose tasks, and verify task execution. Although traditional RPA tools can achieve process automation, they lack flexibility and cannot make dynamic decisions and self-learn.
[0003] Therefore, there is an urgent need for a system that combines multimodal interaction, human-like semantic understanding, dynamic task execution, and continuous learning capabilities to improve the usability and intelligence level of agents in real and complex environments. Summary of the Invention
[0004] In view of the limitations or improvement requirements of existing intelligent assistants, the present invention provides a task understanding and execution system based on a multimodal agent, which realizes task semantic parsing, instruction execution, result verification, knowledge memory, and self-optimization under multimodal inputs such as natural language, speech, and images, and improves the usability and intelligence level of agents in real and complex environments.
[0005] To achieve the above object, the present invention adopts the following technical solutions: The system includes the following modules: Input Module: including speech recognition, text input, and file input interfaces, which are used to receive multimodal input data; Semantic Understanding Module: uses a large language model (LLM), or a multimodal large model (Multimodal LLM), or an agent to perform intent recognition and task decomposition on the input content, generates a structured task queue, and verifies the task execution process. According to the context and expected goals, it analyzes whether the subtasks are successfully completed. If not up to standard, it returns an alternative execution path (such as changing plugins or rewriting execution instructions); Task Scheduling Module: is used to receive the structured task plan generated by the semantic understanding module and is responsible for the orderly organization and execution control of subtasks. This module supports task queue management, concurrent control, task status monitoring, and exception handling. During the execution process, this module coordinates the plugin execution module to complete specific tasks and triggers a retry or interruption mechanism according to the plugin execution status when necessary; Plugin Execution Module: It is used to call the corresponding plugin components according to the execution instructions issued by the Task Scheduling Module to complete specific subtasks, and supports multi-threaded or asynchronous execution of multiple task plugins; it supports dynamic registration, uninstallation, and update of plugins. The plugins are encapsulated with a unified interface specification and can adapt to different platforms or software environments; Result Feedback Module: It is used to feedback the task execution results and their verification information to the user in a multimodal manner, and supports the user to evaluate, correct, or confirm the results. The feedback information can be used to optimize subsequent task executions; Knowledge Memory Module: It is used to record historical task data, user interaction feedback, and execution verification information, construct a knowledge graph and an experience database, and support the Semantic Understanding Module and the Task Scheduling Module to optimize the understanding and execution strategies of the current task based on historical knowledge; Autonomous Learning Module: It is used to dynamically optimize the semantic understanding module, task scheduling module, and plugin execution strategies based on the historical task data and user feedback accumulated in the knowledge memory module, and improve the processing efficiency and accuracy of the system under similar tasks.
[0006] Preferably, the Semantic Understanding Module can be a large language model (LLM), or a multimodal large model (MultimodalLLM), or an agent, which realizes the reception, fusion, and intention reasoning of multi-source information; it can be called through the API cloud service or run through the local deployment method, and these two calling methods can be dynamically switched in the system according to user configuration.
[0007] Preferably, the Semantic Understanding Module can independently develop, install, and upgrade plugins according to requirements.
[0008] Preferably, the Task Scheduling Module supports synchronous and asynchronous subtask execution methods, has task queue management and status tracking functions, and can judge whether to retry or skip tasks based on the execution status; Preferably, the Task Scheduling Module and the Plugin Execution Module run in coordination, support parallel scheduling of multiple tasks, and have the ability to record task execution logs and handle failure tolerance; Preferably, the scheduling decision basis of the Task Scheduling Module is determined by the task execution graph or execution plan provided by the Semantic Understanding Module; Preferably, after the plugin execution is completed, the Task Scheduling Module structurally encapsulates the execution results and feeds them back to the Semantic Understanding Module; the Semantic Understanding Module makes a consistency judgment based on the execution results and the task objectives, and generates subsequent task flow decisions, including continuing, retrying, replacing plugins, or requesting user intervention, and supports multi-round understanding and dynamic adjustment of intermediate results to improve the overall success rate and execution robustness of multi-step tasks.
[0009] Preferably, the plugin execution module accesses multiple functional plugins through a unified call interface. Each plugin is encapsulated as an independent functional unit, supporting functions such as file operations, image processing, table editing, web automation, and system control; The plugin execution module supports the mixed use of local plugins and remote plugins. Remote plugins can send remote instructions and return results through protocols such as HTTP and WebSocket, and support task interruption and fault tolerance mechanisms; When the plugin is first installed or run, the plugin execution module reads the function description file of the plugin (such as JSON or YAML), and passes the information including function name, scope, input and output formats, platform compatibility, etc. to the semantic understanding module; The semantic understanding module performs semantic parsing and vector encoding on the plugin function based on the large language model, and registers the results in the semantic plugin library; In subsequent tasks, the system can automatically retrieve and call the plugin with the most matching function through intent matching, so as to realize intelligent plugin selection and call.
[0010] Preferably, the result feedback module supports multiple output methods such as text, voice, image, and file, and can select the feedback form according to the task type and user preferences; The result feedback module includes a user interaction interface, allowing users to confirm, score, modify or append opinions on the execution results. The feedback data is stored for subsequent task optimization of the system; The result feedback module is connected to the knowledge memory module, and can adjust the task understanding model or plugin selection strategy according to user feedback, realizing autonomous learning and optimization.
[0011] Preferably, the knowledge memory module can record the context information of task execution, user instruction intent, plugin execution path and verification feedback results, and establish classification labels and indexes for different types of tasks; The knowledge memory module constructs a knowledge graph based on a graph database or an embedded vector storage structure, supporting fast matching of similar tasks and intelligent recommendation of plugin combination strategies; The knowledge memory module supports incremental update and version control mechanisms, and can track the changes of models, plugins and tasks, realizing traceable learning process management; The knowledge memory module is linked with the semantic understanding module, introducing historical experience in the task understanding process to improve the understanding accuracy and execution success rate of similar tasks.
[0012] Preferably, the autonomous learning module supports supervised learning, reinforcement learning or rule evolution methods, and optimizes the task processing strategy in combination with the user feedback results; The autonomous learning module periodically evaluates the execution success rate and user satisfaction metrics in historical task data, and dynamically adjusts the plugin selection priority and scheduling logic accordingly; After the cumulative number of system tasks reaches a set threshold, the autonomous learning module extracts historical task records, execution feedback, and user behavior samples from the knowledge memory module to generate a training dataset; the training data can be used to perform online optimization or local fine-tuning on the large model or prompt template called by the semantic understanding module to improve the understanding accuracy and response matching degree of the system in subsequent similar tasks.
[0013] The present invention has the following beneficial effects: 1) Support multi-modal natural interaction and improve the naturalness of human-computer dialogue; 2) Have the ability to understand complex semantics and achieve automatic task decomposition and dynamic control; 3) Have the ability of result verification and self-learning, enhancing the reliability and adaptability of the system; 4) Provide a unified plugin execution framework, with strong scalability and wide adaptability to application scenarios; 5) Can improve the operation efficiency of multiple industries and accelerate the development of social production. Description of the Drawings Figure 1 It is a schematic diagram of the system architecture of the task understanding and execution system based on a multi-modal intelligent agent in an embodiment of the present invention. Figure 2 It is a task scheduling and execution flow chart in an embodiment of the present invention. Figure 3 It is a knowledge memory and autonomous learning architecture diagram in an embodiment of the present invention. Figure 4 It is a plugin installation and registration flow chart in an embodiment of the present invention. Detailed Embodiments
[0014] The following details the system of the present invention in combination with specific embodiments and drawings. Obviously, the described embodiments are a part of the present disclosure, rather than all of the embodiments. It should be understood that the embodiments are only for illustrative purposes and do not constitute a limitation on the protection scope of the present invention.
[0015] It should be understood that the "embodiments" mentioned herein mean that specific features, structures, or operation modes associated with the embodiments may be included in at least one embodiment of the present invention. The "embodiments" that appear multiple times in the specification are not necessarily referring to the same embodiment, and these embodiments do not exclude each other. For those skilled in the art, the features between different embodiments described herein can be freely combined and replaced according to specific application scenarios, still within the protection scope of the present invention.
[0016] The term "and / or" used in this document represents the logical relationship between multiple related objects, and its meaning includes, but is not limited to, the following three situations: only the former exists, only the latter exists, and both the former and the latter exist simultaneously. For example, "A and / or B" should be understood to include the three possibilities of "only A", "only B", and "A and B exist simultaneously". In addition, the character " / " appearing in the specification usually indicates an "or" relationship between the terms before and after it, indicating any one or more of multiple options, depending on the context semantic environment.
[0017] As Figure 1 shown is a schematic diagram of the system architecture of the task understanding and execution system based on a multi-modal agent in an embodiment of the present invention, including an input module, a semantic understanding module, a task scheduling module, a plugin execution module, a result feedback module, a knowledge memory module, and an autonomous learning module. The modules communicate with each other using the standard JSON protocol.
[0018] In some embodiments, an embodiment of the present invention encapsulates the system including the above modules into a computer desktop software. The input module can be a program module developed using programming languages such as Python, C#, and C++. It provides a text input field, a file pick-up button, and a start microphone listening button to meet different input needs of users, obtains the original requirement description data, and sends it to the semantic understanding module.
[0019] Specifically, the semantic understanding module can be Deepseek-R1, or OpenAI o1 / GPT-4o, or other multi-modal large models; it can be deployed locally or can be implemented by calling the cloud API interface; it receives the requirement description data from the input module, performs intent recognition and task decomposition on the input content, generates a structured JSON subtask queue, and sends it to the task scheduling module one by one (sequential tasks) or all (parallel tasks); and verifies the task execution process, analyzes whether the subtasks are successfully completed according to the context and expected goals, and if not up to standard, returns an alternative execution plan (such as changing the plugin or rewriting the execution instruction).
[0020] Specifically, the task scheduling module can be a program module developed using programming languages such as Python, C#, and C++. It receives the structured task plan generated by the semantic understanding module and is responsible for the orderly organization and execution control of subtasks. This module supports task queue management, concurrent control, task status monitoring, and exception handling. During the execution process, this module uses the JSON protocol to input the received subtask parameters into the input interface of the specified plugin, coordinates the plugin execution module to complete specific tasks, and triggers a retry or interruption mechanism according to the plugin execution status when necessary.
[0021] Specifically, the plugin execution module contains many plugins. Each plugin can be a program module developed using programming languages such as Python, C#, C++, Java, JavaScript, etc., and supports functions such as file operations, image processing, table editing, web automation, and system control. When a certain plugin is called by the task scheduling module, it receives the input parameters of the subtask and returns the result to the task scheduling module after execution. The task scheduling module then sends the result to the semantic understanding module to analyze whether it meets the requirements.
[0022] In some embodiments, the plugin registration mechanism is the plugin description file + registration directory method. Each plugin should include: The main execution file (Python / JS / executable program, etc.) The plugin description file (such as plugin.json) Example of the plugin description file: { "name": "ExcelChartGenerator", "description": "Excel plugin for generating charts", "type": "function", "tags": ["excel", "chart", "report"], "entry": "main.py", "platform": ["Windows", "Linux"], "input_format": "Excel path + field", "output_format": "Chart file path", "version": "1.0.2" } All plugins are installed in the specified system plugin directory, such as: / plugins / ├── ExcelChartGenerator / │ ├── plugin.json │ └── main.py ├── FileUploader / ├── PDFConverter / When the system starts, the plugin registration process is executed. The steps are as follows: 1. When the system starts, scan the / plugins directory; 2. For each subdirectory, parse the plugin.json; 3. Register the plugin information into the Plugin Registry; 4. The plugin registry is a globally indexable data structure (it can be in-memory / cache / database); 5. The plugins are indexed by tag, type, platform support, version number, etc. and are retrievable by the semantic understanding module; 6. The plugin entry file path and the call parameter format are uniformly encapsulated and called by the task control module.
[0023] In some embodiments, the registration of the plugin is a semantic-driven plugin registration mechanism, such as using descriptive manifest + embedding retrieval. Compared with the traditional hard-coded registration method, it has stronger flexibility and intelligent adaptation capabilities. When the plugin is installed / loaded, it actively sends information such as its function description, type, applicable scope, call format, etc. to the semantic understanding module, enabling the large model to "understand" the capabilities of the plugin, thus supporting future natural language task calls and matches. The plugin carries structured descriptions (such as JSON). When the system installs or runs, it calls the semantic understanding module, and the model parses out the purpose, input format, applicable scenarios, keyword tags, etc. of the plugin, and finally generates knowledge entries (for subsequent "intention → plugin" matching) compatible with the semantic retrieval system, as Figure 4 shown.
[0024] Specifically, the result feedback module can be a program module developed using programming languages such as Python, C#, C++. It has a multi-modal output interface and supports multiple output methods such as text, voice, image, file, etc. The front end can be implemented using GUI frameworks such as Flutter / Web / Qt, etc., and the feedback form can be selected according to the task type and user preferences. Users are allowed to confirm, score, modify, or append comments on the execution results. The feedback data format is recorded in structured JSON: task ID, plugin ID, execution result, user feedback status (confirmation / negation / modified content). It communicates with the semantic understanding module through an internal event bus (such as a message queue) or function callback, can dynamically switch the feedback presentation method (such as chart preview, editable table, etc.) based on the task type (such as table, chart, text summary), and communicates with the knowledge memory module in standard JSON format; the feedback data is stored for subsequent task optimization of the system.
[0025] In some embodiments, the knowledge memory module can be a database implemented based on a dual mechanism of graph structure and embedding vectors. All historical tasks are stored in the following two forms: 1. Graph database structure (using Neo4j): 1) Node types include: task nodes, plugin nodes, verification nodes, and feedback nodes; 2) Edges represent behaviors such as "invoke", "verification passed", and "user confirmation"; 2. Semantic embedding vectors (such as Sentence - BERT + FAISS): used for fuzzy matching of task intents.
[0026] After the input text of the task is vectorized, a similarity recall operation is performed in the vector database to find the historical task with the highest matching degree (such as score > 0.85); then the plugin execution path of this historical task is extracted from the graph database (such as Plugin_A → Plugin_B → Plugin_C); the task scheduling module uses this plugin path as the preferred execution path. If the plugin interface has changed, the system calls the registry to verify compatibility; after the task is executed and the user confirms, the system will record the task → plugin → verification → user feedback path, update the task nodes in the graph database, and update the vector index of this task text for future recall.
[0027] In some embodiments, the autonomous learning module includes three functional units: a sample generation unit, a policy learning and model fine - tuning unit, and a deployment and version management unit. The sample generation unit is used to extract historical task data from the knowledge memory module, filter task records that meet "execution successful" and "user confirmation or positive feedback", and convert them into a supervised learning sample format. The samples include fields such as input text, structured task description, and plugin call path. This unit can use Python scripts combined with an SQLite database or JSONL format to implement batch sample construction, and the samples can be used for policy training or model fine - tuning. The policy learning and model fine - tuning unit is used to optimize the prompt template, plugin scheduling policy, or semantic model ontology of the semantic understanding module based on the sample data. This unit supports one or a combination of the following optimization methods: 1. Prompt template optimization: According to the statistical rules of the samples, update the template structure in prompt engineering to guide the language model to generate more accurate task decomposition or plugin call instructions; 2. Plugin strategy sorting learning: Use machine learning methods (such as XGBoost, LightGBM, RankNet, reinforcement learning, etc.) to train a plugin selection priority model to improve the accuracy and execution success rate of plugin combinations; 3. Model parameter fine - tuning: When there is the ability to deploy a local large - model, parameter - efficient fine - tuning techniques such as PEFT (such as LoRA) can be used to perform lightweight online tuning on the semantic model to improve the system's understanding ability for specific task types or contexts.
[0028] The deployment and version management unit is used to save, switch, and load the optimization results. This unit can save the optimized prompt template, plugin policy model, or fine-tuned model as a versioned file (such as.json,.pkl,.pt), and establish a version index mechanism. The system supports dynamically selecting the optimal version for execution according to the task type or evaluation metrics, and the user can also choose to roll back to an old version. This unit can also evaluate and eliminate different model versions in combination with indicators such as task success rate and user rating.
[0029] Further, a specific example is used for illustration. For example, in the office automation scenario, the user inputs the instruction by voice: "Please help me organize the PDF contract files downloaded yesterday in the contract folder on drive D and generate an Excel report containing the contract number, company name, and amount." The input module obtains the user's voice and inputs it into the semantic understanding module. The semantic understanding module recognizes the task intention as "extracting PDF content and generating a report"; and decomposes the task into: 1. Find the target file; 2. Extract data; 3. Generate an Excel file; Then generate a structured task queue JSON: { "Execution method": "Sequential", "Step 1": { "Target": "Obtain the file name list", "Call plugin": View file, "Input parameter": "d: / contract" }, "Step 2": { "Target": "Extract the contract number, company name, and amount field values", "Call plugin": PDF check, "Input parameter": "Contract name, contract number, company name, amount" }, "Step 3": { "Target": "Generate an Excel file", "Call plugin": Generate Excel file, "Input parameter": { "Title": "Contract report at a certain time", "Contract number": [Contract number list], "Company name": [Company name list], "Amount": [Amount list] }} } After the semantic understanding module analyzes the task, it concludes that the execution should be sequential. Therefore, first execute subtask step 1, which is to { "Step 1": { "Goal": "Obtain a list of file names", "Call plugin": View files, "Input parameter": "d: / contract" }, } Send it to the task scheduling module. The task scheduling module calls the view file plugin, sends the input parameter "d: / contract" to the view file plugin, monitors the execution status of the plugin execution module, and waits for the return result after execution: a list of contract file names. The task scheduling module sends the list of contract file names to the semantic understanding module. The semantic understanding module determines whether the current result meets the expectations based on the context. If it does not meet the expectations, replace the plugin and regenerate the task queue, and issue subtask step 1; if it meets the expectations, start executing subtask step 2, which is to { "Step 2": { "Goal": "Extract contract numbers, company names, and amount field values", "Call plugin": PDF check, "Input parameter": "List of contract names, contract numbers, company names, amounts" }, } Send it to the task scheduling module. The task scheduling module calls the PDF check plugin, sends the input parameter: "List of contract names, contract numbers, company names, amounts" to the PDF check plugin. After execution, the return result is obtained: a list of contract numbers, a list of company names, and a list of amounts. Then the task scheduling module sends this result to the semantic understanding module. The semantic understanding module checks whether the result meets the expectations after receiving it. If it does not meet the expectations, replace the plugin and regenerate the task queue, and issue subtask step 2; if it meets the expectations, start executing subtask step 3, which is to { "Step 3": { "Goal": "Generate an Excel file", "Call plugin": Generate Excel file, "Input parameter": { "Title": "Contract Report for a Certain Time", "Contract number": [List of contract numbers], "Company name": [List of company names], "Amount": [List of amounts] } } } Send to the task scheduling module, which calls the Excel file generation plugin and takes the input parameters: { "Title": "Contract Report for a Certain Time", "Contract Number": [List of contract numbers], "Company Name": [List of company names], "Amount": [List of amounts] } Send to the Excel file generation plugin. After execution, the returned result is: the contract report Excel file. Then the task scheduling module sends the contract report Excel file to the semantic understanding module. The semantic understanding module determines whether the obtained file meets the expectations according to the context. If it does not meet the expectations, the plugin is replaced to regenerate the task queue and issue sub-task step 3; if it meets the expectations, the contract report Excel file is sent to the result feedback module. The result feedback module previews the file for the user to view, review, and provides a file save button for the user to save the file locally.
[0030] So far, this automated office task has been fully completed. After the user confirms the result, the semantic understanding module sorts out the process outputs of this task, stores them in the knowledge memory module to form an empirical data archive, which can be directly called when encountering similar tasks in the future, or the execution strategy of the current task can be optimized according to historical error information to avoid making the same mistake again. The task execution process is as Figure 2 shown.
[0031] Furthermore, another specific example is used for illustration. This embodiment shows the whole process of user intervention feedback and preference learning when the task result of the system does not meet the user's expectations for the first time. The specific steps are as follows: 1. The user enters a natural language instruction through the desktop input box: "Please convert the document into a table format"; 2. The input module converts the instruction into structured text and transmits it to the semantic understanding module; 3. The multi-modal large model parses the task intention as: 1) Open the specified document (PDF or picture); 2) Extract the content structure; 3) Generate a structured table file (such as.xlsx); 4. The task scheduling module calls the default plugin A to execute the task according to the structured task output by the model; 5. Plugin A successfully outputs a table, but the table structure does not meet the user's expectations (such as data misalignment, information missing); 6. The system displays the preliminary results through the feedback module and provides a "re-execute" option; 7. The user manually switches to Plugin B in the plugin selection interface and re-executes the task; 8. Plugin B executes successfully and outputs a table file that meets the user's intention; 9. The user clicks the "Confirm Results" button, and the feedback module records this user confirmation behavior; 10. The system automatically transfers: 1) The current task context (task content, input type, output format); 2) The user's manual plugin replacement behavior; 3) The user's final confirmed plugin ID to the knowledge memory module; 11. The knowledge memory module marks Plugin B as the "preferred recommended plugin"; 12. The next time the user initiates a similar task, the system defaults to calling Plugin B and asks through the prompt interface: "Do you want to continue using the previous plugin?"; 13. If the user confirms, the plugin selection process is skipped; if cancelled, other plugins can still be selected to continue optimizing the task process.
[0032] Furthermore, another specific example is used for illustration. This embodiment demonstrates the mechanism of how the system collaborates with the mobile terminal across devices to complete tasks. The specific steps are as follows: 1. The user enters a natural language instruction in the desktop system: "Please take a photo of the paper contract with the mobile phone and send it to the computer to generate a PDF"; 2. The input module converts this instruction into text form and transfers it to the semantic understanding module; 3. The multi-modal model parses the task objectives as: 1) Take a photo on the mobile terminal; 2) Transmit the picture back to the desktop; 3) Perform image and text recognition and format conversion on the desktop; 4. The model splits this task into two collaborative subtasks and generates a structured task graph with an "execution platform" label: 1) Subtask 1: Mobile terminal plugin → Open the camera to take a photo; 2) Subtask 2: Desktop terminal plugin → Image recognition → PDF generation; 5. The desktop task scheduling module calls the local area network communication plugin to send a task instruction to the bound mobile terminal; 6. After receiving the request, the mobile terminal plugin invokes the camera application; 7. After the user finishes taking the photo, the image is compressed and uploaded back to the temporary cache path of the desktop system; 8. The desktop plugin performs image OCR (text extraction) and typesetting processing; 9. The PDF plugin converts the typeset content into a standard PDF document; 10. The feedback module prompts in the desktop pop-up window: "The contract PDF has been generated", and provides options to open and save; 11. The entire task process (including the behavioral trajectory of the mobile plugin) is written into the knowledge memory module; 13. If the user confirms satisfaction, the system can save this process as a "cross-device task template" for future quick invocation.
[0033] Further, another specific example is used for illustration. This embodiment shows how the system recalls historical experience, automatically generates a task path, and completes the process of execution without intervention: 1. The user inputs "Please organize the documents and archive them by date naming" through text; 2. The input module receives the text and transfers it to the semantic understanding module; 3. The model analyzes the task objectives as follows: 1) Traverse the document files in the specified directory; 2) Obtain the document creation / modification date; 3) Rename the file according to the date rule; 4) Archive by time to the corresponding subdirectory; 4. The system structures the task into the following steps: 1) Obtain the document path; 2) Extract the document metadata; 3) Match the naming rule; 4) Execute the archiving classification; 5. The semantic understanding module calls the knowledge memory module to retrieve whether there is a similar task history in the system; 6. The system discovers that the user performed a task with the naming rule YYYY-MM-DD_file name in 2024; 7. The system automatically reuses this naming format and the plugin combination path (file scanning + file renaming + archiving plugin); 8. The plugin execution module starts the task chain and performs naming and archiving operations according to the file modification time; 9. All files are moved to the subdirectory named by year or month, and the path example: / Documents / 2025 / 04 / contract_A.pdf; 10. After the execution is completed, the feedback module shows a preview of the archiving structure (such as a file tree structure or a compressed package link); 11. When the user confirms that there are no errors, the system records the task path and plugin combination of this task into the knowledge memory module and marks it as "automatic archiving template". 12. When subsequent similar tasks are triggered, the system will prompt: "Do you want to use the archiving template v2024.12" to achieve automated optimization.
[0034] After the cumulative number of tasks executed by the system reaches a set threshold (such as 100), the autonomous learning module can call the historical task data stored in the knowledge memory module, including task inputs, plugin call chains, execution verification results, and user feedback, to construct a structured training sample set. This sample set can be used to continuously optimize the multi-modal models or prompt templates used in the semantic understanding module, specifically including: adjusting the intent recognition weights, optimizing the task decomposition templates, updating the plugin selection strategies, or performing lightweight model fine-tuning (such as LoRA / PEFT). The optimization methods support two modes: local adaptation and cloud learning. The results of model fine-tuning can be saved versionally to improve the system's response accuracy to specific user usage habits and contexts. In some embodiments, the knowledge memory and autonomous learning architectures are as Figure 4 shown.
[0035] Generally speaking, through the task understanding and execution system based on multi-modal agents proposed by the present invention, the following beneficial effects can be achieved: 1. Based on multi-modal input and semantic understanding capabilities, this system supports the fusion perception and intent reasoning of multi-source information such as voice, files, and text. Combined with a plugin-based execution architecture, it can automatically decompose user tasks and complete end-to-end execution operations without manual intervention, and is widely applicable to application scenarios such as desktop office, process automation, and intelligent assistants. The system has the characteristics of a human-like agent, can understand natural language instructions and automatically complete cross-platform tasks; 2. The system adopts a unified plugin interface specification and dynamic registration mechanism. Plugins can be extended, replaced, or uninstalled at any time, with extremely strong scalability and platform adaptation capabilities. Whether local or remote plugins can be incorporated into unified scheduling, and the invocation of plugins is dynamically decided by the semantic understanding module. The system can continuously optimize the plugin priorities and combination strategies to improve the execution efficiency and success rate; 3. The present invention introduces a knowledge memory module and an autonomous learning module, which can record user behaviors, task execution paths, and verification feedback results to form a sustainable and cumulative empirical knowledge system. The system can perform lightweight fine-tuning of the semantic understanding model or optimize the prompt templates after the task accumulation reaches the set threshold, realizing personalized adaptation to specific users or scenarios and enhancing the intelligent performance in long-term use; 4. The system introduces an execution verification mechanism and a feedback mechanism. The execution results of the plug-ins will be transmitted back to the semantic understanding module for semantic consistency judgment, thereby forming the capabilities of dynamic error correction, rescheduling, and path repair during the task execution process, and having high robustness and stability to adapt to complex and changing actual task environments; 5. This system supports module hot swapping and model configuration switching, does not rely on specific commercial models or closed architectures, can call open-source large models, and can also integrate commercial cloud services, adapts to various computing capabilities and deployment requirements, is suitable for flexible deployment in terminal devices, enterprise local area networks, or cloud platforms, and has strong practicality and promotion value; 6. Different from traditional process scripts or single-assistant systems, this system forms a "virtual humanoid intelligent agent" with understanding, execution, feedback, and self-optimization capabilities through the method of semantic-driven + plug-in call + autonomous learning, and can understand the user context, dynamically process tasks, and continuously evolve like a "digital assistant", and has extremely strong versatility and long-term evolution potential.
[0036] It should be understood that the technical features in the embodiments described herein can be combined arbitrarily according to actual needs. For the sake of simplicity, the specification does not list all possible combinations of technical features in detail, but as long as the combinations of these features do not have a technical conflict, they should be regarded as one of the disclosed contents of this specification. In addition, expressions such as "in some embodiments", "for example", and "such as" used in the specification are intended to illustrate the technical solutions of the present invention by way of example, rather than constituting any limitation to the protection scope of the present invention. Various derivative embodiments, equivalent deformations, replacement solutions, and combinations of technical solutions made by those of ordinary skill in the art after reading this specification without creative efforts shall fall within the protection scope of the present invention.
Claims
1. A task understanding and execution system based on a multi-modal intelligent agent, characterized in that, Including: Input module: It includes speech recognition, text input, and file input interfaces for receiving multi-modal input data; Semantic understanding module: Utilizes a large language model (LLM), or a multi-modal large model (Multimodal LLM), or an agent to perform intent recognition and task decomposition on the input content, generates a structured task queue, and verifies the task execution process. Analyzes whether subtasks are successfully completed based on the context and expected goals. If not up to standard, returns an alternative execution path (such as changing plugins or rewriting execution instructions); Task scheduling module: Used to receive the structured task plan generated by the semantic understanding module, and is responsible for the orderly organization and execution control of subtasks. This module supports task queue management, concurrency control, task status monitoring, and exception handling. During the execution process, this module coordinates the plugin execution module to complete specific tasks, and triggers a retry or interruption mechanism according to the verification results when necessary; Plugin execution module: Used to call the corresponding plugin components to complete specific subtasks according to the execution instructions issued by the task scheduling module, and supports multi-threaded or asynchronous execution of multiple task plugins; supports dynamic registration, uninstallation, and update of plugins. Plugins are encapsulated with a unified interface specification and can be adapted to different platforms or software environments; Result feedback module: Used to feedback the task execution results and their verification information to the user in a multi-modal manner, and supports the user to evaluate, correct, or confirm the results. The feedback information can be used to optimize subsequent task execution; Knowledge memory module: Used to record historical task data, user interaction feedback, and execution verification information, construct a knowledge graph and an experience database, and support the semantic understanding module and the task scheduling module to optimize the understanding and execution strategies of current tasks based on historical knowledge; Autonomous learning module: Used to optimize the semantic understanding module based on the historical task data and user feedback accumulated in the knowledge memory module, and improve the processing efficiency and accuracy of the system under similar tasks.
2. The system according to claim 1, further comprising: The semantic understanding module can be a large language model (LLM), or a multi-modal large model (Multimodal LLM), or an agent, which realizes the reception, fusion, and intent reasoning of multi-source information; all can be called through the API cloud service or run through local deployment, and these two calling methods can be dynamically switched in the system according to user configuration. The semantic understanding module can develop, install, and upgrade plugins autonomously according to requirements.
3. The system according to claim 1, further comprising: The task scheduling module supports synchronous and asynchronous subtask execution methods, has task queue management and status tracking functions, and can judge whether to retry or skip tasks based on the execution status; The task scheduling module and the plugin execution module operate in coordination, support parallel scheduling of multiple tasks, and have the ability to record task execution logs and handle failure tolerance; The scheduling decision basis of the task scheduling module is determined by the task execution graph or execution plan provided by the semantic understanding module; After the plug-in is executed, the task scheduling module will structure the execution result and feed it back to the semantic understanding module; the semantic understanding module will make a consistency judgment based on the execution result and the task goal, and generate subsequent task flow decisions, including continuing, retrying, replacing the plug-in or requesting user intervention, supporting multiple rounds of understanding and dynamic adjustment of intermediate results, and improving the overall success rate and execution robustness of multi-step tasks.
4. The system according to claim 1, Other features include: The plug-in execution module accesses multiple functional plug-ins through a unified calling interface, and each plug-in is encapsulated as an independent functional unit, supporting functions such as file operation, image processing, table editing, web page automation, system control, and software control; The plug-in execution module supports the mixed use of local plug-ins and remote plug-ins. The remote plug-in can issue remote instructions and transmit results through protocols such as HTTP and WebSocket, and supports task interruption and fault tolerance mechanisms; When the plug-in is installed or initialized, the plug-in execution module submits the plug-in description information to the semantic understanding module, including functional purpose, input and output format, applicable scenarios, key parameters, etc.; the semantic understanding module performs semantic analysis and feature encoding on the plug-in description based on the large language model, and registers it in the plug-in index library; the system supports calling the most suitable plug-in to complete the task through semantic matching according to the user's natural language instructions.
5. The system of claim 1, further comprising: The result feedback module supports multiple output modes such as text, voice, image, file, etc., and the feedback form can be selected according to the task type and user preference; The result feedback module includes a user interaction interface, allowing the user to confirm, score, modify or add comments on the execution results; The result feedback module is connected to the knowledge memory module, and can store user feedback data and user behavior samples in the knowledge memory module so as to adjust the task understanding model or plug-in selection strategy accordingly to achieve autonomous learning and optimization.
6. The system according to claim 1, Other features include: The knowledge memory module can record the context information of task execution, user instruction intention, plug-in execution path and verification feedback results, and establish classification labels and indexes for different types of tasks; The knowledge memory module builds a knowledge graph based on a graph database or an embedded vector storage structure, supporting fast similar task matching and intelligent recommendation plug-in combination strategy; The knowledge memory module supports incremental updates and version control mechanisms, can track the changes in models, plug-ins, and tasks, and implement traceable learning process management; The knowledge memory module is linked with the semantic understanding module to introduce historical experience into the task understanding process, thereby improving the understanding accuracy and execution success rate of similar tasks.
7. The system according to claim 1, Other features include: The autonomous learning module supports supervised learning, reinforcement learning or rule evolution, and optimizes task processing strategies in combination with user feedback results; The autonomous learning module periodically evaluates the execution success rate and user satisfaction index in the historical task data, and dynamically adjusts the plug-in selection priority and scheduling logic accordingly; After the cumulative system tasks reach the set threshold, the autonomous learning module extracts historical task records, execution feedback, and user behavior samples from the knowledge memory module to generate a training dataset; the training data can be used to perform online optimization or local fine-tuning on the large model or prompt template called by the semantic understanding module to improve the understanding accuracy and response matching degree of the system in subsequent similar tasks.
Citation Information
Cited By
Voice interaction method and apparatus, and XR device
CN120544579A
Automatic process execution method based on large language model
CN120780437A
Robot and intelligent interaction method thereof
CN120949939A
A robot and an intelligent interaction method thereof
CN120949939B
Task flow control agent for electric power intelligent monitoring and control method
CN121563404A