Agent directional evoking and intelligent matching system and method based on AI large model technology

By invoking a unified AI platform through a global hotkey monitoring module and a multimodal natural language parsing module, and combining a COT knowledge base and a dynamic matching engine, the limitations of the intelligent assistant system's entry point and task execution capabilities are resolved. This enables multi-tool collaborative operation and efficient task execution, improving user interaction efficiency and system security.

CN121704946APending Publication Date: 2026-03-20SHENZHEN YSSTECH INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511867678.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing intelligent assistant systems suffer from closed entry points and limited task execution capabilities in multi-tool collaboration scenarios, causing users to frequently switch application interfaces and be unable to perform substantive tasks, which contradicts the technological trend of one-stop intelligent collaboration.

Method used

A global hotkey monitoring module is used to invoke a unified AI platform. Combined with multimodal natural language parsing, COT knowledge base and dynamic matching engine, it enables multi-application collaborative operation through cross-tool execution module. It supports multimodal commands such as text, image and voice, and optimizes task execution through historical high-quality task image library and reordering model.

Benefits of technology

It provides a global entry point for collaborative operation of multiple tools, improves interaction efficiency and task execution quality, ensures that the AI ​​workbench can be launched with one click from any application interface, supports mixed input of multimodal commands, improves task planning efficiency and security, and ensures a consistent experience across platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121704946A_ABST
    Figure CN121704946A_ABST
Patent Text Reader

Abstract

The invention discloses an Agent directional evoking and intelligent matching system and method based on an AI large model technology, and the system comprises a global hot key monitoring module which is used for evoking a unified AI platform at any application interface through a system-level shortcut key; the multi-modal natural language analysis module is connected with the large language model and the multi-modal large model and is used for disassembling an instruction input by a user or a multi-modal instruction into structured sub-tasks; the COT knowledge base is used for storing a predefined Chain of Though execution chain template; the historical high-quality task mirror image library is used for storing a high-quality task execution scheme and a COT execution chain of the high-quality task execution scheme which are jointly labeled by the user and the AI; the dynamic matching engine is used for matching the COT knowledge base through vector semantic retrieval and outputting an optimal execution scheme through a reordering model; and the cross-tool execution module is used for automatically calling a plurality of independent tools in the operating system according to the execution scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence and operating system integration technology, specifically to an agent-oriented activation and intelligent matching system and method based on AI large model technology. Background Technology

[0002] Existing intelligent assistant systems suffer from three major technical bottlenecks: Firstly, the closed nature of entry points: products like Microsoft Copilot are limited to specific application ecosystems (e.g., only accessible within the Office suite), requiring users to frequently switch application interfaces to use AI functions, resulting in a broken workflow. Secondly, office robots like DingTalk Assistant are also unable to respond to commands outside of their host environment. Lack of task execution capability: Due to security restrictions, general dialogue models such as Wenxin Yiyan cannot access the local file system and operating system API, and can only provide information consultation but cannot perform substantive tasks such as file generation and code debugging. The aforementioned shortcomings force users to manually switch applications and repeatedly configure tasks, which contradicts the technological trend of "one-stop intelligent collaboration." Especially when multiple tools are linked together, such as code development and document processing, existing technologies cannot achieve "single-instruction-triggered full-link automation." Summary of the Invention

[0003] This invention aims to at least partially address one of the technical problems in related technologies. To this end, one objective of this invention is to propose an Agent-based targeted recall and intelligent matching system based on AI large-scale model technology, comprising: A global hotkey monitoring module is used to invoke a unified AI platform from any application interface via system-level shortcut keys; The multimodal natural language parsing module connects the large language model and the multimodal large model, breaking down user input commands or multimodal commands into structured subtasks; The COT knowledge base stores predefined Chain of Thought execution chain templates; A historical high-quality task mirror library stores high-quality task execution schemes and their COT execution chains, jointly annotated by users and AI. The dynamic matching engine matches the COT knowledge base through vector semantic retrieval and outputs the optimal execution plan through a re-ranking model; The cross-tool execution module automatically calls multiple independent tools in the operating system based on the execution plan.

[0004] Preferably, the global hotkey listening module is implemented through hook functions at the operating system level, and the invoked AI workbench is overlaid on the existing application interface in the form of a floating window.

[0005] Preferably, the dynamic matching engine performs the following operations: First, based on user instructions and context, it uses a combination of vector database and large language model to search the historical high-quality task mirror library and determine whether there are reusable existing similar high-quality task solutions. If such a historical task solution exists, then that solution will be used as the optimal execution solution. If it does not exist, the similarity between the subtask and the COT template is calculated using a vector database; The TOP-K candidate solutions are finely ranked using a re-ranking model; The COT execution chain with the highest output score is selected as the final solution.

[0006] Preferably, the cross-tool execution module includes a tool adaptation layer, which converts natural language instructions or model instructions into declarative tool invocation instructions. The instruction format includes action type, target tool, and parameter list.

[0007] Preferably, it also includes a dynamic path adjustment module, which monitors the execution status in real time during the execution of subtasks, and triggers the large language model or multimodal large model to replan the remaining subtasks when an execution deviation is detected.

[0008] Preferably, the system returns the task execution status in real time through a Streamable long connection channel and displays the progress visually within the AI ​​workbench.

[0009] Preferably, the cross-tool execution module supports the invocation of tools including but not limited to code editors, office software, version control systems, command-line terminals, and multimodal processing tools.

[0010] Preferably, it also includes a human-machine collaboration module, which forcibly triggers a confirmation window when performing critical sub-tasks involving file operations, data deletion, or external system access, pausing the task and waiting for the user's explicit decision before continuing execution.

[0011] Preferably, it also includes a real-time human-machine adjustment module, which visually displays the operation rules and intermediate execution status to the user during task execution, and provides a real-time dynamic adjustment interface. The user can intervene at any time to adjust the task parameters or process. The AI ​​system receives the adjustment instructions in real time and responds to execute them. This process does not forcibly pause the task process.

[0012] The Agent-based targeted invocation and intelligent matching method based on AI large model technology is applied to the Agent-based targeted invocation and intelligent matching system based on AI large model technology, including the following steps: S1. Receive user original instructions or multimodal instructions triggered by global hotkeys; S2. Decompose instructions into a sequence of structured subtasks using a large language model or a multimodal large model; S3. Prioritize searching the historical high-quality task image library for matching. If no match is found, then search the COT knowledge base to match each sub-task and generate a complete execution chain. S4. Call multiple utility APIs in the execution chain order to perform operations; S5. Push the execution status to the workbench interface in real time via a long connection channel; S6. At critical nodes, trigger mandatory confirmation based on the strategy or accept real-time adjustment instructions from users.

[0013] The above-mentioned solution of the present invention has at least the following beneficial effects: it breaks through the single interaction modality: it supports mixed input of multimodal commands such as text, image, and voice, solving the limitation of traditional AI systems that only support text commands; users can express their needs in natural ways such as screenshot annotation and voice description, which significantly improves the interaction efficiency; Dual improvement in task execution quality and efficiency: The introduction of a historical high-quality task mirror library enables self-evolution that becomes "smarter with use": Direct reuse of historical high-quality solutions avoids redundant calculations and reduces task planning time; The user rating mechanism drives continuous optimization of the knowledge base, and the execution success rate increases with the frequency of use. Global entry and cross-tool collaboration capabilities: Hotkey global activation ensures that the AI ​​workbench can be launched with one click from any application interface, breaking through the ecosystem barriers of tools such as Microsoft Copilot that are limited to specific applications; Cross-tool execution modules uniformly schedule heterogeneous resources such as code editors, office software, and multimodal processing tools to achieve "one command, automatic collaboration of multiple tools"; System compatibility and security: The tool adaptation layer shields operating system differences, ensuring a consistent experience across platforms; the declarative instruction protocol explicitly defines the scope of tool operations, preventing unauthorized access and improving system security.

[0014] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0016] Figure 1This is a flowchart of an Agent-based targeted recall and intelligent matching system based on AI large model technology provided in an embodiment of the present invention; Figure 2 This is a flowchart of a bias decision tree provided in an embodiment of the present invention; Figure 3 This is a flowchart of an Agent-based targeted recall and intelligent matching method based on AI large model technology provided in an embodiment of the present invention; Figure 4 This is a flowchart of task initiation provided in an embodiment of the present invention; Figure 5 This is the state synchronization diagram for step S5. Detailed Implementation

[0018] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0019] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "circumferential," and "radial," etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the accompanying drawings and are only for the convenience of describing this invention and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.

[0020] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0021] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0022] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can include direct contact between the first and second features, or contact between the first and second features through another feature between them. Furthermore, "above," "over," and "on top" of the second feature includes the first feature directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature includes the first feature directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.

[0023] The Agent-based targeted recall and intelligent matching system based on AI large model technology, an embodiment of the present invention, is described in detail below with reference to the accompanying drawings.

[0024] Please see Figure 1 In this embodiment, the system includes: a global hotkey monitoring module for invoking a unified AI platform from any application interface via system-level shortcut keys; a natural language parsing module connected to a large language model (LLM) to decompose user commands into structured subtasks; a COT knowledge base storing predefined Chain of Thought execution chain templates; a dynamic matching engine that matches the COT knowledge base through vector semantic retrieval and outputs the optimal execution plan through a reordering model; and a cross-tool execution module that automatically calls multiple independent tools in the operating system according to the execution plan. The working principle of this embodiment is as follows: Global hotkey monitoring module: Register system-level hotkeys through operating system-level hook functions (such as SetWindowsHookEx in Windows or CGEventTapCreate in macOS); When a user triggers a hotkey (such as Ctrl+Alt+A) in any application interface, the hook function captures the interrupt signal and invokes a unified AI workbench; This workbench is overlaid on the current application in the form of a floating window, realizing an interactive entry point with zero context switching; Multimodal Natural Language Parsing Module: Connects Large Language Model (LLM) and Multimodal Large Model (MLLM), supporting mixed command inputs such as text, images, and voice; MLLM synchronously parses multimodal information (such as table data in the screenshot + voice command "process this content"), outputting a structured subtask sequence to ensure machine readability and accuracy; Two-layer knowledge collaborative retrieval: COT knowledge base: stores predefined Chain-of-Thought execution chain templates, providing standardized task planning schemes; historical high-quality task mirror library: stores high-quality execution schemes (including success records and scores) jointly annotated by users and AI; the system prioritizes searching the historical library to match similar high-quality schemes, and if no match is found, it degrades to searching the COT library, forming a decision-making mechanism of "experience reuse as the main method and rule generation as the secondary method"; Cross-tool execution module: Through the tool adaptation layer, structured subtasks are converted into declarative tool call instructions (such as {"action type":"create_file", "target tool":"WPS", "parameter list":{...}}), automatically scheduling independent tools in the operating system (such as code editors, office software, multimodal processing tools) to achieve multi-application collaborative operation; Breaking through the limitation of single interaction modality: It supports mixed input of multimodal commands such as text, image, and voice, solving the limitation of traditional AI systems that only support text commands; users can express their needs in natural ways such as screenshot annotation and voice description, which significantly improves interaction efficiency; Dual improvement in task execution quality and efficiency: The introduction of a historical high-quality task mirror library enables self-evolution that becomes "smarter with use": Direct reuse of historical high-quality solutions avoids redundant calculations and reduces task planning time; The user rating mechanism drives continuous optimization of the knowledge base, and the execution success rate increases with the frequency of use. Global entry and cross-tool collaboration capabilities: Hotkey global activation ensures that the AI ​​workbench can be launched with one click from any application interface, breaking through the ecosystem barriers of tools such as Microsoft Copilot that are limited to specific applications; Cross-tool execution modules uniformly schedule heterogeneous resources such as code editors, office software, and multimodal processing tools to achieve "one command, automatic collaboration of multiple tools"; System compatibility and security: The tool adaptation layer shields operating system differences, ensuring a consistent experience across platforms; the declarative instruction protocol explicitly defines the scope of tool operations, preventing unauthorized access and improving system security.

[0025] In this embodiment, the global hotkey listening module is implemented through the hook function at the bottom layer of the operating system, and the AI ​​workbench that is invoked is overlaid on the existing application interface in the form of a floating window. Example 1: Implementation details of global hotkey monitoring and floating workbench I. Implementation mechanism of underlying hook functions: Windows environment: Use SetWindowsHookEx(WH_KEYBOARD_LL) to install a low-level keyboard hook, which can penetrate the application sandbox to capture global hotkeys; The hook procedure filters system key press messages and triggers a workbench wake-up event after recognizing preset key combinations (such as Ctrl+Alt+A).

[0026] macOS environment: Create an event listener using CGEventTapCreate to monitor the key combination kVK_ANSI_A and cmd+opt; After determining if the key combination matches in the event callback function, the NSWindow rendering layer is called.

[0027] II. Floating Window Overlay and Hierarchical Control Window property configuration: Set the top-level window flag Windows: WS_EX_TOPMOST; macOS: NSWindowLevelFloating; Enable the transparent background property (Alpha value 60%) to allow the underlying application content to be seen.

[0028] Event penetration handling: Enable event penetration in non-interactive areas of the workbench (such as transparent overlays): Windows: WS_EX_TRANSPARENT; macOS: ignoresMouseEvents; Only input fields, buttons, and other controls respond to clicks; other operations are directly passed to the underlying application.

[0029] III. Multi-application context capture Synchronously acquire the focused window information when the hook function is triggered: Windows calls GetForegroundWindow and GetWindowThreadProcessId; macOS obtains the current application identifier through NSWorkspace.shared.frontmostApplication; Dynamically load the application context into the workbench: Browser scenario: Inject the current URL and page title; Editor scenario: Get the cursor position and selected text.

[0030] IV. Exception Handling and Compatibility Hook conflict avoidance mechanism: When other applications are detected to have registered the same hotkey, the alternative key combination (such as Ctrl+Alt+Shift+A) will be automatically switched. Release hook resources immediately after the workbench is closed to avoid affecting system performance.

[0031] High DPI screen adaptation: The size of the floating window is dynamically adjusted based on the screen scaling ratio (Windows: GetDpiForWindow | macOS:backingScaleFactor).

[0032] Application scenario example: When designing in Photoshop, pressing Ctrl+Alt+A triggers a low-level keyboard hook to capture the key event, disabling Photoshop's native shortcut response. This creates a floating window that overlays the top layer of the Photoshop interface, with a semi-transparent background that allows the design to be seen through. The input box automatically loads the path of the currently edited file (e.g., / design / poster.psd). The user enters the command: "Export the current layer as a PNG and send it to the client." After the workbench is closed, the Photoshop shortcut functionality is immediately restored.

[0033] In this embodiment, the dynamic matching engine performs the following operations: First, based on user commands and context, it searches a historical high-quality task mirror library using a combination of a vector database and a large language model to determine if there are reusable existing similar high-quality task solutions. If so, the historical task solution is directly adopted as the optimal execution solution. If not, the similarity between the subtask and the COT template is calculated using the vector database. The TOP-K candidate solutions are then finely ranked using a re-ranking model. Finally, the COT execution chain with the highest score is output as the final solution. Example 2: Priority retrieval of historical mirror database and execution chain generation mechanism I. Architecture of Historical High-Quality Task Mirror Repository Data storage structure: Each record includes: original instruction, structured subtask sequence, final execution chain, user rating (1-5 stars), and number of successful executions; Additional metadata: creation time, last used time, utility environment tag.

[0034] Quality labeling mechanism: User explicit rating: A rating dialog box pops up after the task is completed; Implicit quality signals: Tasks with a success rate >95% and no human intervention are automatically marked as high quality.

[0035] II. Two-layer retrieval and acquisition process 1. Receive the sequence of subtasks; 2. Calculate the instruction semantic vector; 3. Retrieve historical mirror database: Matching criteria: Vector similarity > threshold θ and quality score > 4 stars; If a matching solution exists, load the execution chain directly and jump to step 6; 4. If no match is found, retrieve the Top-K candidate chains from the COT knowledge base; 5. The re-ranking model calculates the comprehensive score of the candidate chains; 6. Output the final execution chain.

[0036] III. Optimization of Historical Solution Reuse Parameter adaptive adjustment: When reusing historical schemes, automatically update relevant parameters, such as replacing "last week" with "this week"; Tool usability verification: Before execution, check whether the tools registered in the historical scheme are currently available.

[0037] IV. Mirror Repository Self-Learning Mechanism Successfully implemented new solutions are automatically added to the history database, such as those requiring a user rating of ≥3 stars; Periodically clean up inefficient solutions (those with a success rate of less than 80% or that have not been used within the past 3 months).

[0038] Application scenario example: User command: "Generate a quarterly sales report and email it to management"; Dynamic matching engine execution process: 1. Prioritize historical database retrieval Calculate the instruction vector and retrieve the history library; Three high-quality historical solutions were matched (ratings 4.8-5.0). Select the latest solution (created on 2025-06-30, successfully executed 12 times).

[0039] 2. Direct reuse scheme { Source: Historical Mirror Repository #ID-2076 "Execution chain": [ {"Action": "Export CRM Data", "Tool": "CRM Connector", "Parameters": {"Time Range": "This Quarter"}}, {"Action": "Generate Chart", "Tool": "Excel", "Parameters": {"Template": "Sales Quarterly Report Template.xlsx"}}, {"Action": "Send Email", "Tool": "Outlook", "Parameters": {"Inbox Group": "Management"}} ] }”; Automatic parameter update: Replace "last quarter" with "this quarter"; Verify tool availability: Confirm that Outlook is logged in.

[0040] 3. Feedback on execution results The workbench displays: "Reusing historical high-quality solutions (12 successful records), the estimated time will be reduced by 40%"; After the task is completed, the system will prompt: "Update this plan? The current quarter's data has been added to the cache."

[0041] In this embodiment, the cross-tool execution module includes a tool adaptation layer, which converts natural language instructions or model instructions into declarative tool invocation instructions. The instruction format includes action type, target tool, and parameter list. Example 3: Unified Conversion and Execution Mechanism for Multi-Source Instructions I. Dual-path processing of instruction sources Natural language commands: Text commands directly entered by the user (such as "Create sales report"); Model instructions: Structured intermediate instructions output by LLM / MLLM ({"intent": "create_report", "domain": "sales"}); The tool adaptation layer uniformly receives both types of instructions to ensure consistency in the processing flow.

[0042] II. Declarative Instruction Generation Specification Construct a standardized triplet instruction format: { "Action Type": "create_report", / / Predefined action enumeration value "Target Tool": "WPS", / / Registration Tool Identifier "Parameter list": { / / Tool-related parameters Template: Sales Template.docx "Data source": "CRM system", Output format: PDF } }”; The predefined action type library is extended to multimodal scenarios: Image processing classes: crop_image / recognize_text / convert_format; Audio processing classes: transcribe_audio / add_subtitle.

[0043] III. Examples of Multimodal Command Conversion Input multimodal commands: The user drags an image to the workbench and says, "Extract the data from this chart to Excel"; Conversion process: 1. MLLM parsing of images to generate intermediate instructions: "{"action": "extract_data", "target": "excel", "data_type": "chart_table"}"; 2. The tool adaptation layer is converted into declarative directives: { "Action type": "extract_data", Target tool: Excel "Parameter list": { Image path: " / temp / chart.png", "Output Workbook": "Extract Data.xlsx" } }".

[0044] IV. Cross-tool parameter mapping engine Dynamically resolve parameter dependencies: The parameter "image path" was detected to originate from the output of the previous step. Automatically inject the actual generated path " / temp / chart_20250814.png"; Tool-specific parameter conversion: General parameters "Output format" → WPS tool mapping is: FileFormat=pdf The Python tool is mapped as: format='pdf'.

[0045] Application scenario example: Scenario 1: Natural Language Instruction Conversion 1. User input: "Summarize customer feedback into a Word document"; 2. LLM Model Generation Instructions: "{"action":"aggregate_data","tool":"word","topic": "customer_feedback"}"; 3. The tool adaptation layer outputs declarative instructions: { Action type: "create_document", Target Tool: WPS "Parameter list": { Title: Customer Feedback Summary Report "Data source": "CRM system", "Includes chart": true } }".

[0046] Scenario 2: Multimodal instruction conversion 1. Users upload screenshots of sales data and enter: "Generate Excel in this format"; 2. Output model commands after MLLM parsing and screenshotting: { "action": "replicate_table", "target": "excel", "columns": ["products", "sales volume", "growth rate"], "data_type": "numerical" }” 3. Tool adaptation layer conversion results: { "Action type": "create_table", Target tool: Excel "Parameter list": { "Column Name": ["Product", "Sales Volume", "Growth Rate"], Data type: {"Sales volume": "number", "Growth rate": "percentage"} } }".

[0047] In this embodiment, a dynamic path adjustment module is also included. The dynamic path adjustment module monitors the execution status in real time during the execution of subtasks. When an execution deviation is detected, the large language model or multimodal large model is triggered to replan the remaining subtasks. Example 4: Execution Deviation Detection and Intelligent Reconfiguration Mechanism Based on Multimodal Perception I. Multi-dimensional execution status monitoring Traditional signal monitoring: Continue monitoring tool return codes, output verification, and resource status; Multimodal signal acquisition: Visual signals: Abnormal state of the screenshot recognition tool interface (such as error pop-ups, progress bar lag); Text signals: Log / console information output by real-time parsing tools; Audio signal: Recognize system warning sounds or voice prompts.

[0048] II. Bias Decision-Making Driven by Multimodal Large Models When a potential anomaly is detected, information from multiple sources is aggregated to form a decision context: { "return_code": 0, "screenshot": "base64_encoded_image", "log_output": "Error: File in use by another process", "resource_usage": {"cpu": 85%, "memory": "1.2GB"} }”; Input a large multimodal model for joint judgment and output decision suggestions: { "need_replan": true, "reason": "A file usage error pop-up was detected and CPU usage is abnormal." "suggested_action": "Release the file lock and retry" }".

[0049] III. Context-Aware Reconstruction Based on LLM / MLLM Refactoring request prompts enhances multimodal context: Original command: "[User command]" Completed steps: [Summary of steps 1 to n] Current failed step: [Description of step n+1] Available resources: [List of generated files / data] Multimodal context: - Screen Status: [Describe the key elements in the screenshot] - Error message: [Error details extracted from logs] Please replan your next steps. Multimodal large models generate new execution chains that adapt to the actual environment.

[0050] IV. Multimodal debugging information recording Store enhanced failure cases in the COT knowledge base: { "Failure Step": "Saving WPS File", Error type: "File lock occupation", Visual Evidence: Screenshot of the error pop-up window Solution: Retry after a 500ms delay. }".

[0051] Application scenario example: Scenario 1: Visual Perception Triggers Reconstruction 1. Execution scenario: Steps: Use WPS to save the document; Traditional monitoring: Return code = 0 (success).

[0052] 2. Anomalies detected by multimodal monitoring: The screenshot detected a "Disk space insufficient" pop-up. After analyzing the multimodal large model, it was determined that a new planning was needed.

[0053] 3. LLM generates new solutions Original plan: [Save document → Upload to cloud drive] New plan: [Free up temporary space → Save documents → Upload to cloud drive → Clear cache]".

[0054] Scenario 2: Joint Decision-Making Based on Multiple Signals 1. Execution scenario: The 3D rendering tool timed out, but the return code is still 0 (running). 2. Multimodal monitoring aggregation: Resource status: CPU utilization remained at 95% for 10 minutes. Log analysis: A "Light and Shadow Calculation Blocking" warning was detected; Visual verification: The progress bar is stuck at 78% with no change.

[0055] 3. Restructuring Decisions: The multimodal large model is determined to be in an infinite loop, so the task is terminated and a new solution is generated: Original plan: High-quality rendering (estimated 2 hours) New solution: Reduce image rendering quality (estimated 20 minutes)".

[0056] In this embodiment, the system returns the task execution status in real time through a Streamable long connection channel and displays the progress visually within the AI ​​workbench; Example 5: Real-time Synchronization and Visualization System for Task Execution Status I. Establishment of Long Connection Channel Create a persistent WebSocket connection between the Agent execution module and the AI ​​workbench; State data is encapsulated using a binary protocol (Protobuf format) to reduce transmission overhead; Maintain two-way heartbeat detection, automatically retry and resume transmission of the last status when connection is lost.

[0057] II. Status Data Encapsulation Specification Define a unified state message structure: "{"task_id": "T-202508141203", / / Unique identifier for the task "step_seq": 3, / / Current step number "status_code":"EXECUTING", / / Status codes (PENDING / EXECUTING / SUCCESS / FAILED) "progress": 70, / / Current step progress percentage "artifact": "chart.png", / / Resource identifier has been generated "message": "Installing chart..." / / Readable status description}.

[0058] III. Workbench Visualization Engine Multi-dimensional state rendering: Circular progress bar: Displays the current step's completion status (green: 0-70%, yellow: 70-90%, red: 90-100% risk of timeout); Step flowchart: Highlight the step node that is being executed, gray out the steps that have not started, and mark the failed steps with a red box; Resource Generation Panel: Dynamically displays created files / data objects (such as the sales data.csv icon).

[0059] Enhanced display of abnormal states: The error details for failed steps will be automatically expanded (including tool return codes and repair suggestions); Flash the associated file icon when there is a resource conflict.

[0060] IV. Historical Tracing Function Persistently store end-to-end status messages; Supports replaying the task execution process (accelerating playback / pausing / jumping to key nodes); Export the status log for debugging and analysis.

[0061] Application scenario example: User command: "Merge three Excel files to generate a quarterly report".

[0062] Visualization process: I. Task Initiation (e.g.) Figure 4 (As shown).

[0063] II. Execution Process Update When step 1 is completed, the channel pushes: "{"step_seq":1, "status_code":"SUCCESS", "artifact":"fileA.xlsx"}"; Workbench Update: Step 1: Node turns green (√); Step 2: The progress bar begins to fill (yellow dynamically increasing). The resource panel now includes fileB.xlsx.

[0064] III. Abnormal Handling Step 3 failed push notification: "{"step_seq":3, "status_code":"FAILED", "message":"File C is in use (Error code 5)", "suggestion":"Please close WPS or skip this file"}"; Workbench Response: Step 3: The red frame around the contact point flashes. A repair options window will pop up: "[Retry][Skip file C][Manual intervention]".

[0065] IV. Task Completion Final state: “[1]✓[2]✓[3✓[4]✓”; The resource panel displays the actionable results: Quarterly summary report.xlsx (with preview button).

[0066] In this embodiment, the cross-tool execution module supports the invocation of tools including but not limited to code editors, office software, version control systems, command-line terminals, and multimodal processing tools; Example 6: Integration and Scheduling Mechanism of Multimodal Processing Tools I. Multimodal Tool Registration Extension Add a multimodal tool type and metadata to the global tool registry: Image processing tools: Photoshop (COM interface), GIMP (command line), online AI image creation (REST API); Audio processing tools: Audacity (automation scripts), FFmpeg (command line), speech recognition service (WebSocket stream); Video processing tools: Premiere (COM interface), OpenCV (Python library), CapCut (automation script); Capability description extended to multimodal operations: { Tool Type: Image Processing Supported operations: ["Portrait cutout", "Style transfer", "Resolution enhancement", "Text recognition"], Input format: ["jpg","png","bmp"], Output format: ["jpg","png","pdf"] }".

[0067] II. Multimodal Data Transmission Pipeline Establish a temporary storage area for binary data: Image / audio / video files are cached to %TEMP% / multimedia / Generate a unique hash identifier (such as file_sha256) for each file; Cross-tool data flow: Step 1: Use a speech recognition tool to generate subtitle text → Write it to an SRT file Step 2: The video processing tool reads the SRT file → hardcodes it into the video stream.

[0068] III. Examples of Multimodal Instruction Adaptation Image processing scenarios: User command: "Change the background of this photo to a snow-capped mountain"; Tool adaptation layer conversion: { "Action type": "replace_background", Target tool: Photoshop "Parameter list": { "Input image": " / temp / photo_123.jpg", Background template: "Snow Mountain" Output path: " / temp / photo_123_snow mountain.jpg" } }".

[0069] Audio processing scenarios: User command: "Extract the first 30 seconds of this recording and convert it to text"; Tool adaptation layer conversion: { "Action type": "speech_to_text", "Target Tool": "Whisper" "Parameter list": { "Audio path": " / temp / audio_456.wav", Time range: 0-30s Output format: "txt" } }".

[0070] Application scenario example: Multimodal hybrid task execution User command: "Record a 10-second screen video, add the subtitle 'Demonstration Process', and save it as a GIF"; Execution process: 1. Tool scheduling sequence: "[1] Screen recording tool: capture screen for 10 seconds → output video.mp4; " [2] Subtitle generation tool: Creating SRT files (content: demonstration process); [3] Video processing tools: Combine video and subtitles → Output video_with_subtitle.mp4; [4] Format conversion tool: MP4 to GIF → output demo.gif.

[0071] 2. Cross-modal data transfer: Step 1 outputs video.mp4 → Step 3 inputs; Step 2 outputs subtitle.srt → Step 3 inputs; Step 3 outputs video_with_subtitle.mp4 → Step 4 inputs.

[0072] 3. Multi-tool collaborative support: Automatic downgrade when professional video tools are not installed: "Premiere unavailable → Automatically switch to FFmpeg (command line version)"; Output format adaptive: Automatically select the best parameters based on the capabilities of the target tool (such as GIF color palette optimization).

[0073] In this embodiment, a human-machine collaboration module is also included, which forcibly triggers a confirmation window when performing critical sub-tasks involving file operations, data deletion, or external system access, pausing the task and waiting for the user's explicit decision before continuing execution; Example 7: Security Interception and Mandatory Confirmation Mechanism for Critical Operation Categories I. Refinement of Mandatory Departure Rules Three categories of operations that must be subject to mandatory confirmation are clearly defined: File operations: deleting files (including emptying the recycle bin), overwriting system files, and modifying execute permissions; Data deletion: Clear database tables, erase disk partitions, and batch delete user data; External system access: calling payment interfaces, accessing the production environment, sending emails containing sensitive data; The system scans the execution chain steps in real time, and immediately pauses the process and pops up a forced confirmation window when a rule is matched.

[0074] II. Hierarchical Security Verification Mechanism

[0075] III. Context-Aware Confirmation of Content Analysis of the impact of dynamic pop-up display on operations: "[High-risk operation confirmed]" Execution to be performed: Clear the user log table (user_logs) Scope of impact: • Deleted 1,402,356 records • Impact on services: behavioral analysis, audit trail • Freed up space: 2.1GB Verification method: Enter the access password + dynamic verification code [Cancel] [Enter voucher______]”; Smart prefill safety parameters: Automatically associate with a security whitelist when accessing from outside; Recommended backup path when deleting data.

[0076] IV. Operational Blocking and Recovery Mechanisms After the forced confirmation window pops up: Immediately suspend the currently executing thread; Any automated bypass operations are prohibited.

[0077] After the user makes a decision: Select "Continue": Execution will resume after verifying credentials; Select "Cancel": Record the "Manual Termination" status and clean up intermediate resources; No response timeout (30 seconds): Automatically cancel the operation and notify the user.

[0078] Application scenario example: Scenario 1: Production Environment Database Cleanup 1. Execution chain steps: "[1] Connect to the production database; [2] Execute: TRUNCATE TABLE user_logs → Trigger forced confirmation (data deletion class)”.

[0079] 2. System Response: Immediately pause the task process; A red security confirmation window pops up: "[High-risk operation confirmed!]" Operation: Clear the production database table Table name: user_logs (1,402,356 records) Impact on services: Order inquiry / User behavior analysis Verification requirements: Password + Dynamic verification code [Cancel] [Enter password______][Enter verification code______] .

[0080] 3. Security Management: If the password is entered incorrectly three times consecutively, the task will be locked and a security alert will be triggered. An audit log will be generated upon successful verification. "Operator: Zhang San | Operation Time: 2025-08-14 15:30 | Verification Method: Two-Factor Authentication".

[0081] Scenario 2: External Transfer of Customer Data 1. Execution chain steps: "[1] Packed Customer Data.zip" [2] Call the email API to send to partner@external.com → trigger forced confirmation (external access class)”.

[0082] 2. System Response: The task is paused and a yellow confirmation window pops up: "[! External data transmission confirmed!]" Receiving domain: external.com (not on the whitelist) The attachment contains sensitive fields: mobile phone number / ID card number Security Recommendation: Send to internal security email address first. [Cancel] [Change to secure email] [Persist in sending] .

[0083] 3. Risk avoidance: Select "Change to secure email": It will automatically replace with secxxx@coxxxxx.com; Select "Insist on Sending": You need to enter the reason for risk confirmation and record the audit log.

[0084] In this embodiment, a real-time human-machine adjustment module is also included, which visually displays the operation rules and intermediate execution states to the user during task execution and provides a real-time dynamic adjustment interface. The user can intervene at any time to adjust the task parameters or process. The AI ​​system receives the adjustment instructions in real time and responds to the execution. This process does not forcibly pause the task process. Example 8: Real-time adjustment of interface and dynamic response mechanism I. Non-intrusive adjustment interface design Side-mounted floating control panel: The key parameter controls (slider, drop-down menu, color selector) for the current execution step are continuously displayed on the right side of the workbench. Visual focus guidance: Highlight the actual interface element corresponding to the parameter being adjusted (e.g., when adjusting font size, the corresponding text area flashes a blue box); Real-time preview function: When parameters are adjusted, the effect preview is immediately displayed in the workbench (such as animated demonstrations of color changes and size adjustments).

[0085] II. Dynamically Adjusting Data Transmission Pipelines Establish a low-latency, two-way communication channel: "[User Adjustment of Parameters] → [Workbench UI Update] → [Real-time Command Push] → [Tool Execution Layer]" [Tool Execution Status] ← [Status Feedback] ← [Execution Result] ←”; Incremental parameter update protocol: { "operation": "update_parameter", "step_id": "step_3", "changes": { "font_size": "14px → 16px", / / Only transmits the change "color": "#FF0000 → #00FF00"} }".

[0086] III. Continuity Guarantee During the Adjustment Process Background tasks are not interrupted: When adjusting parameters, the core execution thread continues to run, only injecting the new parameters into subsequent operations; Version control: Each time parameters are adjusted, a version snapshot is generated (e.g., v1.0 → v1.1), and rollback is supported at any time; Conflict resolution mechanisms: When a conflict is detected between user-adjusted parameters and automated execution (such as modifying the same parameter at the same time), the user-input value will be used first.

[0087] Application scenario example: Real-time adjustments during PPT automatic layout process 1. Initial state: The system automatically executes: "Apply the corporate template to these slides"; The side panel displays adjustable parameters: Font size: [14px ━━━━━━━●━━━━━ 18px] Theme colors: [Dropdown menu: Blue / Red / Green] Animation speed: [Slow ━━● ━━━━━ Fast]

[0088] 2. Real-time user intervention: Drag the font size slider to 16px; System response: Instantly update the text style of the current slide (no lag); Record parameter changes: "Font size: 14px → 16px"; Choose "red" as the theme color; System response: Gradual color scheme transition; All subsequent automatically generated slides will use the new color scheme.

[0089] 3. Continuous Implementation Guarantee: During the adjustment process: The system continues to process the remaining slide layout; New parameters are injected into subsequent operation processes in real time.

[0090] Final output: A complete set of slides based on user-adjusted parameters (16px font + red theme). Please see Figure 3In this embodiment, the Agent-based targeted invocation and intelligent matching method based on AI large model technology includes the following steps: S1. Receiving user-triggered raw commands or multimodal commands triggered by a global hotkey; S2. Decomposing the commands into a structured sub-task sequence using a large language model or multimodal large model; S3. Prioritizing the search of a historical high-quality task image library for matching; if no match is found, then searching the COT knowledge base to match each sub-task and generate a complete execution chain; S4. Calling multiple tool APIs to execute operations according to the execution chain sequence; S5. Pushing the execution status to the workbench interface in real time through a long connection channel; S6. Triggering mandatory confirmation or accepting real-time adjustment commands from the user at key nodes according to the strategy. Example 9: Full-process processing method for multimodal instructions I. Multimodal Command Reception and Parsing Input source extension: Text commands: Natural language descriptions entered by the user via keyboard; Image commands: Drag and paste images or screenshots (including OCR text extraction); Voice commands: Speech-to-text input via microphone (integrated real-time ASR); Composite command: Inputting a combination of text and images (e.g., screenshot + text "Process this table").

[0091] Unified resolution mechanism: Multimodal Large Model (MLLM) processes all input sources simultaneously to generate structured instruction representations: { "primary_intent": "process_table", "source_type": ["image", "text"], "content": { "image": " / temp / screenshot_20250814.png", "text": "Export the data here to Excel" } }".

[0092] II. Dual-database search priority process Historical mirror databases are prioritized for matching: 1. Calculate the similarity between the instruction semantic vector and the history database; 2. If a high-quality matching solution exists (similarity > θ and rating > 4 stars), load and execute the chain directly; 3. Mark the source of the task: "Reuse of high-quality historical solutions".

[0093] COT database downgrade retrieval: The COT retrieval process is initiated only when no match is found in the historical database: 1. Vector retrieval of Top-K candidate chains; 2. Overall score of the re-ranking model; 3. Output the optimal execution chain and mark it as "New solution generated".

[0094] III. Multi-tool execution and state synchronization Cross-modal tool scheduling: Example: Image table processing workflow [1] Use an OCR tool to extract table data from the image → generate data.csv; [2] Call Excel to format data → generate formatted.xlsx; [3] Call the email client to send the result → attach formatted file.

[0095] Enhanced real-time status push notifications: Status messages now include multimodal execution information: { "step_seq": 2, "tool_type": "image_processing", / / Add tool type identifier "media_artifact": "formatted.xlsx", / / Identifier for multimedia output "preview_url": "thumbnail.png" / / Visual preview link }".

[0096] Application scenario example: Composite instruction processing flow User actions: 1. Capture the webpage table and drag it to the workbench; 2. Enter the voice command: "Generate this month's sales report in this format".

[0097] Method execution process: S1. Command Reception: Input set: {Screenshot image + Speech-to-text "Generate this month's sales report in this format"} S2 instruction breakdown: MLLM parsing output: { Task Type: Data Report Generation "Reference Template": " / temp / screenshot_20250814.png", Time range: "This month", Output format: Excel }”; S3 Dual-Database Search: 1. Prioritize searching the historical database: Matched the "Sales Report Generation" solution (4.8 stars); 2. Directly reuse the execution chain: "[1] Connect to the CRM system to extract data for this month" [2] Use Excel application template format (template ID: sales_template_v3) [3] Save to the specified path and notify the user.

[0098] S4 Multi-Tool Execution: Step 1: Call the CRM API (HTTP tool); Step 2: Start the Excel COM component (office software tool); Step 3: Call the system notification interface (operating system tool).

[0099] S5 state synchronization (e.g.) Figure 5 (As shown).

[0100] S6 Intelligent Intervention: Automatically insert confirmation points in the "Call CRM API" step (due to the involvement of production data access).

[0101] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0102] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural transformations made using the contents of the present invention's specification and drawings under the inventive concept of the present invention, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.

Claims

1. An agent-oriented recall and intelligent matching system based on AI large model technology, characterized in that, include: A global hotkey monitoring module is used to invoke a unified AI platform from any application interface via system-level shortcut keys; The multimodal natural language parsing module connects the large language model and the multimodal large model, breaking down user input commands or multimodal commands into structured subtasks; The COT knowledge base stores predefined Chain of Thought execution chain templates; A historical high-quality task mirror library stores high-quality task execution schemes and their COT execution chains, jointly annotated by users and AI. The dynamic matching engine matches the COT knowledge base through vector semantic retrieval and outputs the optimal execution plan through a re-ranking model; The cross-tool execution module automatically calls multiple independent tools in the operating system based on the execution plan.

2. The Agent-based targeted recall and intelligent matching system based on AI large model technology according to claim 1, characterized in that, The global hotkey listening module is implemented through the hook function at the bottom layer of the operating system, and the AI ​​workbench that is invoked is overlaid on the existing application interface in the form of a floating window.

3. The Agent-based targeted recall and intelligent matching system based on AI large model technology according to claim 1, characterized in that, The dynamic matching engine performs the following operations: First, based on user commands and context, it uses a combination of vector database and large language model to search the historical high-quality task mirror library and determine whether there are reusable existing similar high-quality task solutions. If such a historical task solution exists, then that solution will be used as the optimal execution solution. If it does not exist, the similarity between the subtask and the COT template is calculated using a vector database; The TOP-K candidate solutions are finely ranked using a re-ranking model; The COT execution chain with the highest output score is selected as the final solution.

4. The Agent-based targeted recall and intelligent matching system based on AI large model technology according to claim 1, characterized in that, The cross-tool execution module includes a tool adaptation layer, which converts natural language instructions or model instructions into declarative tool call instructions. The instruction format includes action type, target tool, and parameter list.

5. The Agent-based targeted recall and intelligent matching system based on AI large model technology according to claim 1, characterized in that, It also includes a dynamic path adjustment module, which monitors the execution status in real time during the execution of subtasks. When an execution deviation is detected, it triggers the large language model or multimodal large model to replan the remaining subtasks.

6. The Agent-based targeted recall and intelligent matching system based on AI large model technology according to claim 1, characterized in that, The system returns the task execution status in real time through a Streamable long connection channel and displays the progress visually within the AI ​​workbench.

7. The Agent-based targeted recall and intelligent matching system based on AI large model technology according to claim 1, characterized in that, The cross-tool execution module supports the use of tools including but not limited to code editors, office software, version control systems, command-line terminals, and multimodal processing tools.

8. The Agent-based targeted recall and intelligent matching system based on AI large model technology according to claim 1, characterized in that, It also includes a human-machine collaboration module, which forcibly triggers a confirmation window when performing critical sub-tasks involving file operations, data deletion, or access to external systems, pausing the task and waiting for the user's explicit decision before it can continue.

9. The Agent-based targeted recall and intelligent matching system based on AI large model technology according to claim 1, characterized in that, It also includes a real-time human-machine adjustment module, which visually displays the operation rules and intermediate execution status to the user during task execution and provides a real-time dynamic adjustment interface. Users can intervene at any time to adjust task parameters or processes. The AI ​​system receives adjustment instructions in real time and responds to execute them. This process does not forcibly pause the task process.

10. A method for targeted agent activation and intelligent matching based on AI large model technology, applied to the agent targeted activation and intelligent matching system based on AI large model technology as described in any one of claims 1 to 9, characterized in that, Includes the following steps: S1. Receive user-generated commands or multimodal commands triggered by global hotkeys; S2. Decompose instructions into a sequence of structured subtasks using a large language model or a multimodal large model; S3. Prioritize searching the historical high-quality task image library for matching. If no match is found, then search the COT knowledge base to match each sub-task and generate a complete execution chain. S4. Call multiple utility APIs in the execution chain order to perform operations; S5. Push the execution status to the workbench interface in real time via a long connection channel; S6. At critical nodes, trigger mandatory confirmation based on the strategy or accept real-time adjustment instructions from users.