Intelligent task processing method and system based on external state file driving

By adopting an intelligent task processing method driven by external state files, the problems of illusion and inconsistent state management in large language models in complex tasks are solved, and the reliable, auditable and accurate execution of tasks is achieved, thereby improving the stability and reliability of the system.

CN120973500AActive Publication Date: 2025-11-18WELLCOME (SHENZHEN) INTELLIGENT CO LTD

Patent Information

Application Number
CN202511492287.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

In existing technologies, large language models are prone to hallucinations when dealing with complex tasks, have inconsistent state management, are difficult to audit and review end-to-end, and cannot effectively grasp the overall picture of the task or make precise interventions.

Method used

An intelligent task processing method based on external state files is adopted. It creates an external state file by receiving task requests, decomposes tasks using a large language model, reads the state file to generate action instructions in each loop, executes and updates the state file in conjunction with the tool interface, supports human-machine collaborative intervention, generates structured output, and forms an auditable task file after the task is terminated.

Benefits of technology

It ensures the accuracy and reliability of output, avoids inconsistencies in state management, improves system stability and reliability, provides full auditability and traceability, and ensures complete record and reliable output of task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973500A_ABST
    Figure CN120973500A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent task processing method and system based on external state file driving. The method comprises the steps that an external state file is created and initialized based on a task request of a user; performing task decomposition through the large language model to generate at least one to-be-executed sub-task, and updating the sub-task to an external state file; entering a task execution loop, and in each loop, reading the current content of the external state file, and generating an action instruction of the next step through a large language model; updating an observation result after the tool interface executes the action instruction and a corresponding subtask state to an external state file; when it is determined that the task termination condition is satisfied, a structured output is generated based on the final content of the external state file. According to the invention, by physically separating the fact knowledge from the reasoning process, the possibility that the LLM generates output which does not accord with the fact is eliminated, and the accuracy and reliability of system output are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to an intelligent task processing method and system driven by external state files. Background Technology

[0002] With the rapid development of Large Language Modeling (LLM) technology, its application in automated systems is becoming increasingly widespread. Current technologies directly apply LLM to handle complex tasks, such as in-depth research or emergency response decision-making; however, LLM can exhibit "illusion" phenomena when dealing with tasks requiring precise facts, outputting content that does not conform to reality. When handling long-term, multi-step tasks, it is prone to forgetting key objectives or state information due to context length limitations. Existing technologies construct multi-agent systems that iteratively decompose tasks through a master agent, typically storing the task's execution state implicitly and unstructuredly in the master agent's dialogue history or internal memory. This state management approach has drawbacks: lack of a human-machine collaboration interface, inability to grasp the overall task picture, and inability to precisely intervene in task planning; inconsistencies or loss of state management; and difficulty in end-to-end auditing and review. Summary of the Invention

[0003] The purpose of this invention is to provide an intelligent task processing method and system based on external state files, so as to solve the problems mentioned in the background art, such as the tendency to have "illusion" phenomena, inconsistent state management, and insufficient auditability.

[0004] To achieve the above objectives, the present invention adopts the following technical solution: According to one aspect of the present invention, an intelligent task processing method based on an external state file is provided, the method comprising: Receive a task request input by the user, and create and initialize an external state file based on the task request; Based on the task request and the current content of the external state file, the task is decomposed using a large language model to generate at least one subtask to be executed, and the subtask is updated in the external state file. Enter the task execution loop. In each loop: read the current content of the external state file and generate the next action instruction through the large language model; execute the action instruction through the tool interface; update the observation results and corresponding subtask states after executing the action instruction to the external state file. When the task termination condition is met, structured output is generated based on the final content of the external status file.

[0005] Based on the aforementioned scheme, the task execution loop also includes a human-machine collaborative intervention step: Receive natural language intervention instructions from external input; Based on the intervention instruction and the current content of the external state file, modification suggestions for the external state file are generated through the large language model; Update the external state file according to the proposed modifications to modify the task execution flow.

[0006] Based on the aforementioned scheme, executing the action instruction through the tool interface includes calling a deterministic external tool function and obtaining a structured observation object, wherein the observation object contains at least one or more of the following: execution status, master data, and error information.

[0007] Based on the aforementioned solution, the generation of structured output, when the structured output is a presentation, includes: The content recorded in the external state file is transformed into a structured design intent specification by generating an intelligent agent based on the design intent. Based on the design intent specifications, a layout template is matched from the metadata template library by a layout matching agent; Based on the matched layout template and content, a series of deterministic editing instructions are generated by the intelligent agent through editing instructions; The instruction execution agent executes the editing instructions to generate the presentation.

[0008] Based on the aforementioned scheme, if an error occurs when the instruction execution agent is executing an editing instruction, it will send the error information back to the editing instruction generation agent to correct the instruction and re-execute it.

[0009] Based on the aforementioned scheme, after the task is terminated, the method further includes packaging the final version of the external status file, the final structured output, and related metadata together to form an auditable task file and storing it.

[0010] Based on the aforementioned scheme, the task decomposition using a large language model includes: Construct a prompt word, which includes at least system instructions, the current contents of the external status file, and a list of available tools; The large language model performs inference based on the prompt words and outputs structured task decomposition results; The task decomposition results are updated to the external status file as new subtasks, where each subtask is assigned a status identifier.

[0011] Based on the aforementioned scheme, the action instructions executed through the tool interface in the task execution loop include at least one of the following: The fusion knowledge base query tool is invoked to obtain domain knowledge and historical data from the fusion knowledge base; Call the real-time data acquisition tool to obtain dynamically changing real-time data streams through the real-time data bus.

[0012] According to another aspect of the present invention, an intelligent task processing system based on an external state file is provided. The system includes: a state file management module, a planning engine module, a context builder, a tool interface module, a fusion knowledge base, and a real-time data bus. The status file management module is used to create, read, and update external status files; The planning engine module is used to run large language models for task decomposition and reasoning; The context builder is used to construct structured input prompts for the planning engine module, the core content of which originates from the external state file; The tool interface module is used to execute action commands output by the planning engine module; The fusion knowledge base is used to store domain-specific structured and unstructured knowledge; The real-time data bus is used to access and provide dynamically changing real-time data streams.

[0013] Based on the aforementioned scheme, a multi-agent document generation module is also included, comprising an agent for generating design intent, an agent for matching layout, an agent for generating editing instructions, an agent for executing instructions, and an agent for compiling. The design intent generation agent is used to generate structured design intent specifications; The layout matching agent is used to match layouts in the metadata template library; The editing instruction generating agent is used to generate deterministic editing instructions; The instruction execution agent is used to execute editing instructions to generate a single-page document file; The compiler agent is used to merge the single-page document files and output the final document.

[0014] As can be seen from the above technical solution, this invention has at least the following advantages and positive effects compared with the prior art: By physically separating factual knowledge from the reasoning process, it forces the LLM to obtain factual data by querying the fusion knowledge base, completely eliminating the possibility of the LLM producing outputs that are inconsistent with the facts, and ensuring the accuracy and reliability of the system output. By using the external state file as a unique, complete, and persistent state representation throughout the entire task lifecycle, it avoids the complex state synchronization problem of distributed systems, greatly improving system stability and reliability. The external state file is a complete, chronologically recorded, and readable archive, in which the entire task's thinking and execution chain is fully presented, providing auditability. In the document generation scenario, by executing deterministic editing instructions to modify the template, it ensures that the preset formatting technical parameters in the template are consistently reproduced in the output file.

[0015] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings: Figure 1 This is a schematic diagram of an intelligent task processing method based on an external state file driven by the present invention. Figure 2 This is a flowchart of an intelligent task processing method based on an external state file driven by the present invention; Figure 3 This is a schematic diagram of the task execution loop of the present invention; Figure 4 This is a schematic diagram of an intelligent task processing system based on an external state file driven by the present invention. Detailed Implementation

[0017] To more clearly illustrate the purpose, technical solutions, and advantages of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein. On the contrary, these embodiments are provided so that the present invention will be more comprehensive and complete, and fully convey the concept of the exemplary embodiments to those skilled in the art.

[0018] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a full understanding of embodiments of the invention. However, those skilled in the art will recognize that the technical solutions of the invention can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of the invention.

[0019] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0020] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0021] The present invention will now be described in detail with reference to specific embodiments: Example 1

[0022] like Figure 1 , 2 As shown, this embodiment provides an intelligent task processing method based on an external state file. The specific steps of this method are as follows: S1: Receive a task request input by the user, and create and initialize an external state file based on the task request.

[0023] In this embodiment, the task state of the large language model is managed based on an external state file, and structured output is generated. The external state file, also known as a dynamic plan file or DPF file, evolves dynamically as the task is executed. Initially, it may only contain a simple task description, but as the large language model's tasks are decomposed and the system is executed, it will gradually grow into a complete plan containing detailed sub-tasks, execution results, current state, and subsequent plans.

[0024] After receiving a task request, the system creates and initializes an independent external state file for the task (for example, it can be named DPF_<taskID>.md). This file is stored in a persistent storage device, making it an external state, thereby achieving the separation of state from computation.

[0025] The initial content of this external status file follows a predefined structure. For example, its initial content includes at least: a file header containing metadata such as a unique task identifier, creation timestamp, and initial status; and a task list, whose initial status includes a pending task item generated based on a user request, which can be labeled as: [P] <task description>.

[0026] Specifically, the system first receives task requests via a user interface (such as a chat box or button trigger) or an automated monitoring system (such as sensor alarms). This task request can be a question described in natural language (e.g., "Analyze the flood risk of the G4 Expressway K150 section") or a structured instruction (e.g., {"task_type": "generate_ppt", "topic":"Work plan for the second half of the year"}). Further, an external state file is created; a separate, persistent text file, such as DPF_task_001.md, is created and initialized for this task. In all subsequent steps, the system's decisions, execution, and updates are strictly based on the latest content of the external state file, ensuring a strong consistency in the understanding of the task state across all components. The created external status file includes at least: task description, initial status, and metadata; the task description records the user's original request; the initial status, for example, is a top-level task to be solved, marked as: [P] Task: Understand and decompose the user query: "Deduce the main risk points of G4 Expressway K150 section under 3 hours of 150mm rainfall." ([P] stands for Pending / Pending); the metadata includes the task's unique identifier, creation time, status (such as Status: ACTIVE), etc.

[0027] S2: Based on the task request and the current content of the external state file, the task is decomposed using a large language model to generate at least one subtask to be executed, and the subtask is updated in the external state file.

[0028] The external state file (DPF) containing the initial task description created in step S1 is transformed into a structured, step-by-step action plan; Specifically, the process begins by reading the current content of the external state file. At this point, the external state file typically contains only one top-level task marked as [P] (Pending). Further, a context is constructed, packaging the current DPF content, system-preset instructions (such as "You are a planning engine, please decompose the task"), and a list of available tools into a structured and clear prompt. The list of available tools includes tools for querying the fusion knowledge base (such as `query_knowledge_base`) and tools for obtaining real-time data (such as `get_realtime_data`). This prompt is then sent to the large language model, which outputs a structured task decomposition result. The large language model's task is not to directly answer the initial question, but to analyze the task intent and decompose it into a series of discrete subtasks that can be executed sequentially or in parallel. Finally, the output of the large language model (i.e., the task decomposition result) is received as new subtasks and updated in the external state file. These newly generated subtasks are also marked as [P], while the original top-level task may be marked as [D] (Completed) or [Decomposed], etc.

[0029] When a task item with a status of 'Pending' ([P]) and related to 'planning decomposition' exists in the external state file, the task decomposition process is automatically triggered. The input to the large language model includes reading the entire current content of the external state file as the core context; embedding an instruction that defines the role and output format of the large language model, such as: 'Your task is to decompose the top-level task in the current external state file into specific, executable sub-steps; the output must be only a list of sub-tasks'; and then attaching a list of available utility functions and their descriptions, such as query_asset_database, run_simulation, generate_report, etc., to guide the large language model in planning within its known capabilities.

[0030] It should be noted that the large language model performs context-based reasoning, analyzes the final goal of the current task, and decomposes it into a series of sequential or parallel atomic subtasks. Each subtask should, as far as possible, match an available tool function, forming a clear action such as 'calling tool X and inputting parameter Y'. The output of the large language model strictly adheres to a preset format, for example, returning a JSON array where each array element represents a subtask, containing fields such as description and expected_tool. After receiving the structured decomposition suggestions from the large language model, the original top-level task in the external state file with a state of [P] is marked as completed ([D]), and optionally 'decomposed' is recorded. All subtask items generated by the large language model are appended to the task list in the external state file with a state of [P]. The updated external state file clearly shows the evolution from 'goal' to 'path', and all newly added subtasks become the execution objects of subsequent control loops.

[0031] For example, an initial task: "[P] Task: Deduce the risk of G4 Expressway K150 section under 3 hours of 150mm rainfall.", after inference by the large language model, the external state file is updated to: [D] Task: Understand and break down user queries... [P]Task: Query asset information for section G4 K150 from the knowledge base.

[0032] [P]Task: Call the hydrological simulation tool and input the rainfall parameters.

[0033] [P]Task: Based on the simulation results, retrieve the relevant SOP clauses.

[0034] [P]Task: Generate a comprehensive risk assessment report. By breaking down a fuzzy and complex task into specific subtasks, this embodiment transforms nondeterministic large language model reasoning into a list of deterministic steps; subsequent steps only require performing these explicit actions according to the list, which greatly improves the reliability and predictability of the system.

[0035] Preferably, if the output of the large language model does not follow a predetermined format or contains unrecognized tool calls, the tool interface will capture this exception and mark the status of the decomposition task as failed ([F]) in the external status file, while recording error details in the observation results; this can be designed to trigger a retry or notify manual intervention, thereby preventing erroneous plans from flowing into the execution phase.

[0036] S3: Enter the task execution loop. In each loop: read the current content of the external state file and generate the next action instruction through the large language model; execute the action instruction through the tool interface; update the observation results and corresponding subtask states after executing the action instruction to the external state file.

[0037] like Figure 3 As shown, the triggering condition for the task execution loop is: there is at least one subtask marked as "to be executed" ([P]) in the external state file; each loop cycle begins with an access to the external state file's state storage, reading all the current contents of the external state file to obtain the current complete state of the task execution, ensuring that every decision is based on the latest, global context.

[0038] In this embodiment, after the original task is decomposed into a series of subtasks, the external state file is continuously scanned to locate subtasks in the pending state. Subtasks marked [P] (pending) that require data acquisition are selected from the external state file. A prompt word is constructed for the large language model based on this specific subtask, such as "Currently, we need to query asset information for section K150; please generate a precise query instruction based on available tools." The prompt word template includes: system role instructions, such as, "You are a planning engine, and your responsibility is to suggest the best tool to be called next based on the current external state file status"; the current external state file status, i.e., the entire content of the external state file; a list of available tools, including all callable deterministic tools and their parameter descriptions; and an output format instruction, requiring the output of the large language model to be in a predefined structured format.

[0039] The large language model outputs structured, machine-readable action instructions based on the prompt word. The tool interface receives and parses the instructions generated by the large language model, and verifies the validity of the tool name and parameters of the external tool. It maps the tool name to the corresponding deterministic tool function and executes it (such as executing a specific database SQL query or calling an API to obtain sensor data). After the tool function is executed, the tool interface fully captures the execution result and encapsulates the result (successfully obtained data, empty data, or an exception) into a structured observation object. The fields of the observation object include, but are not limited to, execution status, primary data, warning information, and error information. Execution status (execution_status) is such as "SUCCESS", "PARTIAL_SUCCESS", "FAILED", primary data (primary_data) is the main data result returned by the tool call, warning information (warnings) is an array of additional information that needs attention, and error information (errors) is an array of error information that caused the operation to fail.

[0040] When the action command executed by the tool interface is to query the fusion knowledge base, the system retrieves structured information such as domain knowledge, historical data, and device parameters from the fusion knowledge base by calling a deterministic query tool. When the action command is to obtain real-time data, the tool interface accesses the real-time data stream through the real-time data bus to obtain dynamic information such as sensor readings, device status, and environmental monitoring data.

[0041] Furthermore, a deterministic update operation is performed on the observed object: the corresponding sub-task item executed in the current cycle is found in the external state file, and the status identifier of the task item is updated from pending execution ([P]) to completed ([D]) or failed ([F]). At the end of the description of the task item, a concise summary or complete content of the observed object is appended or embedded in the format of (observation result: ...) as an additional comment. This update operation indicates the end of the current cycle, and the updated external state file becomes the input for the next cycle, starting a new task execution loop. This complete recording method provides indispensable factual basis for the large language model to make complex decisions (such as fault handling and path replanning) in the next cycle, and is auditable.

[0042] It should be noted that the large language model transforms the natural language descriptions in the external state file into structured, deterministic calling instructions that the underlying tools can understand. The large language model itself does not hold or output factual answers, but only outputs "action requests". This achieves a physical separation between the reasoning process and the source of facts, thereby isolating the propagation path of "illusions".

[0043] Throughout the task execution process, a loop mechanism continuously retrieves background knowledge and domain specifications from the fusion knowledge base; retrieves the latest status information from the real-time data bus; updates the external status file with the retrieved data as observation results; and makes the next round of decisions based on the updated status.

[0044] Optionally, this embodiment also provides human-machine collaborative intervention, which can be triggered in parallel and integrated into the task execution loop. The operator inputs a natural language command, i.e., an external intervention command, through a user interface (such as a chat window or command panel). This command can be a query for the current task, a modification of the execution process, or the injection of new knowledge. For example, when monitoring flood control emergency response, if the operator finds that the system's planned deployment of drones poses a risk, they can input: "Suspend the deployment of drone UAV_02, prioritize notifying the road administration department to close the K150-K152 entrance, and add a task to check the status of the backup power supply." Upon receiving the external intervention command, the original external status file will not be interrupted or reset. Instead, it will be treated as a new, high-priority input and incorporated into the next task execution loop for processing. Based on the complete content of the current external status file and the newly received external intervention command, a corresponding command is constructed. For example, "An operator has given the following command. Please understand its intent and, based on the current external status file status, generate specific modification suggestions for the external status file to execute the command." A list of available tools is also provided. The large language model performs reasoning based on this context and outputs modification suggestions for the external status file. The tool interface parses the output and treats it as a deterministic operation to directly modify the external status file (such as pausing a task, inserting a new task, or modifying the task priority).

[0045] Through human-machine collaborative intervention, operators can achieve precise control using natural language; all interventions are transformed into explicit modifications to external status files, avoiding ambiguity and recording the entire operation; the operator's external interventions are seamlessly integrated into the existing automated processes and status management framework, maintaining the consistency of the system status; every human-machine interaction, including intervention commands, the system's understanding of them (generated modification suggestions), and the final external status file changes, is fully recorded in the historical versions of the external status file or the system log, providing clear audit clues for subsequent process optimization and review.

[0046] S4: When the task termination condition is met, generate structured output based on the final content of the external status file.

[0047] The task termination conditions include: no subtasks in the external status file are in the 'pending execution' state; a task termination command is received from an external input; and all subtasks in the external status file are in the 'complete' or 'failed' state. After the task execution loop completes all data acquisition and logical reasoning subtasks, it enters the structured output generation step. The input to this step is an external status file that records the complete task execution history and data observation results, and the output is a structured document that conforms to preset specifications. The corresponding generation process is initiated according to the output type; in this embodiment, there are two paths based on the output type: generating a text report and outputting a presentation document.

[0048] When an external state file contains a task item for generating a text report (such as a risk assessment report), the observation results and key data of all completed tasks are extracted from the external state file. The content of the external state file and the report generation instructions (such as "Please write a structured risk assessment report based on the following task records and data") are packaged and sent to the large language model. The structured report draft generated by the large language model is formatted by a deterministic template (such as an XML or Markdown template) and finally rendered into a document in the specified format (such as PDF or DOCX) to complete the output.

[0049] When an external state file contains a task item for generating a presentation, the multi-agent collaborative pipeline is activated, which includes, in sequence: design intent generation agent, layout matching agent, editing instruction generation agent, instruction execution agent, and compilation agent; The design intent generation agent reads the parts of the external state file related to the presentation outline and content, and through natural language processing, transforms the unstructured text content of each page into a machine-readable structured design intent specification. The design intent specification at least defines the content element types and design constraints of the page; and transforms the communication intent into layout technical parameters. The layout matching agent accepts the design intent specification and searches and matches it in the metadata template library based on its content. The metadata template library stores the technical parameters of each layout template, such as the shape ID, precise coordinates, font style, and color RGB value of the placeholder. The layout matching agent outputs a unique identifier for the layout template with the highest compatibility with the design intent specification. The editing instruction generation agent receives the specific page content and the unique identifier of the layout template. Based on the metadata of the layout template, it generates a series of deterministic, atomic editing instructions. These editing instructions precisely specify the identifier of the element to be operated on and the operation content. The instruction execution agent executes the sequence of editing instructions one by one, modifying the template by calling the underlying document operation interface. An error correction mechanism is also included: when the instruction execution agent encounters an exception during execution, it generates feedback information containing the error type and error identifier, and returns this feedback information to the instruction generation agent. The instruction generation agent corrects the editing instructions based on the feedback information and resubmits them to the instruction execution agent for execution until the entire instruction sequence is successfully executed, generating a single-page presentation file. The compiler merges all successfully generated single-page presentation files in sequence, applies a unified theme style, and outputs the final presentation document.

[0050] Furthermore, this embodiment also includes task archiving and auditing; when all subtask statuses in the external status file are determined to be "complete" ([D]) or "failed" ([F]) or an external termination instruction is received, the task termination condition is met; the task status field (Status) in the file header metadata of the external status file is updated from "active" to "closed" or "terminated"; the final version of the external status file, the final output result, and the task metadata index are combined to form a complete task archive and stored in a persistent archive database. The final version of the external status file records all state evolution records, action instruction sequences, and observation results from task initialization to completion; the final output result is the structured output generated in step S5, such as a text report or presentation file; the metadata index includes the task's unique identifier, creation timestamp, end timestamp, key performance indicators, etc.

[0051] The task archive, consisting of the final external status file, final output results, and task metadata index, provides a complete and tamper-proof execution chain from task request to final output, making any decision or output traceable to its original data source and decision context.

[0052] Example 2

[0053] like Figure 4 As shown in the figure, this embodiment exemplifies an intelligent task processing system driven by an external state file, including a state file management module, a context builder, a planning engine module, a tool interface module, a fusion knowledge base, and a real-time data bus.

[0054] The status file management module is used to create, read, and update external status files. It creates external status files based on task requests. The external status file is a structured text file (such as Markdown format) that records the complete lifecycle of the task, including the initial goal, the decomposed subtasks, the execution status of each task (such as [P] pending execution, [D] completed), the observation results of tool execution, and records of human intervention. It is the system's only source of facts.

[0055] The context builder is used to construct structured input prompts for the planning engine module. The core content of the input prompts comes from external state files. At the beginning of each work cycle, an unambiguous and structured prompt is constructed from external state files, system instructions, and a list of available tools based on a template, which greatly improves the stability and relevance of the output of the large language model.

[0056] The planning engine module is used to run large language models for task decomposition and reasoning. It does not interact directly with users or external tools, but only outputs suggestions for the next action (such as tool call instructions) based on the context provided by the context builder in each work cycle.

[0057] The tool interface module is used to execute action instructions output by the planning engine module; it is a deterministic code module responsible for parsing and executing action instructions output by the large language model (such as calling APIs or querying databases).

[0058] The integrated knowledge base is used to store domain-specific structured and unstructured knowledge; static and authoritative factual knowledge within the storage domain, such as infrastructure design drawings, operating procedures (SOPs), and equipment parameters; the tool interface module provides factual evidence to the large language model by querying the integrated knowledge base to eliminate illusions.

[0059] A real-time data bus is used to access and provide dynamically changing real-time data streams, such as sensor readings, video streams, and meteorological data, enabling the system to perceive instantaneous changes in the real world.

[0060] In another implementation, a multi-agent document generation module is also included, comprising an agent for generating design intent, an agent for matching layout, an agent for generating editing instructions, an agent for executing instructions, and an agent for compiling. The design intent generation agent is used to generate structured design intent specifications; the layout matching agent is used to match layouts in the metadata template library according to the design intent specifications; the editing instruction generation agent is used to generate deterministic editing instructions; the instruction execution agent is used to execute editing instructions to generate single-page manuscript files and achieve closed-loop error correction; and the compilation agent is used to merge single-page manuscripts and output the final manuscript document.

[0061] Optionally, this embodiment also includes a human-machine collaboration interface. The system provides a natural language interaction interface, allowing operators to view external status files in real time to grasp the overall situation and to input natural language commands for intervention. The system understands the intent of the commands through a large language model and transforms them into deterministic modifications to the external status files, thereby achieving precise and secure human-machine collaboration.

[0062] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims. It should be understood that the invention is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A smart task processing method based on external state files, characterized in that, The method includes: Receive a task request input by the user, and create and initialize an external state file based on the task request; Based on the task request and the current content of the external state file, the task is decomposed using a large language model to generate at least one subtask to be executed, and the subtask is updated in the external state file. Enter the task execution loop. In each loop: read the current content of the external state file and generate the next action instruction through the large language model; execute the action instruction through the tool interface; update the observation results and corresponding subtask states after executing the action instruction to the external state file. When the task termination condition is met, structured output is generated based on the final content of the external status file.

2. The intelligent task processing method based on external state file driving according to claim 1, characterized in that, The task execution loop also includes a human-machine collaborative intervention step: Receive natural language intervention instructions from external input; Based on the intervention instruction and the current content of the external state file, modification suggestions for the external state file are generated through the large language model; Update the external status file according to the proposed modifications.

3. The intelligent task processing method based on external state file driving according to claim 1, characterized in that, The execution of the action instruction through the tool interface includes calling a deterministic external tool function and obtaining a structured observation object, which contains at least one or more of the following: execution status, master data, and error information.

4. The intelligent task processing method based on external state file driving according to claim 1, characterized in that, The generation of structured output, when the structured output is a presentation, includes: The content recorded in the external state file is transformed into a structured design intent specification by generating an intelligent agent based on the design intent. Based on the design intent specifications, a layout template is matched from the metadata template library by a layout matching agent; Based on the matched layout template and content, a series of deterministic editing instructions are generated by the intelligent agent through editing instructions; The instruction execution agent executes the editing instructions to generate the presentation.

5. The intelligent task processing method based on external state file driving according to claim 4, characterized in that, If an error occurs while the instruction execution agent is executing an editing instruction, it will send the error information back to the editing instruction generation agent to correct the instruction and re-execute it.

6. The intelligent task processing method based on external state file driving according to claim 1, characterized in that, After the task is terminated, the final content of the external status file, the structured output, and related metadata are packaged together to form an auditable task file and stored.

7. The intelligent task processing method based on external state file driving according to claim 1, characterized in that, The task decomposition using a large language model includes: Construct a prompt word, which includes at least system instructions, the current contents of the external status file, and a list of available tools; The large language model performs inference based on the prompt words and outputs structured task decomposition results; The task decomposition results are updated to the external status file as new subtasks, where each subtask is assigned a status identifier.

8. The intelligent task processing method based on external state file driving according to claim 1, characterized in that, In the task execution loop, the action instructions executed through the tool interface include at least one of the following: The fusion knowledge base query tool is invoked to obtain domain knowledge and historical data from the fusion knowledge base; Call the real-time data acquisition tool to obtain dynamically changing real-time data streams through the real-time data bus.

9. An intelligent task processing system driven by an external state file, characterized in that, It includes a status file management module, a context builder, a planning engine module, a tool interface module, a fusion knowledge base, and a real-time data bus; The status file management module is used to create, read, and update external status files; The context builder is used to construct structured input prompts for the planning engine module, the core content of which originates from the external state file; The planning engine module is used to run a large language model for task decomposition and reasoning; The tool interface module is used to execute action commands output by the planning engine module; The fusion knowledge base is used to store domain-specific structured and unstructured knowledge; The real-time data bus is used to access and provide dynamically changing real-time data streams.

10. The intelligent task processing system based on external state file driving according to claim 9, characterized in that, It also includes a multi-agent document generation module, including an agent for generating design intent, an agent for matching layout, an agent for generating editing instructions, an agent for executing instructions, and an agent for compiling; The design intent generation agent is used to generate structured design intent specifications; The layout matching agent is used to match layouts in the metadata template library; The editing instruction generating agent is used to generate deterministic editing instructions; The instruction execution agent is used to execute editing instructions to generate a single-page document file; The compiler agent is used to merge the single-page document files and output the final document.

Citation Information

Patent Citations

  • Complex task decomposition and dynamic optimization method and device based on large language model

    CN119883549A

  • Intelligent agent-based big language model retrieval enhancement generation system and method

    CN120470088A

  • Task disassembly and multi-agent arrangement execution system and method based on large language model

    CN120560815A

  • Generating and maintaining composite actions utilizing large language models

    US20250111149A1

  • Task-oriented assistant using language models

    US20250265422A1

Cited By

  • Task decomposition method and terminal

    CN121525886A

  • Method and device for adding context of large language model

    CN122114186A