Intelligent agent dynamic decision network generation method based on reinforcement learning
Through the dynamic decision network generation method of agents based on reinforcement learning, the challenges of enterprise-level customers when integrating artificial intelligence technology are solved, efficient generation and dynamic adjustment of agent workflow are achieved, and the work efficiency and accuracy of agents are improved.
Patent Information
- Application Number
- CN202510679404.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-26
AI Technical Summary
The existing technology is difficult to effectively solve the challenges of enterprise-level customers when integrating artificial intelligence technology into business processes, including the balance between model inference efficiency and accuracy, video memory or memory bottlenecks, and business-level adaptation issues.
Adopt the dynamic decision network generation method of agents based on reinforcement learning, and feature analysis is performed by receiving user input information demand data, extracting business goals, constraints and key parameters, generating agent workflows, and building a hierarchical reinforcement learning framework to optimize the nodes and paths of agent workflows.
It realizes efficient generation and dynamic adjustment of the agent's workflow, improves the work efficiency and accuracy of the agent when handling complex tasks, and ensures that the agent can respond efficiently and flexibly to various business scenarios.
Smart Images

Figure CN120197644A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method for generating an intelligent agent dynamic decision-making network based on reinforcement learning. Background Art
[0002] With the continuous innovation of large artificial intelligence model technology, generative artificial intelligence is gradually being applied in various fields. However, despite the growing demand for artificial intelligence technology, there is a relative lack of intelligent agent automated development tools that are both easy to use and user-friendly for enterprise-level customers. When most enterprise developers attempt to integrate artificial intelligence technology into business processes, they often encounter a series of challenges. Taking model inference as an example, developers face many specific technical difficulties. For example, in scenarios such as chatbots and intelligent assistants, the model often requires low latency and waiting time; the model is too large, resulting in video memory or memory bottlenecks; it is difficult to optimize between the inference efficiency of large models and model accuracy; on the other hand, existing technologies also involve business-level adaptation issues, specifically reflected in how to seamlessly connect intelligent technology with existing business processes and ensure that the technology application meets the specific needs and compliance requirements of the enterprise. Summary of the Invention
[0003] Based on this, it is necessary to provide a method for generating an intelligent agent dynamic decision-making network based on reinforcement learning to solve at least one of the above technical problems.
[0004] To achieve the above object, a method for generating an intelligent agent dynamic decision-making network based on reinforcement learning, the method includes the following steps: Step S1: Receive the information requirement data input by the user; perform feature parsing on the information requirement data, and extract the business objective, constraint conditions, and key parameters to obtain the information requirement parsing result; identify the task process corresponding to the information requirement parsing result, and match the API call chain according to the task process to obtain the task requirement technical blueprint; Step S2: Generate an intelligent agent workflow according to the task requirement technical blueprint and using a preset dynamic workflow engine; perform context analysis on the language requirement parsing result to obtain context information; divide the intelligent agent workflow into an ultra-long thought chain based on the context information, and determine the nodes and paths of the intelligent agent workflow according to the ultra-long thought chain; Step S3: Collect the intelligent agent business requirement data; dynamically adjust the nodes and paths in the intelligent agent workflow according to the intelligent agent business requirement data; Step S4: Construct a hierarchical reinforcement learning framework for the agent, including a strategic layer, a tactical layer, and an execution layer; in the strategic layer, use a reinforcement learning algorithm to predict the long-term business goal achievement path of the agent's business demand data and determine the quarterly-level optimization strategy; in the tactical layer, use the proximal policy optimization algorithm to adjust the agent's workflow node configuration and resource allocation plan on a weekly basis; in the execution layer, apply the deep Q-network algorithm to optimize the single-task execution strategy of the agent's workflow in real time; Step S5: During the process of the agent executing tasks, construct and generate a dynamic decision-making network for the agent based on the hierarchical reinforcement learning framework of the agent.
[0005] Preferably, in step S1, perform feature analysis on the information demand data, and extract business goals, constraints, and key parameters, including: Extract features from the information demand data, and convert the extracted demand features into an information demand vector set; Input the information demand vector set into a preset large model for intention recognition, and output the intention recognition matching result; Divide the intention recognition matching result into business goals, constraints, and key parameters, and record their positions in the information demand data to obtain the information demand analysis result.
[0006] Preferably, in step S1, identify the task process corresponding to the information demand analysis result, and match the API call chain according to the task process, including: Decompose the business goal in the information demand analysis result into multiple subtasks, where each subtask corresponds to a business operation information; Perform feature recognition on each subtask to determine its sequence and dependency relationship in the task process; According to the sequence and dependency relationship of the subtasks, construct a task flow diagram, where the task flow diagram includes task nodes and task edges, the task nodes represent specific subtasks, and the task edges represent the sequence and dependency relationship between subtasks; For each task node in the task flow diagram, find the matching API interface type; According to the task edges in the task flow diagram, determine the call sequence between API interfaces to form an API call chain; Verify the API call chain to detect whether the input and output of each API interface conform to the task process sequence; Bind the verified API call chain to the task flow diagram to form a complete mapping relationship between the task process and the API call chain, so as to obtain the technical blueprint of the task requirements.
[0007] Preferably, in step S2, generate the agent's workflow according to the technical blueprint of the task requirements and using a preset dynamic workflow engine, including: Extract the start node and end node of the task from the task requirement technical blueprint; Identify the key task nodes in the technical blueprint. Each key task node corresponds to a specific business operation decision point, and record the name and function of each key task node; Assign a unique identifier to each key task node and record its position coordinates in the technical blueprint; According to the order and dependency of the key task nodes, draw a preliminary framework diagram of the workflow; the framework diagram includes nodes and connection lines, the connection lines represent the execution order between nodes, and arrows are used to indicate the direction; Configure execution parameters for each node in the preliminary framework diagram according to the technical blueprint. The execution parameters include execution time, resource requirements, and priority; In the preset dynamic workflow engine, instantiate the configured nodes and connection lines to generate an executable intelligent agent workflow.
[0008] Preferably, in step S2, the intelligent agent workflow is divided into ultra-long thought chains based on context information. Determining the nodes and paths of the intelligent agent workflow includes: Extract the context information of each task node from the intelligent agent workflow, including input parameters, output parameters, and execution environment of the task node; Encode the context information of each task node and convert the input parameters, output parameters, and execution environment into vector form; Calculate the cosine similarity between vectors to determine the similarity between adjacent task nodes; Group the task nodes according to the similarity between adjacent task nodes, and group the task nodes with similarity higher than the preset similarity threshold into the same thought unit; For each thought unit, determine the execution order of the task nodes according to the dependency and priority of the task nodes; Connect all thought units in series according to the execution order of the task nodes to form an ultra-long thought chain; Optimize the ultra-long thought chain by adjusting the order and path of the task nodes to ensure the coherence and efficiency of the thought chain; Determine the nodes and paths of the intelligent agent workflow according to the task nodes and path connection methods in the ultra-long thought chain.
[0009] Preferably, step S3 includes the following steps: Step S31: Through a preset data acquisition interface, collect intelligent agent business requirement data from multiple data sources. The data sources include user input, sensor data, and historical business records; Step S32: Preprocess the intelligent agent business requirement data, including data cleaning, format conversion, and data annotation, to obtain standard intelligent agent business requirement data; Step S33: Input the standard intelligent agent business requirement data into the dynamic adjustment module, and use the dynamic adjustment module to identify the nodes and paths that need to be adjusted in the standard intelligent agent business requirement data; Step S34: Dynamically adjust the nodes and paths that need to be adjusted through the dynamic adjustment module for task priority, resource requirements, and execution time, and generate business requirement adjustment instructions; Step S35: Reconfigure the nodes in the intelligent agent workflow according to the business requirement adjustment instructions, including adding, deleting, or modifying the attributes of the nodes; Step S36: Re-plan the paths in the intelligent agent workflow according to the business requirement adjustment instructions, including adjusting the path order and optimizing the path connection method.
[0010] Preferably, in step S4 at the strategic level, the reinforcement learning algorithm is adopted to predict the long-term business goal achievement path of the intelligent agent business requirement data and determine the quarterly-level optimization strategy, including the following steps: At the strategic level, measure the task completion rate of the intelligent agent business requirement data and record it as the strategic business indicator; Input the strategic business indicator into the reinforcement learning algorithm, and initialize the environmental state and the initial strategy of the intelligent agent; Through the interaction between the simulation environment and the intelligent agent, collect the business execution results of the intelligent agent under different strategies, and record the reward value of each strategy; According to the business execution results and the reward values of each strategy, use the policy iteration method in the reinforcement learning algorithm to update the strategy of the intelligent agent; Based on the updated strategy, predict the long-term business goal achievement path of the intelligent agent business requirement data and generate a quarterly-level optimization strategy.
[0011] Preferably, in step S4 at the tactical level, the proximal policy optimization algorithm is used to adjust the intelligent agent workflow node configuration and resource allocation plan in the weekly dimension, including: At the tactical level, extract the current node configuration and resource allocation situation from the intelligent agent workflow; among them, the current node configuration includes the type, quantity, and connection relationship of the nodes, and the resource allocation situation includes the resource type, quantity, and allocation ratio of each node; Input the current node configuration and resource allocation situation into the proximal policy optimization algorithm, initialize the policy parameters, and set the running environment of the algorithm, including the initial state of the simulation environment and the initial strategy of the intelligent agent; On a weekly basis, evaluate the resource allocation of each node, record the resource utilization efficiency and task completion status of each node to generate a resource allocation evaluation result; where the resource utilization efficiency includes the type, quantity, and usage time of the resources, and the task completion status includes the type, quantity, and completion time of the tasks. Adjust the resource allocation of each node using the proximal policy optimization algorithm according to the resource allocation evaluation result, optimize the resource utilization efficiency, and update the policy parameters.
[0012] Preferably, in the execution layer of step S4, applying the deep Q-network algorithm to optimize the single-task execution policy of the agent workflow in real time includes: In the execution layer, extract the current single-task execution policy from the agent workflow, including the input parameters, output parameters, and execution environment of the task; Input the extracted single-task execution policy into the deep Q-network algorithm to initialize the weights and biases of the Q-network; In the execution layer, monitor the execution process of each task in real time and record the execution status and results of the task; According to the execution status and results of the task, calculate the Q-value of each task using the deep Q-network algorithm to evaluate the pros and cons of the task execution policy; According to the calculated Q-value, adjust the execution policy of the task to optimize the execution efficiency and results of the task; Verify the adjusted execution policy to ensure that it meets the business requirements and actual operating conditions; According to the verification results, fine-tune the execution policy to further optimize the execution efficiency and results of the task; Apply the fine-tuned execution policy to the agent workflow and update the single-task execution policy of the workflow.
[0013] Preferably, step S5 includes the following steps: Step S51: During the process of the agent executing tasks, monitor the input and output data of the agent in real time and record the intermediate state of task execution; Step S52: Based on the hierarchical reinforcement learning framework, decompose the task execution process into multiple subtasks, and each subtask corresponds to a specific decision point; Step S53: Assign an independent reinforcement learning module to each subtask, and each module is responsible for optimizing the decision-making strategy of the subtask; Step S54: At each decision point, use the reinforcement learning module to calculate the optimal action in the current state and generate a decision instruction; Step S55: Send the generated decision instruction to the agent to guide the next action of the agent; Step S56: According to the execution results of the agent, update the parameters of the reinforcement learning module and optimize the decision-making strategy; Step S57: Integrate the decision-making strategies of all subtasks to construct an intelligent agent dynamic decision-making network.
[0014] The unique technical effects of the present invention are as follows: By receiving the information requirement data input by the user and performing feature parsing, it is possible to accurately extract the business objectives, constraints, and key parameters, thereby obtaining an accurate information requirement parsing result. This enables the system to clarify the core requirements and limiting conditions of the user, providing a solid foundation for the subsequent identification of the task process and the matching of the API call chain. The task requirement technical blueprint obtained by matching according to the task process corresponding to the information requirement parsing result can provide clear and accurate technical architecture guidance for the generation of the agent workflow, ensuring that the generation of the agent workflow highly conforms to the user's requirements and laying a foundation for the subsequent efficient operation of the agent. Using the preset dynamic workflow engine to generate the agent workflow according to the task requirement technical blueprint can quickly construct a workflow framework that adapts to the task requirements based on the established technical blueprint. Conducting context analysis on the language requirement parsing result to obtain context information, and then dividing the agent workflow into ultra-long thinking chains based on this, and determining the nodes and paths of the agent workflow. This process can fully consider the context relevance of the language requirements, making the structure of the agent workflow more reasonable and the logic more coherent, effectively improving the work efficiency and accuracy of the agent when dealing with complex tasks. Collecting the agent business requirement data and dynamically adjusting the nodes and paths in the agent workflow accordingly realizes the dynamic adaptability of the agent workflow. This process can optimize and adjust the workflow in a timely manner according to the changes in the actual business requirements, ensuring that the agent workflow always remains consistent with the business requirements, thereby improving the flexibility and adaptability of the agent in the face of different business scenarios and enhancing the business processing ability of the agent. The constructed agent hierarchical reinforcement learning framework, including the strategic layer, tactical layer, and execution layer, can optimize the agent's workflow from different dimensions. At the strategic layer, a reinforcement learning algorithm is used to predict the long-term business objective achievement path of the agent business requirement data and determine the quarterly-level optimization strategy, enabling the agent to grasp the long-term direction of business development from a macro perspective, plan the optimal path in advance, and ensure that the agent always maintains an efficient and stable development trend during the long-term business development process. At the tactical layer, the proximal policy optimization algorithm is used to adjust the agent workflow node configuration and resource allocation plan on a weekly basis, which can timely adjust the node configuration and resource allocation of the workflow according to the changes in short-term business requirements, improving the resource utilization efficiency and task execution effect of the agent during the short-term business execution process. At the execution layer, the deep Q-network algorithm is applied to optimize the single-task execution strategy of the agent workflow in real time, enabling the agent to quickly and accurately make the optimal decision when executing specific tasks, and improving the execution quality and efficiency of single tasks. During the process of the agent executing tasks, an agent dynamic decision-making network is constructed and generated based on the agent hierarchical reinforcement learning framework, enabling the agent to make decisions in real time according to the hierarchical reinforcement learning framework during the task execution process.The construction of this dynamic decision-making network enables the agent to make decisions that adapt to the current task requirements quickly and accurately based on real-time business requirement data and workflow status, further enhancing the agent's decision-making ability and task execution efficiency in complex business environments and ensuring that the agent can complete various tasks efficiently and stably. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic flow chart of the steps of a method for generating a dynamic decision-making network of an agent based on reinforcement learning; Figure 2 is Figure 1 a detailed schematic flow chart of step S3 in Figure 3 is Figure 1 a detailed schematic flow chart of step S5 in The realization, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0017] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0018] It should be understood that although the terms "first", "second", etc. may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.
[0019] To achieve the above object, please refer to Figures 1 to 3, a method for generating an intelligent agent dynamic decision network based on reinforcement learning, the method comprising the following steps: Step S1: Receive the information requirement data input by the user; perform feature parsing on the information requirement data, and extract the business objective, constraint conditions, and key parameters to obtain an information requirement parsing result; identify the task process corresponding to the information requirement parsing result, and match the API call chain according to the task process to obtain a technical blueprint of the task requirement; Step S2: Generate an intelligent agent workflow according to the technical blueprint of the task requirement and using a preset dynamic workflow engine; perform context analysis on the language requirement parsing result to obtain context information; divide the intelligent agent workflow into an ultra-long thought chain based on the context information, and determine the nodes and paths of the intelligent agent workflow according to the ultra-long thought chain; Step S3: Collect the intelligent agent business requirement data; dynamically adjust the nodes and paths in the intelligent agent workflow according to the intelligent agent business requirement data; Step S4: Construct a hierarchical reinforcement learning framework for the intelligent agent, including a strategic layer, a tactical layer, and an execution layer; in the strategic layer, use a reinforcement learning algorithm to predict the long-term business objective achievement path of the intelligent agent business requirement data and determine the quarterly-level optimization strategy; in the tactical layer, use the proximal policy optimization algorithm to adjust the intelligent agent workflow node configuration and resource allocation plan on a weekly basis; in the execution layer, apply the deep Q-network algorithm to optimize the single-task execution strategy of the intelligent agent workflow in real time; Step S5: During the process of the intelligent agent executing tasks, construct and generate an intelligent agent dynamic decision network based on the hierarchical reinforcement learning framework of the intelligent agent.
[0020] In the embodiment of the present invention, with reference to Figure 1 shown in the figure, it is a schematic diagram of the step flow of a method for generating an intelligent agent dynamic decision network based on reinforcement learning according to the present invention. In this example, the method for generating an intelligent agent dynamic decision network based on reinforcement learning includes the following steps: Step S1: Receive the information requirement data input by the user; perform feature parsing on the information requirement data, and extract the business objective, constraint conditions, and key parameters to obtain an information requirement parsing result; identify the task process corresponding to the information requirement parsing result, and match the API call chain according to the task process to obtain a technical blueprint of the task requirement; In the embodiments of the present invention, the system receives the information requirement data input by the user through the natural language interface and completely records it in the system log. The BERT model is used to perform intent recognition and entity extraction on the natural language text input by the user. The BERT model encodes the input text through its pre-trained weight parameters and extracts the word embedding vectors in the text. The system sets a threshold parameter. When the vector similarity between the word embedding vector and the vectors of predefined business objectives, constraints, and key parameters exceeds this threshold (for example, 0.8), it is considered that the corresponding features are successfully extracted. Based on the generation ability of GPT, the business objectives, constraints, and key parameters obtained from the semantic parsing layer are transformed into a technical blueprint including task processes, API call chains, and evaluation metrics. The GPT model generates a detailed technical specification text according to the input structured information, combined with its pre-trained parameters and generation logic. The system sets a keyword matching threshold, such as 80% keyword matching degree. When the keyword matching degree between the keywords in the technical specification and the keywords in the task process template exceeds this threshold, it is considered that the matching is successful. The system further analyzes the logical relationship of the successfully matched keywords through a logical reasoning algorithm to ensure that the generated task process conforms to the business logic. According to each link in the task process, the predefined API interface library is called to generate a specific API call chain; according to the key parameters in the technical specification, the parameters of each API interface in the API call chain are configured. The system stores the task process and the API call chain in a structured manner in the database to form a task requirement technical blueprint. The format of the technical blueprint is JSON format, including each step of the task process, detailed information of the API call chain, and evaluation metrics, etc. The system stores the generated technical blueprint in the task blueprint library for subsequent use by the agent construction and optimization module. The storage path of the technical blueprint is " / task_blueprint / {task_id}.json", where {task_id} is the unique identifier of the task.
[0021] Step S2: Generate an agent workflow according to the task requirement technical blueprint and using a preset dynamic workflow engine; perform context analysis on the language requirement parsing result to obtain context information; divide the agent workflow into an ultra-long thought chain based on the context information, and determine the nodes and paths of the agent workflow according to the ultra-long thought chain; In the embodiments of the present invention, when the dynamic workflow engine is initialized, it loads the task flow and API call chain information in the task requirement technical blueprint. The engine sets the maximum number of nodes in the workflow to 128 to support the generation of ultra-long thought chains. The engine sequentially converts each step in the task flow and its corresponding API call chain into workflow nodes according to the step sequence in the task flow. Each node includes a node ID, a node type (such as API call, decision point, result output, etc.), and the input and output parameters of the node. The context analysis module extracts semantic information from the natural language text input by the user, and at the same time combines the business objectives, constraints, and key parameters in the task requirement technical blueprint to generate context information. For example, if the information requirement input by the user is "Build a supply chain risk warning system", the context information includes keywords such as "supply chain management", "risk assessment", and "real-time monitoring". The module analyzes the extracted context information to determine the specific position and role of each keyword in the task flow. For example, the keyword "supply chain management" is related to the demand forecasting and supplier negotiation steps, and the keyword "risk assessment" is related to the demand forecasting and logistics scheduling steps. The system divides the agent workflow into multiple sub-chains according to the keywords and logical relationships in the context information. Each sub-chain contains a series of related nodes, forming a logically coherent thought chain. For example, the ultra-long thought chain of the supply chain risk warning system is divided into a "demand forecasting sub-chain" (nodes 1 - node 2) and a "logistics scheduling sub-chain" (nodes 2 - node 3). The system determines the nodes and paths in each sub-chain according to the logical relationships in the context information. For example, in the "demand forecasting sub-chain", the output of node 1 (demand forecasting) is used as the input of node 2 (supplier negotiation), forming a path from node 1 to node 2. In the "logistics scheduling sub-chain", the output of node 2 (supplier negotiation) is used as the input of node 3 (logistics scheduling), forming a path from node 2 to node 3.
[0022] Step S3: Collect agent business requirement data; dynamically adjust the nodes and paths in the agent workflow according to the agent business requirement data; In the embodiments of the present invention, the data acquisition module configures data source access parameters. For example, for business system logs, the log file path is set to " / var / log / business_system.log", and for user interaction records, the database connection parameters are set to "host:port:database:user:password". According to business requirements, the data acquisition frequency is set. For example, for business data with high real-time requirements, the acquisition frequency is set to once every 5 minutes; for log data, the acquisition frequency is set to once every 30 minutes. The acquired data passes through the preprocessing module for data cleaning and formatting. The preprocessing module removes noise and redundant information in the data to ensure data accuracy and consistency. For example, for user interaction records, the preprocessing module removes duplicate records and converts the data into a unified JSON format. The intelligent agent workflow adjustment module parses the acquired business requirement data and extracts key information, such as changes in business objectives and new constraint conditions. For example, if the business requirement data contains the information "add new supplier evaluation indicators", then the change requirement is extracted. According to the parsed business requirement data, the intelligent agent workflow adjustment module dynamically adjusts the nodes in the workflow. For example, if new supplier evaluation indicators need to be added, a new evaluation indicator sub-node is added to the supplier negotiation node (node 2). The adjustment module updates the input and output parameters of the node to ensure the compatibility of the new node with the existing nodes. The intelligent agent workflow adjustment module dynamically adjusts the paths in the workflow according to the business requirement data. For example, if the new business requirement requires adding a supplier credit evaluation step after supplier negotiation, a new supplier credit evaluation node (node 2.1) is inserted between the supplier negotiation node (node 2) and the logistics scheduling node (node 3), and the path is updated to node 2 → node 2.1 → node 3. When adjusting nodes and paths, the intelligent agent workflow adjustment module dynamically configures the parameters of relevant nodes according to the business requirement data. For example, for the newly added supplier credit evaluation node (node 2.1), evaluation indicator parameters such as credit score threshold and evaluation period are configured.
[0023] Step S4: Construct an intelligent agent hierarchical reinforcement learning framework, including a strategic layer, a tactical layer, and an execution layer; in the strategic layer, a reinforcement learning algorithm is used to predict the long-term business goal achievement path of the intelligent agent's business requirement data and determine the quarterly-level optimization strategy; in the tactical layer, the proximal policy optimization algorithm is used to adjust the intelligent agent workflow node configuration and resource allocation plan on a weekly basis; in the execution layer, the deep Q-network algorithm is applied to optimize the single-task execution strategy of the intelligent agent workflow in real time; In the embodiments of the present invention, the system calls the model-based reinforcement learning (MBRL) algorithm module at the strategic layer. This module loads historical business data and the execution results of the agent workflow to construct a world model. The world model describes the transition probability of the agent between different states and the reward value in different states through a state transition probability matrix and a reward function. The system sets the dimension of the state transition probability matrix to n×n, where n is the number of states, and the value range of the reward function is . Using the MBRL algorithm and combining with the Monte Carlo tree search (MCTS) technology, a large number of paths are simulated, and the optimal path is selected as the prediction result of the long-term business goal achievement path. According to the prediction result, the system generates a quarterly-level optimization strategy, and the strategy content includes specific measures such as task allocation, resource allocation, and workflow optimization within a quarter. The system calls the proximal policy optimization (PPO) algorithm module at the tactical layer. This module initializes the policy network, sets the hidden layer size to 256 neurons, the learning rate to 0.001, and the discount factor γ to 0.99. Taking a week as a unit, according to the execution results of the agent workflow and business requirement data, the PPO algorithm calculates the policy gradient, updates the policy network parameters, and dynamically adjusts the workflow node configuration and resource allocation scheme. For example, according to the weekly business data, the resource allocation ratio of each node in the workflow is adjusted to optimize the task execution efficiency. The system adjusts the resource allocation scheme according to the optimization result of the PPO algorithm. For example, according to the priority and resource requirements of the task, the allocation ratio of resources such as CPU and memory is dynamically adjusted to ensure the efficient use of resources. The system calls the deep Q-network (DQN) algorithm module at the execution layer. This module initializes the Q-network, sets the hidden layer size to 128 neurons, the learning rate to 0.0005, and the discount factor γ to 0.95. Through the DQN algorithm, the value function of the state-action pair is learned. The DQN algorithm uses the experience replay technology to store and replay the interaction experience of the agent and optimize the Q-network parameters. For example, according to the execution results of the agent in different states, the weights of the Q-network are updated to optimize the single-task execution strategy. The system adjusts the single-task execution strategy of the agent workflow in real time according to the optimization result of the DQN algorithm. For example, according to the current state of the task, the optimal action is selected to optimize the task execution efficiency and response time.
[0024] Step S5: During the process of the agent executing tasks, a dynamic decision network of the agent is constructed and generated based on the hierarchical reinforcement learning framework of the agent.
[0025] In the embodiments of the present invention, first, real-time data generated by the agent during task execution is obtained from the task execution monitoring module. This data includes key information such as task progress, resource consumption, and environmental feedback. This data is completely recorded and stored in the task execution log database. The table structure of the database is pre-designed, including fields such as task ID, timestamp, progress percentage, resource usage, and environmental feedback type, to ensure the integrity and traceability of the data. Subsequently, the system calls the decision network construction module, which constructs the agent's dynamic decision network based on the three levels of the hierarchical reinforcement learning framework, namely, the strategic level, the tactical level, and the execution level. At the strategic level, the system uses the previously predicted long-term business goal achievement path and quarterly-level optimization strategies, combined with the real-time data of the current task execution, to update the agent's long-term strategy through the policy iteration technique in the reinforcement learning algorithm. Specifically, the system sets the convergence threshold of the policy iteration to 0.01. When the policy change between two consecutive iterations is less than this threshold, the policy is considered to converge, thus obtaining a stable long-term strategy. This strategy will serve as the top-level guidance of the decision network to determine the agent's behavior direction at the macroscopic level. At the tactical level, the system adjusts the agent's short-term strategy using the Proximal Policy Optimization algorithm (PPO) according to the agent's workflow node configuration and resource allocation plan at the weekly dimension, as well as the real-time task execution data. The system sets the parameters of the PPO algorithm, including a learning rate of 0.001, a discount factor γ of 0.99, a batch size of 64, etc. Through these parameters, the PPO algorithm can dynamically adjust the transition strategy between nodes and the resource allocation strategy according to the agent's current workflow node state and resource usage situation, so as to improve the agent's task execution efficiency and resource utilization efficiency in the short term. For example, if the task execution progress of a certain node lags behind, the PPO algorithm will adjust the resource allocation, increase the computing resources of this node, and at the same time optimize the transition strategy between nodes to reduce unnecessary waiting time. At the execution level, the system optimizes the agent's immediate behavior based on the Deep Q-Network (DQN) algorithm. The DQN algorithm module receives real-time data from the task execution monitoring module, including information such as the current state of the task, available actions, and immediate rewards. The system sets the parameters of the DQN algorithm, such as a learning rate of 0.0005, a discount factor γ of 0.95, and the size of the experience replay pool of 10,000, etc. The DQN algorithm continuously updates the Q value by learning the value function of the state-action pair and combining the experience replay technique, so as to select the optimal action for the agent at each time step. For example, in a specific task execution scenario, when the agent faces multiple optional actions, the DQN algorithm will select the action with the highest expected reward according to the current state and the estimated Q value to achieve the efficient completion of the task. Finally, the system integrates the strategies of the strategic level, the tactical level, and the execution level to form a complete agent dynamic decision network.The decision-making network is stored in the decision-making network database in the form of a graph. Each node represents a decision point, the edge represents the transition from one decision point to another, and the weight of the edge represents the probability or priority of the transition. The table structure of the decision-making network database is reasonably designed and can store information such as the topological structure of the decision-making network, node attributes, and edge attributes. During the process of the agent executing tasks, the system queries the decision-making network database in real time and dynamically adjusts the behavior of the agent according to the current state and the guidance of the decision-making network to ensure that the agent can efficiently and flexibly handle various complex situations and achieve the task goals.
[0026] Preferably, in step S1, the information requirement data is subjected to feature analysis, and the business objectives, constraint conditions, and key parameters are extracted, including: Feature extraction is performed on the information requirement data, and the extracted requirement features are converted into an information requirement vector set; The information requirement vector set is input into a preset large model for intention recognition, and an intention recognition matching result is output; The intention recognition matching result is divided into business objectives, constraint conditions, and key parameters, and their positions in the information requirement data are recorded to obtain an information requirement analysis result.
[0027] In an embodiment of the present invention, after the system receives the information requirement data in the form of natural language input by the user, it calls the feature extraction module. This module uses word embedding methods in natural language processing technologies, such as Word2Vec or BERT, to extract features from the information requirement data. Taking BERT as an example, the system takes the information requirement data as input and, through the encoder part of the BERT model, converts each word or phrase into a vector with a fixed dimension. In specific operations, the system sets the parameters of the BERT model, such as the hidden layer dimension is 768 and the maximum sequence length is 512. In this way, each word or phrase in the information requirement data is converted into a 768-dimensional vector, forming an information requirement vector set. At the same time, the system records the position information of each vector in the original information requirement data for subsequent processing. The system inputs the generated information requirement vector set into a preset large intent recognition model. This large model is based on the Transformer architecture and processes the input vector set through multiple layers of self-attention mechanisms and feed-forward neural networks. During the processing, the system sets the parameters of the large model, such as the number of layers is 12, the number of attention heads is 12, and the hidden layer dimension is 768. The large model performs deep learning and pattern recognition on the information requirement vector set according to these parameters and outputs an intent recognition matching result. The output result is presented in the form of a probability distribution, indicating the matching degree between the information requirement data and various predefined intents. For example, if the information requirement data is "Construct a supply chain risk warning system", the large model will output a probability of 0.9 for matching the "system construction" intent and a probability of 0.7 for matching the "risk assessment" intent, etc. After the system receives the intent recognition matching result output by the large model, it calls the intent parsing module. This module divides the intent recognition matching result into business objectives, constraint conditions, and key parameters according to predefined intent classification rules. In specific operations, the system sets an intent classification threshold, such as 0.5, and divides the intent recognition results with a matching probability higher than this threshold into corresponding categories. For business objectives, the system extracts the results matching intents such as "system construction" and "function implementation"; for constraint conditions, it extracts the results matching intents such as "time limit" and "cost control"; for key parameters, it extracts the results matching intents such as "technical requirements" and "performance indicators". At the same time, the system determines the specific positions of these categories in the original information requirement data according to the previously recorded position information. Finally, the system integrates the divided business objectives, constraint conditions, and key parameters and their position information into an information requirement parsing result and stores it in the database in a structured form to provide a basis for subsequent agent construction and optimization.
[0028] Preferably, in step S1, identifying the task process corresponding to the information requirement parsing result, and matching the API call chain according to the task process includes: Decompose the business objectives in the information requirement parsing result into multiple subtasks, where each subtask corresponds to a business operation information; Identify the features of each subtask and determine its sequence and dependencies in the task process; Construct a task flow chart based on the sequence and dependencies of the subtasks. The task flow chart includes task nodes and task edges. The task nodes represent specific subtasks, and the task edges represent the sequence and dependencies between subtasks; For each task node in the task flow chart, find the matching API interface type; Determine the call sequence between API interfaces according to the task edges in the task flow chart to form an API call chain; Verify the API call chain and check whether the input and output of each API interface conform to the task process sequence; Bind the verified API call chain to the task flow chart to form a complete mapping relationship between the task process and the API call chain, thereby obtaining the technical blueprint of the task requirements.
[0029] In an embodiment of the present invention, business objective fields are extracted from the information requirement analysis results, such as "building a supply chain risk warning system". According to the business process specifications in the domain knowledge base, the business objective is decomposed into multiple subtasks. For example, "building a supply chain risk warning system" is decomposed into subtasks such as "demand forecasting", "supplier evaluation", "risk analysis", and "warning generation". The mapping relationship between business objectives and subtasks is predefined in the domain knowledge base, and the system decomposes according to this mapping relationship. The decomposed subtasks are stored in a task decomposition table, and the table structure includes fields such as subtask ID, subtask name, and subtask description. For example, the subtask with subtask ID 1 has a subtask name of "demand forecasting" and a description of "forecast supply chain demand". Keywords are extracted from the description of each subtask, and the TF-IDF algorithm is used to extract keywords. For example, keywords "demand" and "forecast" are extracted for the "demand forecasting" subtask. Using the RAG retrieval technology, semantic analysis is performed on the subtask descriptions, and the semantic similarity between subtasks is calculated. The semantic similarity threshold is set to 0.7. When the semantic similarity between two subtasks exceeds this threshold, they are considered to have a dependency relationship. According to the keyword matching and RAG retrieval results, the dependency relationship between subtasks is determined. For example, "demand forecasting" depends on the completion of "data collection", and "supplier evaluation" depends on the result of "demand forecasting". The dependency relationship is stored in a dependency relationship table, and the table structure includes fields such as subtask ID, dependent subtask ID, and dependency type. A task node is created for each subtask, and a unique node ID is assigned. For example, the node ID of the subtask "data collection" is 1, and the node ID of "demand forecasting" is 2. According to the dependency relationship table, task edges are created to represent the order and dependency relationship between subtasks. For example, a task edge 1→2 is created, indicating that "demand forecasting" is executed after "data collection" is completed. The task flow chart is stored in a task flow database in a graphical form, and the database table structure includes fields such as node ID, node name, edge ID, edge start point, and edge end point. For example, the node table records the node ID and node name, and the edge table records the edge ID, edge start point, and edge end point. According to the subtask name and description, the matching API interface type is searched in the predefined API interface library. The API interface library contains information such as the name, function description, input parameters, and output parameters of the API interface. For example, the API interface matched by the "data collection" subtask is "DataCollector". The matching result is stored in an API matching table, and the table structure includes fields such as subtask ID, API interface name, and API interface description. For example, the API interface name for subtask ID 1 is "DataCollector", and the description is "collect supply chain-related data". According to the task edges in the task flow chart, the call order of the API interfaces is determined. For example, according to the task flow Figure 1→ 2 → 3, generate an API call chain as "DataCollector → DemandPredictor → SupplierEvaluator". Store the API call chain in the API call linked list, and the table structure includes fields such as subtask ID, API interface name, call order, etc. For example, the API interface name for subtask ID 1 is "DataCollector" and the call order is 1. According to the input and output parameter definitions in the API interface library, verify whether the input and output of each API interface conform to the task process order. For example, verify whether the output of "DataCollector" is the input of "DemandPredictor", and whether the output of "DemandPredictor" is the input of "SupplierEvaluator". Set verification rules, such as input and output parameter type matching, data format consistency, etc. The verification result is output in the form of a boolean value. If the verification passes, continue with the subsequent operations; if the verification fails, return an error message and prompt to re-adjust the API call chain. Associate each task node with the corresponding API interface to form a mapping relationship between the complete task process and the API call chain. For example, task node 1 is bound to the API interface "DataCollector", and task node 2 is bound to the API interface "DemandPredictor". Store the binding result in the task requirement technology blueprint table, and the table structure includes fields such as task node ID, API interface name, input and output parameters, etc. For example, the API interface name for task node ID 1 is "DataCollector", the input parameter is "supply chain data", and the output parameter is "demand prediction data". The finally generated task requirement technology blueprint is stored in JSON format, containing the detailed information of the task flow chart and the API call chain, providing a basis for subsequent agent construction and optimization.
[0030] Preferably, in step S2, generating an agent workflow according to the task requirement technology blueprint and using a preset dynamic workflow engine includes: Extract the start node and end node of the task from the task requirement technology blueprint; Identify the key task nodes in the technology blueprint. Each key task node corresponds to a specific business operation decision point, and record the name and function of each key task node; Assign a unique identifier to each key task node and record its position coordinates in the technology blueprint; According to the order and dependency of the key task nodes, draw a preliminary framework diagram of the workflow; the framework diagram includes nodes and connection lines, and the connection lines represent the execution order between the nodes and use arrows to indicate the direction; Execute parameters for each node configuration in the preliminary framework diagram according to the technical blueprint, where the execution parameters include execution time, resource requirements, and priority; In the preset dynamic workflow engine, instantiate the configured nodes and connection lines to generate an executable intelligent agent workflow.
[0031] In the embodiments of the present invention, the system calls the technical blueprint parsing module to read the data stored in the task requirement technical blueprint table. The technical blueprint table contains fields such as task node ID, node name, node type, input and output parameters, etc. The system identifies the starting node through the node type field. The starting node is usually marked as the "START" type. For example, the node with node ID 1 has a node name of "Data Collection" and a node type of "START". The system identifies the ending node through the node type field. The ending node is usually marked as the "END" type. For example, the node with node ID 10 has a node name of "Result Output" and a node type of "END". Store the information of the starting node and the ending node in the workflow configuration table, and the table structure includes fields such as node ID, node name, node type, position coordinates, etc. Call the key node identification module to identify the key task nodes through the node function description field. The key nodes usually contain decision logic or important business operations. For example, the node with node ID 3 has a node name of "Demand Forecasting" and a function description of "Forecast future demand based on historical data". Record the names and functions of the key task nodes in the workflow configuration table. For example, the node with node ID 3 has a node name of "Demand Forecasting" and a function description of "Forecast future demand based on historical data". Call the unique identifier generation module to generate a unique identifier for each key task node. For example, the node with node ID 3 generates a unique identifier of "UUID-001". According to the layout information in the technical blueprint, record the position coordinates of each key task node. For example, the position coordinates of the node with node ID 3 are (100, 200). Update the unique identifier and the position coordinates to the workflow configuration table. For example, the unique identifier of the node with node ID 3 is "UUID-001" and the position coordinates are (100, 200). Call the workflow framework diagram drawing module to initialize a blank framework diagram. The framework diagram is stored in a graphical form and contains nodes and connection lines. According to the information in the workflow configuration table, add each key task node to the framework diagram. The nodes are represented in a graphical form and contain the node name, unique identifier, and position coordinates. According to the dependency relationship in the technical blueprint, draw connection lines to represent the execution order between nodes. The connection lines use arrows to indicate the direction. For example, the node "Data Collection" with node ID 1 is connected to the node "Demand Forecasting" with node ID 2, and draw an arrow from (100, 100) to (200, 200). Store the preliminary framework diagram in the workflow framework diagram database, and the database table structure includes fields such as node ID, node name, unique identifier, position coordinates, connection line start point, connection line end point, etc. Extract the execution parameters of each node from the technical blueprint table. For example, the execution time of the node "Demand Forecasting" with node ID 3 is 10 minutes, the resource requirements are 2 CPU cores and 4 GB of memory, and the priority is high. Call the parameter configuration module to configure the extracted execution parameters into the workflow configuration table.For example, the execution time of the node "Demand Forecasting" with node ID 3 is 10 minutes, the resource requirements are 2 CPU cores and 4 GB of memory, and the priority is high. Update the configured execution parameters to the workflow configuration table. For example, the execution time of the node "Demand Forecasting" with node ID 3 is 10 minutes, the resource requirements are 2 CPU cores and 4 GB of memory, and the priority is high. Call the dynamic workflow engine to initialize the workflow instantiation module. The workflow engine supports multiple task scheduling algorithms, such as priority scheduling, resource scheduling, etc. According to the information in the workflow configuration table, each node is instantiated as a task object. The task object contains information such as the node name, unique identifier, execution parameters, etc. For example, the node "Demand Forecasting" with node ID 3 is instantiated as a task object, which includes the unique identifier "UUID-001", the execution time of 10 minutes, the resource requirements of 2 CPU cores and 4 GB of memory, and the priority is high. According to the connection line information in the workflow configuration table, the connection lines are instantiated as dependencies between tasks. The dependencies are stored in the form of references to task objects. For example, the task object "Data Collection" depends on the task object "Demand Forecasting". Assemble the instantiated nodes and connection lines into a complete intelligent agent workflow. The workflow is stored in a graphical form and supports real-time monitoring and dynamic adjustment. Store the generated intelligent agent workflow in the workflow instance database. The database table structure includes fields such as task object ID, task name, unique identifier, execution parameters, dependencies, etc. For example, the task object with task object ID 3 has the task name "Demand Forecasting", the unique identifier is "UUID-001", the execution time is 10 minutes, the resource requirements are 2 CPU cores and 4 GB of memory, and the dependency is the "Data Collection" with task object ID 1.
[0032] Preferably, in step S2, the intelligent agent workflow is divided into an ultra-long thinking chain based on the context information, and determining the nodes and paths of the intelligent agent workflow includes: Extract the context information of each task node from the intelligent agent workflow, including the input parameters, output parameters, and execution environment of the task node; Encode the context information of each task node, and convert the input parameters, output parameters, and execution environment into vector form; Calculate the cosine similarity between the vectors to determine the similarity between adjacent task nodes; Group the task nodes according to the similarity between adjacent task nodes, and group the task nodes with similarity higher than the preset similarity threshold into the same thinking unit; For each thinking unit, determine the execution order of the task nodes according to the dependencies and priorities of the task nodes; Connect all the thinking units in series according to the execution order of the task nodes to form an ultra-long thinking chain; Optimize the ultra-long thought chain by adjusting the order and path of task nodes to ensure the coherence and efficiency of the thought chain; Determine the nodes and paths of the agent workflow according to the task nodes and path connection methods in the ultra-long thought chain.
[0033] In the embodiments of the present invention, the system call context extraction module reads the task node information stored in the workflow configuration table. The workflow configuration table includes fields such as task node ID, node name, input parameters, output parameters, execution environment, etc. For example, for the task node with ID 3, the node name is "Demand Forecasting", the input parameter is "Historical Sales Data", the output parameter is "Future Demand Forecast", and the execution environment is "2 CPU cores, 4GB of memory". The extracted context information is stored in the context information table, and the table structure includes fields such as task node ID, input parameters, output parameters, execution environment, etc. For example, for the task node with ID 3, the input parameter is "Historical Sales Data", the output parameter is "Future Demand Forecast", and the execution environment is "2 CPU cores, 4GB of memory". The system call vector encoding module converts the context information into vector form using predefined encoding rules. For example, the input parameter "Historical Sales Data" is converted into the vector 1, 0, 0, the output parameter "Future Demand Forecast" is converted into the vector 0, 1, 0, and the execution environment "2 CPU cores, 4GB of memory" is converted into the vector 0, 0, 1. The encoded vectors are stored in the context vector table, and the table structure includes fields such as task node ID, input parameter vector, output parameter vector, execution environment vector, etc. For example, for the task node with ID 3, the input parameter vector is 1, 0, 0, the output parameter vector is 0, 1, 0, and the execution environment vector is 0, 0, 1. The system call cosine similarity calculation module calculates the cosine similarity between the input parameter vector, output parameter vector, and execution environment vector of each task node. The cosine similarity is determined by calculating the ratio of the dot product of two vectors to the product of their magnitudes. For example, the cosine similarity between the input parameter vector 1, 0, 0 of the task node with ID 3 and the input parameter vector 1, 0, 0 of the task node with ID 4 is 1. The calculated cosine similarity is stored in the similarity table, and the table structure includes fields such as task node ID1, task node ID2, similarity, etc. For example, the similarity between the task node with ID 3 and the task node with ID 4 is 1. The system sets the similarity threshold to 0.8. When the cosine similarity between two task nodes is higher than this threshold, they are grouped into the same thinking unit. The system call grouping module groups the task nodes according to the data in the similarity table. For example, the cosine similarity between the task node with ID 3 and the task node with ID 4 is 1, which is higher than the threshold 0.8, so they are grouped into the same thinking unit. The grouping result is stored in the grouping information table, and the table structure includes fields such as thinking unit ID, task node ID, etc. For example, the thinking unit ID 1 includes the task node ID 3 and the task node ID 4. The system extracts the dependency relationship and priority of each task node from the workflow configuration table. For example, the task node with ID 3 has a high priority and depends on the task node with ID 2. The system call execution order determination module determines the execution order of the task nodes according to the dependency relationship and priority.For example, the priority of task node with ID 2 is medium, the priority of task node with ID 3 is high, and task node with ID 3 depends on task node with ID 2. Therefore, the execution order is 2 → 3. Store the determined execution order in the execution order table, and the table structure includes fields such as thinking unit ID, task node ID, execution order, etc. For example, the execution order of the task node with thinking unit ID 1 is 2 → 3. The system calls the concatenation module and concatenates all thinking units according to the data in the execution order table. For example, the execution order of the task node with thinking unit ID 1 is 2 → 3, and the execution order of the task node with thinking unit ID 2 is 4 → 5. The concatenated ultra-long thinking chain is 2 → 3 → 4 → 5. Store the ultra-long thinking chain in the ultra-long thinking linked list, and the table structure includes fields such as task node ID, execution order, etc. For example, the ultra-long thinking linked list records that the execution order of task node with ID 2 is 1, the execution order of task node with ID 3 is 2, the execution order of task node with ID 4 is 3, and the execution order of task node with ID 5 is 4. The system sets optimization rules, such as reducing the waiting time between task nodes and optimizing resource utilization. The system calls the optimization module and adjusts the order and path of task nodes according to the optimization rules. For example, by adjusting the order of task nodes, the waiting time between task nodes is reduced, and the resource utilization efficiency is improved. Store the optimized ultra-long thinking chain in the optimized ultra-long thinking linked list, and the table structure includes fields such as task node ID, optimized execution order, etc. For example, the optimized ultra-long thinking linked list records that the optimized execution order of task node with ID 2 is 1, the optimized execution order of task node with ID 3 is 2, the optimized execution order of task node with ID 4 is 3, and the optimized execution order of task node with ID 5 is 4. The system calls the node and path determination module and determines the nodes and paths of the agent workflow according to the data in the optimized ultra-long thinking linked list. For example, the node name of task node with ID 2 is "data collection", the node name of task node with ID 3 is "demand prediction", and the path is 2 → 3. Update the determined nodes and paths to the agent workflow to form the final agent workflow. The final agent workflow is stored in a graphical form and supports real-time monitoring and dynamic adjustment. Store the final agent workflow in the workflow instance database, and the database table structure includes fields such as task node ID, task name, execution order, dependency relationship, etc. For example, the task name of task node with ID 2 is "data collection", the execution order is 1, and the dependency relationship is none; the task name of task node with ID 3 is "demand prediction", the execution order is 2, and the dependency relationship is task node with ID 2.
[0034] As an example of the present invention, refer to Figure 2 shown. In this example, step S3 includes: Step S31: Collect intelligent agent business requirement data from multiple data sources through a preset data collection interface. The data sources include user input, sensor data, and historical business records; Step S32: Preprocess the intelligent agent business requirement data, including data cleaning, format conversion, and data annotation, to obtain standard intelligent agent business requirement data; Step S33: Input the standard intelligent agent business requirement data into the dynamic adjustment module, and use the dynamic adjustment module to identify the nodes and paths that need to be adjusted in the standard intelligent agent business requirement data; Step S34: Dynamically adjust the nodes and paths that need to be adjusted in terms of task priority, resource requirements, and execution time through the dynamic adjustment module, and generate a business requirement adjustment instruction; Step S35: Reconfigure the nodes in the intelligent agent workflow according to the business requirement adjustment instruction, including adding, deleting, or modifying the attributes of the nodes; Step S36: Re-plan the paths in the intelligent agent workflow according to the business requirement adjustment instruction, including adjusting the path order and optimizing the path connection method.
[0035] In the embodiments of the present invention, the system configures a data collection interface to support the access of multiple data sources. For example, user inputs are collected through a Web interface, sensor data is collected through an Internet of Things (IoT) interface, and historical business records are collected through a database interface. User input data is received in real time through the Web interface, such as the business requirements filled in by users in a Web form. Sensor data is received through the IoT interface, such as real-time data from temperature sensors, pressure sensors, etc. Historical business records are read from the enterprise database through the database interface, such as past sales data, inventory records, etc. The collected data is stored in a preset data warehouse, which supports large-scale data storage and fast query. The system invokes a data cleaning module to remove noise and redundant information from the data. For example, duplicate user input records are removed, and outliers in sensor data are filtered out. The system invokes a format conversion module to convert data in different formats into a unified format. For example, JSON-format data input by users is converted into XML format, and binary-format sensor data is converted into text format. The system invokes a data annotation module to annotate the data for subsequent processing. For example, business requirement data input by users is annotated as "high priority" or "low priority", and sensor data is annotated as "normal" or "abnormal". The preprocessed data is stored in a preprocessing data table, and the table structure includes fields such as data ID, data content, data format, data annotation, etc. The preprocessed data is input into a dynamic adjustment module, and the dynamic adjustment module reads and analyzes the data. The dynamic adjustment module identifies the nodes and paths that need to be adjusted through a preset rule engine. For example, if the business requirement marked as "high priority" in the data does not match a certain node in the existing workflow, then that node is identified as the node that needs to be adjusted. The identification results are stored in an adjustment requirement table, and the table structure includes fields such as data ID, node ID to be adjusted, path ID to be adjusted, etc. The dynamic adjustment module adjusts the task priorities of relevant nodes according to the priority annotation in the business requirement data. For example, the priority of a node marked as "high priority" is adjusted from "medium" to "high". According to the resource requirement annotation in the business requirement data, the resource requirements of relevant nodes are adjusted. For example, the CPU resource requirement of a certain node is increased from 2 cores to 4 cores. According to the execution time annotation in the business requirement data, the execution time of relevant nodes is adjusted. For example, the execution time of a certain node is adjusted from 10 minutes to 15 minutes. The dynamic adjustment module generates business requirement adjustment instructions, and the instructions contain information such as the node ID to be adjusted, the adjusted task priority, resource requirements, and execution time. The generated adjustment instructions are stored in an adjustment instruction table, and the table structure includes fields such as instruction ID, node ID, adjusted task priority, resource requirements, execution time, etc. The system invokes a workflow configuration module to reconfigure the nodes in the workflow according to the adjustment instructions.For example, according to the adjustment instruction, a new node is added, a node that is no longer needed is deleted, or attributes such as the task priority, resource requirements, and execution time of an existing node are modified. The reconfigured node information is stored in the workflow configuration table, and the table structure includes fields such as node ID, node name, task priority, resource requirements, and execution time. The system calls the workflow path planning module to re-plan the paths in the workflow according to the adjustment instruction. For example, the path order is adjusted to adapt to the new task priority, and the path connection method is optimized to reduce the waiting time between nodes. The re-planned path information is stored in the workflow path table, and the table structure includes fields such as path ID, start node ID, end node ID, and path order.
[0036] Preferably, in step S4 at the strategic level, the reinforcement learning algorithm is adopted to predict the long-term business goal achievement path of the intelligent agent's business demand data and determine the quarterly-level optimization strategy, including the following steps: At the strategic level, the task completion rate of the intelligent agent's business demand data is measured and recorded as a strategic business indicator; The strategic business indicator is input into the reinforcement learning algorithm, and the environmental state and the initial strategy of the intelligent agent are initialized; By simulating the interaction between the environment and the intelligent agent, the business execution results of the intelligent agent under different strategies are collected, and the reward value of each strategy is recorded; According to the business execution results and the reward value of each strategy, the policy iteration method in the reinforcement learning algorithm is used to update the strategy of the intelligent agent; Based on the updated strategy, the long-term business goal achievement path of the intelligent agent's business demand data is predicted, and a quarterly-level optimization strategy is generated.
[0037] In the embodiments of the present invention, the system call task completion rate measurement module calculates the task completion rate according to the task execution records in the agent business requirement data. The task completion rate is determined by the ratio of the number of completed tasks to the total number of tasks. For example, if the total number of tasks is 100 and the number of completed tasks is 80, the task completion rate is 80%. The task completion rate is used as a strategic business indicator and stored in the strategic business indicator table. The strategic business indicator table includes fields such as indicator ID, indicator name, and indicator value. For example, the indicator name for indicator ID 1 is "task completion rate" and the indicator value is 80%. The system selects the model-based reinforcement learning (MBRL) algorithm as the optimization algorithm for the strategic layer. The system initializes the environmental state, including information such as the current task completion rate, resource usage, and time progress. For example, the current task completion rate is 80%, the resource utilization rate is 60%, and the time progress is the second quarter. The system initializes the initial policy of the agent, and the initial policy is generated based on historical data and preset rules. For example, the initial policy stipulates that resource investment is increased when the task completion rate is lower than 70%, and resource investment is reduced when the task completion rate is higher than 90%. The system sets up a simulation environment, which includes a task generator, a resource allocator, and a time controller. The task generator generates tasks according to business requirements, the resource allocator allocates resources according to the agent's policy, and the time controller controls the simulation time progress. The system runs the simulation environment, and the agent interacts with the environment according to the initial policy. For example, in the first interaction, the agent selects to increase resource investment according to the initial policy, and the task completion rate increases to 85% and the resource utilization rate increases to 70%. The system records the reward value of each policy, and the reward value is comprehensively calculated based on the task completion rate, resource utilization rate, and time progress. For example, the reward value calculation formula is: reward value = task completion rate × 0.5 + resource utilization rate × 0.3 + time progress × 0.2. In the first interaction, the reward value is 0.85 × 0.5 + 0.70 × 0.3 + 0.5 × 0.2 = 0.735. The system sets the parameters of policy iteration, including the learning rate, discount factor, and number of iterations. For example, the learning rate is 0.01, the discount factor is 0.99, and the number of iterations is 1000. The system runs the policy iteration algorithm and updates the agent's policy according to the reward value of each interaction. For example, in the first iteration, the agent adjusts the policy according to the reward value of 0.735 and increases the resource investment ratio when the task completion rate is lower than 75%. The system determines whether the policy converges, and the convergence condition is that the policy change in 10 consecutive iterations is less than 0.001. If the policy converges, the iteration stops; otherwise, the next iteration continues. The system uses the updated policy and combines it with the Monte Carlo tree search (MCTS) technology to predict the long-term business goal achievement path of the agent business requirement data. For example, it is predicted that within the next 4 quarters, through policy adjustments each quarter, the task completion rate will gradually increase to 95%.The system generates quarterly optimization strategies based on the predicted long-term business goal achievement path. The quarterly optimization strategies include task allocation, resource allocation, and policy adjustment suggestions for each quarter. For example, it is recommended to increase resource investment in the first quarter to improve the task completion rate, and it is recommended to optimize task allocation in the second quarter to improve resource utilization efficiency. The generated quarterly optimization strategies are stored in the optimization strategy table, and the table structure includes fields such as strategy ID, quarter, task allocation, resource allocation, and policy adjustment suggestions. For example, for strategy ID 1, the quarter is the first quarter, the task allocation is "increase the resource investment for task X", the resource allocation is "increase 2 cores of CPU resources", and the policy adjustment suggestion is "increase resource investment when the task completion rate is lower than 75%".
[0038] Preferably, in step S4 at the tactical layer, using the proximal policy optimization algorithm, adjusting the intelligent agent workflow node configuration and resource allocation scheme in the weekly dimension includes: At the tactical layer, extract the current node configuration and resource allocation situation from the intelligent agent workflow; among them, the current node configuration includes the type, quantity, and connection relationship of the nodes, and the resource allocation situation includes the resource type, quantity, and allocation ratio of each node; Input the current node configuration and resource allocation situation into the proximal policy optimization algorithm, initialize the policy parameters, and set the running environment of the algorithm, including the initial state of the simulation environment and the initial policy of the intelligent agent; In the weekly dimension, evaluate the resource allocation of each node, record the resource utilization efficiency and task completion situation of each node to generate a resource allocation evaluation result; among them, the resource utilization efficiency includes the type, quantity, and usage time of the resources, and the task completion situation includes the type, quantity, and completion time of the tasks; According to the resource allocation evaluation result, use the proximal policy optimization algorithm to adjust the resource allocation of each node, optimize the resource utilization efficiency, and update the policy parameters.
[0039] In the embodiments of the present invention, the system calls the node configuration extraction module to read the node configuration information in the agent workflow. The node configuration information includes the type, quantity, and connection relationship of the nodes. For example, the workflow contains three types of nodes: data processing nodes, analysis nodes, and output nodes, with quantities of 5, 3, and 2 respectively. The connection relationship is that the data processing nodes are connected to the analysis nodes, and the analysis nodes are connected to the output nodes. The system calls the resource allocation extraction module to read the resource allocation situation of each node. The resource allocation situation includes the resource type, quantity, and allocation ratio. For example, the data processing node is allocated 2 cores of CPU resources and 4GB of memory, with an allocation ratio of 50%; the analysis node is allocated 4 cores of CPU resources and 8GB of memory, with an allocation ratio of 30%; the output node is allocated 2 cores of CPU resources and 4GB of memory, with an allocation ratio of 20%. The extracted node configuration and resource allocation situation are stored in the node configuration table and the resource allocation table. The node configuration table contains fields such as node ID, node type, quantity, and connection relationship; the resource allocation table contains fields such as node ID, resource type, quantity, and allocation ratio. The system selects the Proximal Policy Optimization (PPO) algorithm as the optimization algorithm for the tactical layer. The system initializes the policy parameters of the PPO algorithm, including the learning rate, discount factor, batch size, etc. For example, the learning rate is 0.001, the discount factor is 0.99, and the batch size is 64. The system sets the running environment of the algorithm, including the initial state of the simulation environment and the initial policy of the agent. The initial state of the simulation environment includes the current node configuration and resource allocation situation, and the initial policy of the agent is generated based on historical data and preset rules. For example, the initial policy stipulates that the resource allocation is increased when the resource utilization efficiency is lower than 60%, and the resource allocation is reduced when the task completion rate is higher than 80%. The system sets the evaluation period to once a week. The system calls the resource utilization efficiency evaluation module to calculate the resource utilization efficiency of each node. The resource utilization efficiency includes the type, quantity, and usage time of the resources. For example, the CPU resource utilization efficiency of the data processing node is 70%, the memory resource utilization efficiency is 65%, and the usage time is 40 hours. The system calls the task completion situation evaluation module to record the task completion situation of each node. The task completion situation includes the type, quantity, and completion time of the tasks. For example, the data processing node has completed 10 data cleaning tasks, and the completion time is 30 hours; the analysis node has completed 5 data analysis tasks, and the completion time is 20 hours. The resource utilization efficiency and task completion situation are stored in the resource allocation evaluation table, and the table structure includes fields such as node ID, resource type, resource utilization efficiency, task type, task quantity, and completion time. The system sets the adjustment strategy of the PPO algorithm, including the resource adjustment threshold and the task completion rate threshold. For example, the resource adjustment threshold is 60%, and the task completion rate threshold is 80%. The system runs the PPO algorithm and adjusts the resource allocation of each node according to the resource allocation evaluation results.For example, the resource utilization efficiency of the data processing node is 70%, which is higher than the resource adjustment threshold of 60%. Therefore, its resource allocation ratio is increased from 50% to 60%. The task completion rate of the analysis node is 75%, which is lower than the task completion rate threshold of 80%. Therefore, its resource allocation ratio is decreased from 30% to 25%. The system updates the policy parameters of the PPO algorithm according to the adjusted resource allocation. For example, the updated learning rate is 0.0008, the discount factor is 0.98, and the batch size is 128. The adjusted resource allocation and the updated policy parameters are stored in the resource allocation adjustment table and the policy parameter table. The resource allocation adjustment table contains fields such as node ID, resource type, adjusted quantity, adjusted allocation ratio, etc.; the policy parameter table contains fields such as parameter name and parameter value.
[0040] Preferably, in the execution layer in step S4, applying the deep Q-network algorithm to optimize the single-task execution policy of the agent workflow in real time includes: In the execution layer, extract the current single-task execution policy from the agent workflow, including the input parameters, output parameters, and execution environment of the task; Input the extracted single-task execution policy into the deep Q-network algorithm to initialize the weights and biases of the Q-network; In the execution layer, monitor the execution process of each task in real time and record the execution status and results of the task; According to the execution status and results of the task, use the deep Q-network algorithm to calculate the Q value of each task and evaluate the pros and cons of the task execution policy; According to the calculated Q value, adjust the execution policy of the task to optimize the execution efficiency and results of the task; Verify the adjusted execution policy to ensure that it meets the business requirements and actual operating conditions; According to the verification results, fine-tune the execution policy to further optimize the execution efficiency and results of the task; Apply the fine-tuned execution policy to the agent workflow to update the single-task execution policy of the workflow.
[0041] In the embodiments of the present invention, the system call policy adjustment module adjusts the execution policy of tasks according to the Q value. For example, the Q value of the task "data cleaning" with task ID 1 is 0.885, which is lower than the preset Q value threshold of 0.9. Therefore, the execution policy needs to be adjusted. The adjustment policies include increasing resource allocation, optimizing the task execution process, etc. For example, increase the CPU resource allocation of the task "data cleaning" from 2 cores to 4 cores and optimize the task execution process to reduce the execution time. The system updates the execution policy of the task according to the adjusted policy. For example, the execution policy of the task "data cleaning" with task ID 1 is updated to: the input parameter is "raw data file", the output parameter is "cleaned data file", and the execution environment is "CPU 4 cores, memory 4GB". Store the adjusted execution policy in the execution policy table, and the table structure includes fields such as task ID, task name, input parameter, output parameter, execution environment, etc. For example, the task name of the task with task ID 1 is "data cleaning", the input parameter is "raw data file", the output parameter is "cleaned data file", and the execution environment is "CPU 4 cores, memory 4GB". The system configuration verification module verifies whether the adjusted execution policy meets the business requirements and actual operating conditions. The verification module checks whether the execution result of the task meets the expectations by simulating the execution of the task. The system operation verification module simulates the execution of the task "data cleaning" with task ID 1 and checks the execution result of the task. For example, the execution result of the task "data cleaning" is successful, the output parameter is "cleaned data file", the task completion time is 2024-10-01 09:00, and the task execution result meets the expectations. Store the verification result in the verification result table, and the table structure includes fields such as task ID, verification result, etc. For example, the verification result of the task with task ID 1 is successful. The system configuration fine-tuning module fine-tunes the execution policy according to the verification result. The fine-tuning module further optimizes the execution efficiency and result of the task by adjusting the resource allocation and execution process of the task. The system operation fine-tuning module fine-tunes the task "data cleaning" with task ID 1. For example, the fine-tuning module adjusts the CPU resource allocation of the task "data cleaning" from 4 cores to 3 cores and optimizes the task execution process to reduce resource waste. Store the fine-tuned execution policy in the execution policy table, and the table structure includes fields such as task ID, task name, input parameter, output parameter, execution environment, etc. For example, the task name of the task with task ID 1 is "data cleaning", the input parameter is "raw data file", the output parameter is "cleaned data file", and the execution environment is "CPU 3 cores, memory 4GB". The system call workflow update module applies the fine-tuned execution policy to the agent workflow. For example, the execution policy of the task "data cleaning" with task ID 1 is updated to: the input parameter is "raw data file", the output parameter is "cleaned data file", and the execution environment is "CPU 3 cores, memory 4GB".Store the updated execution policy in the workflow configuration table, and the table structure includes fields such as task ID, task name, input parameters, output parameters, execution environment, etc. For example, for the task with task ID 1, the task name is "Data Cleaning", the input parameter is "raw data file", the output parameter is "cleaned data file", and the execution environment is "3 CPU cores and 4GB of memory".
[0042] As an example of the present invention, refer to Figure 3 As shown, in this example, step S5 includes: Step S51: During the process of the agent executing the task, monitor the input and output data of the agent in real time and record the intermediate state of the task execution; Step S52: Based on the hierarchical reinforcement learning framework, decompose the task execution process into multiple subtasks, and each subtask corresponds to a specific decision point; Step S53: Assign an independent reinforcement learning module to each subtask, and each module is responsible for optimizing the decision-making strategy of the subtask; Step S54: At each decision point, use the reinforcement learning module to calculate the optimal action in the current state and generate a decision instruction; Step S55: Send the generated decision instruction to the agent to guide the next action of the agent; Step S56: According to the execution result of the agent, update the parameters of the reinforcement learning module and optimize the decision-making strategy; Step S57: Integrate the decision-making strategies of all subtasks to construct a dynamic decision-making network for the agent.
[0043] In the embodiments of the present invention, the system configures a real-time monitoring module to monitor the input and output data of the agent. The monitoring module records the execution status of the task through the task execution log, including the task start time, end time, execution progress, resource usage, etc. The system records the input and output data of the agent in real time. For example, for the task "data cleaning" with task ID 1, the input data is "raw data file", and the output data is "cleaned data file". The system records the intermediate state of the task execution. For example, during the execution of the task "data cleaning" with task ID 1, the intermediate states include data reading completed, data preprocessing completed, data cleaning completed, etc. The input and output data and the intermediate state of the task are stored in the task execution log table, and the table structure includes fields such as task ID, task name, input data, output data, intermediate state, etc. For example, for the task with task ID 1, the task name is "data cleaning", the input data is "raw data file", the output data is "cleaned data file", and the intermediate state includes data reading completed, data preprocessing completed, data cleaning completed. The system configures a task decomposition module to decompose the task execution process into multiple subtasks according to predefined business rules and the domain knowledge base. For example, the task "data cleaning" is decomposed into three subtasks: "data reading", "data preprocessing", and "data cleaning". Each subtask corresponds to a specific decision point. For example, the decision point of the "data reading" subtask is which data reading method to choose, the decision point of the "data preprocessing" subtask is which preprocessing algorithm to choose, and the decision point of the "data cleaning" subtask is which cleaning strategy to choose. The decomposed subtasks are stored in the subtask table, and the table structure includes fields such as subtask ID, subtask name, decision point, etc. For example, for the subtask with subtask ID 1, the subtask name is "data reading", and the decision point is "choose data reading method". The system configures an independent reinforcement learning module for each subtask. For example, a reinforcement learning module is configured for the "data reading" subtask, and another reinforcement learning module is configured for the "data preprocessing" subtask. The system initializes the parameters of each reinforcement learning module, including the learning rate, discount factor, experience replay pool size, etc. For example, the learning rate is 0.001, the discount factor is 0.99, and the experience replay pool size is 10,000. The configuration information of each reinforcement learning module is stored in the reinforcement learning module table, and the table structure includes fields such as module ID, subtask ID, learning rate, discount factor, experience replay pool size, etc. For example, for the module with module ID 1, the subtask ID is 1, the learning rate is 0.001, the discount factor is 0.99, and the experience replay pool size is 10,000. The system senses the execution state of the current task. For example, at the decision point of the "data reading" subtask, the current state includes information such as the format and size of the input data. The system calls the reinforcement learning module to calculate the optimal action according to the current state. For example, at the decision point of the "data reading" subtask, the most suitable data reading method is selected according to the format and size of the input data.The system generates decision instructions to guide the next actions of the agent. For example, the generated decision instruction is "Select the data reading method in CSV format". The generated decision instructions are stored in the decision instruction table, and the table structure includes fields such as decision point ID and decision instructions. For example, the decision instruction for decision point ID 1 is "Select the data reading method in CSV format". The system sends the decision instructions to the agent through a preset communication interface. The communication interface supports real-time data transmission to ensure that the instructions can be delivered in a timely manner. The agent receives the decision instructions and executes the next actions according to the instructions. For example, after receiving the instruction "Select the data reading method in CSV format", the agent starts to execute the data reading operation. The system records the results of the agent's execution of the decision instructions. For example, if the agent successfully executes the data reading operation, the execution result is recorded as successful. The system evaluates the results of the agent's execution of the decision instructions. For example, it evaluates the execution efficiency and resource usage of the data reading operation. The system updates the parameters of the reinforcement learning module using the parameter update method in the reinforcement learning algorithm according to the execution results. For example, it adjusts parameters such as the learning rate and discount factor according to the execution results to optimize the decision-making strategy. The updated parameters are stored in the reinforcement learning module table, and the table structure includes fields such as module ID, sub-task ID, learning rate, discount factor, and experience replay pool size. For example, for module ID 1, the sub-task ID is 1, the updated learning rate is 0.0008, the discount factor is 0.98, and the experience replay pool size is 12,000. The system integrates the decision-making strategies of each sub-task into a complete decision-making network. For example, it integrates the decision-making strategies of the three sub-tasks of "data reading", "data preprocessing", and "data cleaning" into a dynamic decision-making network. The system constructs the agent's dynamic decision-making network, which is stored in a graphical form and includes nodes and edges. The nodes represent sub-tasks, and the edges represent the execution order and dependency relationships between sub-tasks. The constructed agent's dynamic decision-making network is stored in the decision-making network database, and the database table structure includes fields such as node ID, node name, edge start point, and edge end point. For example, for node ID 1, the node name is "data reading", the edge start point is 1, and the edge end point is 2, indicating that the "data preprocessing" sub-task is executed after the "data reading" sub-task is completed.
[0044] Therefore, in all respects, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be encompassed within the present invention.
[0045] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating an intelligent agent dynamic decision-making network based on reinforcement learning, characterized in that It includes the following steps: Step S1: Receive the information requirement data input by the user; perform feature parsing on the information requirement data, and extract the business objective, constraints, and key parameters to obtain the information requirement parsing result; Identify the task process corresponding to the information requirement parsing result, and match the API call chain according to the task process to obtain the technical blueprint of the task requirements; Step S2: Generate the agent workflow according to the technical blueprint of the task requirements and using a preset dynamic workflow engine; perform context analysis on the language requirement parsing result to obtain context information; Divide the agent workflow into an ultra-long thinking chain based on the context information, and determine the nodes and paths of the agent workflow according to the ultra-long thinking chain; Step S3: Collect the agent business requirement data; dynamically adjust the nodes and paths in the agent workflow according to the agent business requirement data; Step S4: Build an agent hierarchical reinforcement learning framework, including a strategic layer, a tactical layer, and an execution layer; in the strategic layer, use a reinforcement learning algorithm to predict the long-term business objective achievement path of the agent business requirement data and determine the quarterly-level optimization strategy; In the tactical layer, use the proximal policy optimization algorithm to adjust the agent workflow node configuration and resource allocation plan on a weekly basis; In the execution layer, apply the deep Q-network algorithm to optimize the single-task execution strategy of the agent workflow in real time; Step S5: During the process of the agent executing tasks, build and generate an agent dynamic decision network based on the agent hierarchical reinforcement learning framework.
2. The method for generating an intelligent agent dynamic decision network based on reinforcement learning according to claim 1, wherein The feature parsing of the information requirement data in Step S1 and the extraction of the business objective, constraints, and key parameters include: Extract the features of the information requirement data, and convert the extracted requirement features into an information requirement vector set; Input the information requirement vector set into a preset large model for intention recognition, and output the intention recognition matching result; Divide the intention recognition matching result into a business objective, constraints, and key parameters, and record their positions in the information requirement data to obtain the information requirement parsing result.
3. The method for generating an intelligent agent dynamic decision network based on reinforcement learning according to claim 2, wherein The identification of the task process corresponding to the information requirement parsing result in Step S1 and the matching of the API call chain according to the task process include: Decompose the business objective in the information requirement parsing result into multiple subtasks, where each subtask corresponds to a business operation information; Perform feature recognition on each subtask to determine its order and dependency relationship in the task process; According to the order and dependency relationship of the subtasks, build a task flow chart, where the task flow chart includes task nodes and task edges, the task nodes represent specific subtasks, and the task edges represent the order and dependency relationship between subtasks; For each task node in the task flow chart, find the matching API interface type; According to the task edges in the task flow chart, determine the call order between API interfaces to form an API call chain; Verify the API call chain to detect whether the input and output of each API interface conform to the task process order; Bind the verified API call chain to the task flow chart to form a complete mapping relationship between the task process and the API call chain, so as to obtain the technical blueprint of the task requirements.
4. The method for generating an intelligent agent dynamic decision network based on reinforcement learning according to claim 1, wherein In step S2, generating the agent workflow according to the task requirement technical blueprint and using the preset dynamic workflow engine includes: Extracting the start node and end node of the task from the task requirement technical blueprint; Identifying the key task nodes in the technical blueprint, where each key task node corresponds to a specific business operation decision point, and recording the name and function of each key task node; Assigning a unique identifier to each key task node and recording its position coordinates in the technical blueprint; Drawing a preliminary framework diagram of the workflow according to the order and dependency relationship of the key task nodes; the framework diagram includes nodes and connection lines, the connection lines represent the execution order between nodes, and arrows are used to indicate the direction; Configuring execution parameters for each node in the preliminary framework diagram according to the technical blueprint, and the execution parameters include execution time, resource requirements, and priority; In the preset dynamic workflow engine, instantiating the configured nodes and connection lines to generate an executable agent workflow.
5. The method for generating an intelligent agent dynamic decision network based on reinforcement learning according to claim 1, wherein In step S2, dividing the agent workflow into an ultra-long thinking chain based on the context information, and determining the nodes and paths of the agent workflow includes: Extracting the context information of each task node from the agent workflow, including the input parameters, output parameters, and execution environment of the task node; Encoding the context information of each task node and converting the input parameters, output parameters, and execution environment into vector form; Calculating the cosine similarity between vectors to determine the similarity between adjacent task nodes; Grouping the task nodes according to the similarity between adjacent task nodes, and classifying the task nodes with similarity higher than the preset similarity threshold into the same thinking unit; For each thinking unit, determining the execution order of the task nodes according to the dependency relationship and priority of the task nodes; Connecting all thinking units in series according to the execution order of the task nodes to form an ultra-long thinking chain; Optimizing the ultra-long thinking chain by adjusting the order and path of the task nodes to ensure the coherence and efficiency of the thinking chain; Determining the nodes and paths of the agent workflow according to the task nodes and path connection methods in the ultra-long thinking chain.
6. The method for generating an intelligent agent dynamic decision network based on reinforcement learning according to claim 1, wherein Step S3 includes the following steps: Step S31: Collecting agent business requirement data from multiple data sources through a preset data collection interface, and the data sources include user input, sensor data, and historical business records; Step S32: Preprocessing the agent business requirement data, including data cleaning, format conversion, and data annotation, to obtain standard agent business requirement data; Step S33: Inputting the standard agent business requirement data into the dynamic adjustment module, and using the dynamic adjustment module to identify the nodes and paths that need to be adjusted in the standard agent business requirement data; Step S34: Dynamically adjusting the task priority, resource requirements, and execution time of the nodes and paths that need to be adjusted through the dynamic adjustment module, and generating a business requirement adjustment instruction; Step S35: Reconfiguring the nodes in the agent workflow according to the business requirement adjustment instruction, including adding, deleting, or modifying the attributes of the nodes; Step S36: Re-plan the paths in the agent workflow according to the business requirements adjustment instructions, including adjusting the path order and optimizing the path connection method.
7. The method for generating an intelligent agent dynamic decision network based on reinforcement learning according to claim 1, wherein In step S4, at the strategic level, a reinforcement learning algorithm is adopted to predict the long-term business goal achievement path of the agent business requirement data and determine the quarterly-level optimization strategy, including the following steps: At the strategic level, measure the task completion rate of the agent business requirement data and record it as a strategic business indicator; Input the strategic business indicator into the reinforcement learning algorithm and initialize the environmental state and the initial strategy of the agent; Through simulating the interaction between the environment and the agent, collect the business execution results of the agent under different strategies and record the reward value of each strategy; According to the business execution results and the reward values of each strategy, use the policy iteration method in the reinforcement learning algorithm to update the strategy of the agent; Based on the updated strategy, predict the long-term business goal achievement path of the agent business requirement data and generate a quarterly-level optimization strategy.
8. The method for generating an intelligent agent dynamic decision network based on reinforcement learning according to claim 1, wherein In step S4, at the tactical level, use the Proximal Policy Optimization (PPO) algorithm to adjust the agent workflow node configuration and resource allocation plan on a weekly basis, including: At the tactical level, extract the current node configuration and resource allocation situation from the agent workflow; among them, the current node configuration includes the type, quantity, and connection relationship of the nodes, and the resource allocation situation includes the resource type, quantity, and allocation ratio of each node; Input the current node configuration and resource allocation situation into the Proximal Policy Optimization algorithm, initialize the policy parameters, and set the running environment of the algorithm, including the initial state of the simulation environment and the initial strategy of the agent; On a weekly basis, evaluate the resource allocation of each node, record the resource utilization efficiency and task completion situation of each node to generate a resource allocation evaluation result; among them, the resource utilization efficiency includes the type, quantity, and usage time of the resources, and the task completion situation includes the type, quantity, and completion time of the tasks; According to the resource allocation evaluation result, use the Proximal Policy Optimization algorithm to adjust the resource allocation of each node, optimize the resource utilization efficiency, and update the policy parameters.
9. The method for generating an intelligent agent dynamic decision network based on reinforcement learning according to claim 1, wherein In step S4, at the execution level, apply the Deep Q-Network (DQN) algorithm to optimize the single-task execution strategy of the agent workflow in real time, including: At the execution level, extract the current single-task execution strategy from the agent workflow, including the input parameters, output parameters, and execution environment of the task; Input the extracted single-task execution strategy into the Deep Q-Network algorithm and initialize the weights and biases of the Q-network; At the execution level, monitor the execution process of each task in real time and record the execution status and results of the tasks; According to the execution status and results of the tasks, use the Deep Q-Network algorithm to calculate the Q value of each task and evaluate the pros and cons of the task execution strategy; According to the calculated Q value, adjust the execution strategy of the task, optimize the execution efficiency and results of the task; Verify the adjusted execution strategy to ensure that it meets the business requirements and actual operating conditions; According to the verification results, fine-tune the execution strategy to further optimize the execution efficiency and results of the task; Apply the fine-tuned execution strategy to the agent workflow and update the single-task execution strategy of the workflow.
10. The method for generating an intelligent agent dynamic decision network based on reinforcement learning according to claim 1, characterized in that Step S5 includes the following steps: Step S51: During the process of the agent executing the task, monitor the input and output data of the agent in real time and record the intermediate state of task execution; Step S52: Based on the hierarchical reinforcement learning framework, decompose the task execution process into multiple subtasks, and each subtask corresponds to a specific decision point; Step S53: Assign an independent reinforcement learning module to each subtask, and each module is responsible for optimizing the decision-making strategy of the subtask; Step S54: At each decision point, use the reinforcement learning module to calculate the optimal action in the current state and generate a decision instruction; Step S55: Send the generated decision instruction to the agent to guide the next action of the agent; Step S56: According to the execution result of the agent, update the parameters of the reinforcement learning module and optimize the decision-making strategy; Step S57: Integrate the decision-making strategies of all subtasks to construct a dynamic decision-making network for the agent.
Citation Information
Patent Citations
Task execution method and related device, equipment, storage medium and agent platform
CN118689563A
Logistics operation multi-target supervision method based on complex logistics field group
CN119338356A
Emergency fire protection hidden danger troubleshooting method based on multi-mode AI large model identification technology
CN119646271A
Intelligent agent automatic generation and scheduling system based on artificial intelligence large language model
CN119690536A
Dynamic data pipeline construction method based on artificial intelligence and multi-modal data processing
CN119830200A
Cited By
Event-driven business agent dynamic construction method and device
CN120407241A
Intelligent agent decision path optimization method fusing explicit topology and implicit semantics
CN120409874A
A method for agent decision-making path optimization that integrates explicit topology and implicit semantics
CN120409874B
Method and system for realizing enterprise collaborative decision-making based on AI algorithm and government-enterprise API
CN120743394A
Enterprise operation and maintenance method and equipment based on AI intelligent agent, and medium
CN120851940A