Complex task management method and device based on large model, medium and program product

Through complex task management methods based on large models and real-time adjustments of multi-agent systems, the shortcomings of traditional task management methods in large-scale, changing tasks and dynamic environments are solved, and more efficient task planning management is achieved.

CN120106539APending Publication Date: 2025-06-06ZIGUANG HENGYUE TECH CO LTD

Patent Information

Application Number
CN202510585869.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When traditional task management methods face large-scale and changing tasks and dynamic environments, it is difficult to ensure the accuracy, flexibility and real-time nature of task planning.

Method used

A complex task management method based on large models is adopted to collect and preprocess historical execution data, extract feature data sets, and generate preliminary planning schemes using pre-trained large models. At the same time, multi-agent systems are used to implement and adjust the planning scheme in real time, and dynamically adjust it based on real-time available resources and external environment data.

Benefits of technology

It improves the accuracy, flexibility and real-time nature of task planning, and can handle task management in complex scenarios more effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106539A_ABST
    Figure CN120106539A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a complex task management method and device based on a large model, a medium and a program product, and relates to the technical field of task management. The method comprises the steps of collecting historical execution data and performing feature extraction; obtaining basic data of the current task, and generating a preliminary plan scheme based on the feature data and the basic data by using the plan customization large model; executing a current plan scheme by using the multi-agent system and acquiring task state data; and dynamically adjusting the currently generated plan scheme based on the task state data, the real-time available resource data and the real-time external environment data by using the plan customization large model. According to the embodiment of the invention, the large model and the multi-agent system are used for learning the relationship between various factors and the task and generating and dynamically adjusting the planning scheme of the current task, so that the accuracy, the flexibility and the real-time performance of making the task plan are improved in a scene in which various consideration factors are involved and the external environment is dynamically changed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of task management, and in particular, to a complex task management method, device, medium and program product based on a large model. Background Art

[0002] Traditional solutions often rely on manual experience or simple rule systems to plan tasks. These solutions may work in small-scale, static environments, but when the scale of tasks increases, the types of resources increase, and the external environment changes dynamically, traditional methods can no longer meet the needs of efficient, accurate, and real-time task customization.

[0003] Therefore, in scenarios involving multiple considerations and dynamically changing external environments, how to improve the accuracy, flexibility, and real-time nature of task planning has become a difficult problem that needs to be solved urgently. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a complex task management method, device, medium and program product based on a large model, so as to improve the accuracy, flexibility and real-time performance of task planning in scenarios involving multiple considerations and dynamically changing external environments.

[0005] In a first aspect, an embodiment of the present application provides a complex task management method based on a large model, including: Collecting historical execution data corresponding to various task types; wherein the historical execution data includes task goal description, task execution steps, resource consumption records, time consumption information and external environmental factors; Preprocessing the historical execution data, and extracting features from the preprocessed historical execution data to obtain a feature data set; Obtain basic data corresponding to the current task to be processed, and use the pre-trained plan customization model to generate a preliminary plan based on the feature data set and the basic data; wherein the basic data includes task objectives, task available resources and external environment data; Utilize a pre-built multi-agent system to execute the currently generated plan in real time, and obtain task status data of the multi-agent system during the execution of the plan; Real-time available resource data and real-time external environment data are acquired, and the plan customization model is used to dynamically adjust the currently generated plan based on the task status data, the real-time available resource data and the real-time external environment data.

[0006] In an embodiment of the present application, a large model is used to learn the relationship between various factors and tasks based on historical execution data, and a corresponding plan is generated based on the current multiple complex factors. At the same time, a multi-agent system is used to process the current plan in real time, and the current plan is dynamically adjusted based on real-time collected status, resources, external environment and other data, thereby improving the accuracy, flexibility and real-time performance of task planning in scenarios involving multiple considerations and dynamically changing external environments.

[0007] In some possible embodiments, the multi-agent system includes a task management agent, a resource management agent, an environment monitoring agent, and a decision coordination agent; The task management agent is used to perform task decomposition, task allocation and task progress monitoring on the currently executed plan; The resource management agent is used to perform resource scheduling, resource allocation and resource status update on the currently executed plan; The environment monitoring agent is used to collect and analyze the current external environment data in real time; The decision-making coordination agent is used to integrate information of various agents and coordinate the behaviors among various agents.

[0008] In an embodiment of the present application, a multi-agent system is set up to automatically execute and manage the current plan in a refined manner, thereby effectively improving the task processing efficiency and resource allocation efficiency, thereby further improving the task planning management efficiency in complex scenarios.

[0009] In some possible embodiments, the complex task management method based on a large model further includes: Iteratively training the multi-agent system based on a preset reward function and a reinforcement learning algorithm to obtain an optimized and adjusted multi-agent system; The reward function is composed of multiple reward items and their corresponding weight coefficients, and the multiple reward items include a task completion reward item, a resource utilization efficiency reward item, a time cost reward item, and an inter-agent collaboration efficiency reward item; The task completion reward item is used to characterize the number of tasks completed by the intelligent agent, the quality of task completion and the importance of tasks; the resource utilization efficiency reward item is used to characterize the degree of resource conservation, efficiency improvement and resource sustainability; the time cost reward item is used to characterize the degree of shortening of the planned execution time and the deviation between the actual execution time and the expected time; the inter-agent collaboration efficiency reward item is used to characterize the timeliness of information exchange between intelligent agents, the effectiveness of information exchange and the efficiency of collaborative problem solving.

[0010] In an embodiment of the present application, the multi-agent system is iteratively trained by setting a reward function and a reinforcement learning algorithm, thereby further improving the intelligence and flexibility of the multi-agent system in managing task plans.

[0011] In some possible embodiments, the multi-agent system is iteratively trained based on a preset reward function and a reinforcement learning algorithm to obtain an optimized and adjusted multi-agent system, including: Initialize the parameters of the state and decision-making strategy of the pre-built multi-agent system; Acquire the experience data of the multi-agent system during the execution of the current plan, and store the experience data in the experience playback buffer; wherein the experience data includes the initial state, execution action, reward value and update state of each agent after the execution of the plan, and the reward value is calculated based on the reward function; Based on the experience data in the experience replay buffer, a preset reinforcement learning algorithm is used to update the parameters of the decision strategy of the multi-agent system; The steps of acquiring experience data and updating parameters are executed cyclically until a preset round end condition is reached to complete one round of training; Acquire the performance index of the multi-agent system, adjust the weight coefficient corresponding to each reward item in the reward function according to the performance index, and adjust the model learning rate and strategy update step size of each agent according to the performance index; Repeat the steps of performing a single round of training and obtaining the performance indicator until a preset number of training rounds is reached or the performance indicator meets a preset condition.

[0012] In an embodiment of the present application, by combining the empirical data such as the state, action and reward of the multi-agent system during the execution of the plan as the knowledge to guide each round of training, the efficiency and accuracy of the multi-agent system training are further improved.

[0013] In some possible embodiments, the acquiring of real-time available resource data and real-time external environment data, using the plan customization big model, and dynamically adjusting the currently generated plan scheme based on the task status data, the real-time available resource data, and the real-time external environment data, includes: Acquire in real time the task status time series data, available resource time series data and external environment time series data corresponding to the multi-agent system in the process of executing the currently generated plan; Inputting the task status time series data, the available resource time series data and the external environment time series data into a pre-trained time series prediction model to obtain a time series prediction result output by the time series prediction model; Generate plan adjustment auxiliary information corresponding to the time series prediction result; The plan customization model is used to dynamically adjust the currently generated plan based on the task status data, the real-time available resource data, the real-time external environment data and the plan adjustment auxiliary information.

[0014] In an embodiment of the present application, various time series data of the multi-agent system in the process of executing the plan are acquired in real time, and then a time series prediction model is used to perform data prediction based on these time series data, and the prediction results are used as one of the bases for guiding the large model to adjust the plan in real time, thereby further improving the real-time and accuracy of task planning management.

[0015] In some possible embodiments, the time series prediction model adopts a long short-term memory network model, and the training process of the time series prediction model includes: Preprocessing the historical execution data to obtain time series data; wherein the preprocessing includes time serialization and data cleaning; Dividing the time series data into a training set, a validation set, and a test set; Initialize and set the model parameters of the long short-term memory network model; wherein the model parameters include the number of layers, the number of hidden units in each layer, the activation function, the optimizer and the loss function; Iteratively training the long short-term memory network model using the training set, and updating the model parameters of the long short-term memory network model based on a back propagation algorithm and a gradient descent method; Based on the validation set, a model performance evaluation index of the long short-term memory network model is obtained, and the model parameters and hyperparameters of the long short-term memory network model are adjusted according to the model performance evaluation index to obtain a trained time series prediction model.

[0016] In an embodiment of the present application, by using a long short-term memory network model for training based on historical execution data, the efficiency of the model in acquiring timing features can be improved, thereby further improving the accuracy and real-time performance of task planning management.

[0017] In some possible embodiments, the task type includes at least one of a production and manufacturing task, a logistics and distribution task, a project management task, a resource scheduling task, and an emergency response task; the external environmental factors include at least one of weather conditions, market demand information, policy changes, and supply chain status.

[0018] In the embodiment of the present application, by considering multiple task types and multiple external environmental factors, the large model can better learn key features such as task characteristics, resource requirements and external environmental influences, thereby further improving the accuracy and flexibility of task planning management.

[0019] In a second aspect, an embodiment of the present application provides a complex task management device based on a large model, including: A data collection module is used to collect historical execution data corresponding to various task types; wherein the historical execution data includes task goal description, task execution steps, resource consumption records, time consumption information and external environmental factors; A feature extraction module, used to preprocess the historical execution data and extract features from the preprocessed historical execution data to obtain a feature data set; A plan generation module is used to obtain basic data corresponding to the current task to be processed, and use the pre-trained plan customization model to generate a preliminary plan based on the feature data set and the basic data; wherein the basic data includes task objectives, task available resources and external environment data; A plan execution module is used to use a pre-built multi-agent system to execute the currently generated plan in real time and obtain task status data of the multi-agent system during the execution of the plan; The plan adjustment module is used to obtain real-time available resource data and real-time external environment data, and use the plan customization model to dynamically adjust the currently generated plan based on the task status data, the real-time available resource data and the real-time external environment data.

[0020] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor can implement the method described in any embodiment of the first aspect when executing the program.

[0021] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any embodiment of the first aspect can be implemented.

[0022] In a fifth aspect, an embodiment of the present application provides a computer program product, wherein the computer program product includes a computer program, wherein when the computer program is executed by a processor, the method described in any embodiment of the first aspect can be implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0024] Figure 1 A flowchart of a complex task management method based on a large model provided in an embodiment of the present application; Figure 2 A schematic diagram of the structure of a complex task management device based on a large model provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0026] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0027] It should be noted that in modern society, all walks of life are facing increasingly complex task planning and resource management problems. With the expansion of task scale and complexity, traditional planning methods have been unable to meet the needs of efficiency, accuracy and real-time. Especially in scenarios involving multiple tasks, multiple resources, multiple goals and dynamic external environments, how to formulate complex plans that meet task requirements, make full use of resources and reduce time costs has become a problem that needs to be solved urgently.

[0028] Traditional planning methods often rely on manual experience or simple rule systems. However, manual experience is difficult to fully consider all factors, and is limited by personal ability and experience, and often cannot formulate the best plan. Simple rule systems lack flexibility and adaptability and cannot cope with complex and changing actual situations.

[0029] In response to the problems existing in the above-mentioned prior art, an embodiment of the present application provides a complex task management method based on a large model. By utilizing the large model to learn the relationship between various factors and tasks based on historical execution data, a preliminary plan is generated according to real-time data, and then a multi-agent system is used to execute the plan in real time. During the execution of the plan, the plan is adjusted in real time according to the latest multi-factor data, thereby improving the accuracy, flexibility and real-time performance of task planning.

[0030] like Figure 1 As shown, the embodiment of the present application provides a complex task management method based on a large model, which may include the following steps: S1. Collect historical execution data corresponding to various task types; wherein the historical execution data includes task goal description, task execution steps, resource consumption records, time consumption information and external environmental factors.

[0031] Furthermore, the task types include at least one of production and manufacturing tasks, logistics and distribution tasks, project management tasks, resource scheduling tasks and emergency response tasks; the external environmental factors include at least one of weather conditions, market demand information, policy changes and supply chain status.

[0032] Specifically, we first collect historical execution data covering various types of tasks, including task goal descriptions, task execution steps, resource consumption records (such as human, material, equipment and other resource consumption), time consumption information and external environmental factors, etc. Task types can include production and manufacturing tasks, logistics and distribution tasks, project management tasks, resource scheduling tasks and emergency response tasks, etc. External environmental factors can include weather conditions, market demand information, policy changes and supply chain status, etc., depending on the task type. The specific data types of external environmental factors can be set according to demand. These collected data will provide a solid data foundation for subsequent model training.

[0033] S2. Preprocess the historical execution data, and extract features from the preprocessed historical execution data to obtain a feature data set.

[0034] Exemplarily, the preprocessing operations on the collected historical execution data may include data cleaning, data deduplication, format normalization, data standardization, etc.; data cleaning can remove invalid or erroneous data; data deduplication can avoid duplicate data and save computing resources; format normalization can ensure data consistency; and data standardization can eliminate dimensional differences.

[0035] Then, statistical analysis, machine learning algorithms, etc. can be used to extract features from the preprocessed data to form a feature data set that can reflect key features such as task characteristics, resource status, and external environment changes. These feature data can be used to train the large customized planning model.

[0036] S3. Obtain the basic data corresponding to the current task to be processed, and use the pre-trained plan customization model to generate a preliminary plan based on the feature data set and basic data; the basic data includes task objectives, task available resources and external environment data.

[0037] According to the current task to be processed, the corresponding basic data is obtained, including the target description of the current task, the available resources required to execute the task, and the data of relevant external environmental factors affecting the task, etc.; the pre-built and pre-trained plan customization model is used to generate a preliminary plan corresponding to the current task based on the feature data set and the above-obtained basic data (real-time data). Among them, the plan customization model uses deep learning technology, such as neural networks, transformers, etc., to capture the complex relationship between tasks and resources.

[0038] For example, a deep learning model based on the Transformer architecture can be used as a large model for planning customization. The Transformer model is good at processing serialized task steps and parallel resource allocation problems, and can effectively capture the dependencies between tasks and the constraints between resources.

[0039] Among them, the input feature data set is used to plan the customization big model to capture the dependencies and constraints between tasks, resources, etc. Through training and optimization, the planning customization big model can learn the optimal planning strategy under different task types and different external environmental conditions.

[0040] By inputting the preprocessed feature data set into the large model and combining it with basic data such as the current task objectives, resource status, and external environmental factors, the planning customization model will output a preliminary plan containing detailed task allocation, time nodes, and resource allocation information.

[0041] S4. Use the pre-built multi-agent system to execute the currently generated plan in real time, and obtain the task status data of the multi-agent system during the execution of the plan.

[0042] Based on the pre-built multi-agent system, the currently generated preliminary plan can be processed in real time, in which each agent will be assigned the responsibility of task management or resource management in a specific link. The agents can exchange and coordinate information through predefined communication protocols to ensure the rationality of task decomposition and resource allocation.

[0043] Then, the task status data of the multi-agent system during the execution of the plan can be collected in real time or obtained at a set period. These task status data will serve as one of the data bases for subsequent real-time adjustment of the plan.

[0044] S5. Obtain real-time available resource data and real-time external environment data, use the plan customization model, and dynamically adjust the currently generated plan based on task status data, real-time available resource data, and real-time external environment data.

[0045] By obtaining real-time available resource data (real-time available resource data) and external environment data (real-time external environment data), and then combining these data with the real-time task status data obtained in the previous step, these data are input into the large plan customization model. The model can dynamically adjust the last generated plan (if it is the first adjustment, it is the preliminary plan) according to the latest data, and output a plan that is more in line with the current actual situation.

[0046] Based on this, complex plans can be customized in real time based on the trained multi-agent system and the plan customization model. The system can automatically generate and execute the optimal plan based on the current task status, resource availability, and external environment data. During the execution of the plan, the multi-agent system maintains continuous monitoring of the environment and task status, and dynamically adjusts the plan through the plan customization model when necessary, ultimately ensuring the smooth execution of the plan and the achievement of the goal.

[0047] The embodiments of the present application utilize a large model to learn the relationship between various factors and tasks based on historical execution data, and generate corresponding plans based on the current multiple complex factors. At the same time, a multi-agent system is used to process the current plan in real time, and the current plan is dynamically adjusted according to real-time collected status, resources, external environment and other data, thereby improving the accuracy, flexibility and real-time performance of task planning in scenarios involving multiple considerations and dynamically changing external environments.

[0048] In some possible embodiments, the multi-agent system includes a task management agent, a resource management agent, an environment monitoring agent, and a decision coordination agent; The task management agent is used to decompose tasks, assign tasks, and monitor the progress of tasks for the currently executed plan; The resource management agent is used to schedule resources, allocate resources, and update resource status for the currently executed plan; The environment monitoring agent is used to collect and analyze the current external environment data in real time; The decision-making coordination agent is used to integrate information among various agents and coordinate the behaviors among various agents.

[0049] It should be noted that the architecture of the multi-agent system can include task management agents, resource management agents, environmental monitoring agents and decision-making and coordination agents; among them, the task management agent is responsible for the decomposition, allocation and progress monitoring of tasks; the resource management agent is responsible for the scheduling, allocation and status update of resources; the environmental monitoring agent is responsible for real-time collection and analysis of external environmental data, such as weather forecasts, changes in market demand, policy trends and supply chain conditions; the decision-making and coordination agent is responsible for integrating information between agents to coordinate the actions of the agents and ensure the consistency and feasibility of the plan.

[0050] Among them, each intelligent agent can exchange information and work together through an efficient message queue or distributed communication framework to achieve rapid response to tasks and optimal allocation of resources.

[0051] Based on this, a multi-agent system is set up to automatically execute and manage the current plan in a refined manner, effectively improving the task processing efficiency and resource allocation efficiency, thereby further improving the efficiency of task planning management in complex scenarios.

[0052] In some possible embodiments, the complex task management method based on a large model may further include: S6. Iteratively train the multi-agent system based on a preset reward function and reinforcement learning algorithm to obtain an optimized and adjusted multi-agent system; The reward function consists of multiple reward items and their corresponding weight coefficients. The multiple reward items include task completion reward items, resource utilization efficiency reward items, time cost reward items, and inter-agent collaboration efficiency reward items. The task completion reward item is used to characterize the number of tasks completed by the agent, the quality of task completion and the importance of tasks; the resource utilization efficiency reward item is used to characterize the degree of resource conservation, efficiency improvement and resource sustainability; the time cost reward item is used to characterize the degree of shortening of the planned execution time and the deviation between the actual execution time and the expected time; the inter-agent collaboration efficiency reward item is used to characterize the timeliness of information exchange between agents, the effectiveness of information exchange and the efficiency of collaborative problem solving.

[0053] It should be noted that by designing a reward function, the reward function can comprehensively consider multiple dimensions such as task completion, resource utilization efficiency, time cost, and collaboration efficiency between agents, so as to quantitatively evaluate various performances in the execution of the plan. Then, reinforcement learning algorithms such as Q-learning and deep reinforcement learning are used to iteratively train the multi-agent system based on the above quantitative evaluation indicators as guidance information. Through continuous trial and error and learning, the multi-agent system can gradually learn to make better decisions in a complex and changing environment, and improve the intelligence level and robustness of plan making.

[0054] Specifically, the reward function can be composed of multiple sub-reward items. During the application process, the total reward value is obtained by weighted summation. The specific formula can be expressed as: R = w1*R_task + w2*R_resource+w3*R_time + w4 * R_collaboration Among them: R_task represents the reward item for task completion, which is quantified based on the number, quality and importance of completed tasks. This reward item can encourage agents to prioritize key tasks. R_resource represents the reward item for resource utilization efficiency, which is quantified based on the degree of resource conservation, efficiency improvement and sustainability of resource use. This reward item can encourage agents to optimize resource allocation as much as possible while ensuring task completion. R_time represents the reward item for time cost, which is quantified based on the degree of reduction in planned execution time (before and after agent optimization) and the deviation between actual and expected time. This reward item can motivate agents to improve the timeliness of planned execution. R_collaboration represents the reward item for collaborative efficiency between agents, which is quantified based on the timeliness, effectiveness of information exchange between agents and the efficiency of collaborative problem solving. This reward item can promote good cooperation between agents.

[0055] It should be noted that w1, w2, w3, and w4 are the weight coefficients of each sub-reward item, which can be dynamically adjusted according to actual task requirements and training effects.

[0056] It should be noted that reinforcement learning algorithms such as Deep Q-Network, policy gradient method, and PPO (Proximal Policy Optimization) algorithm can be used to iteratively train multi-agent systems.

[0057] In some possible embodiments, step S6, iteratively training the multi-agent system based on a preset reward function and a reinforcement learning algorithm to obtain an optimized and adjusted multi-agent system, may include: S601, initializing parameters of the state and decision-making strategy of the pre-built multi-agent system; Specifically, step S601 is to initialize the state and decision-making strategy of the multi-agent system. Each agent (task management agent, resource management agent, environment monitoring agent, decision coordination agent) is assigned a corresponding initial state according to its responsibilities, and the parameters of the decision-making strategy of each agent are randomly initialized.

[0058] S602, obtaining the experience data of the multi-agent system in the process of executing the current plan, and storing the experience data in the experience playback buffer; wherein the experience data includes the initial state, execution action, reward value and update state of each agent after executing the plan, and the reward value is calculated based on the reward function; S603, based on the experience data in the experience replay buffer, using a preset reinforcement learning algorithm to update the parameters of the decision-making strategy of the multi-agent system; It should be noted that steps S602 to S603 are executed in each training round. Specifically, each agent executes the actions in the plan according to the current decision strategy, such as task allocation, resource scheduling, environmental monitoring, etc., obtains the reward value and new state after execution, and stores the experience data of each execution (initial state, execution action, reward value and updated state) in the experience replay buffer; then, based on the data in the experience replay buffer, the parameters of the decision strategy are updated through the PPO algorithm. Among them, the amplitude (strategy update step size) used to limit the strategy update is set to ensure the stability and convergence of the strategy. The goal of iterative training is to maximize the expected value of the cumulative reward, that is, to optimize the long-term behavior performance of the agent.

[0059] S604, looping through the step of acquiring experience data (S602) and the step of updating parameters (S603) until a preset round end condition is reached (for example, reaching a maximum number of time steps or completing all tasks), so as to complete one round of training; S605, obtaining the performance index of the multi-agent system, adjusting the weight coefficient corresponding to each reward item in the reward function according to the performance index, and adjusting the model learning rate and strategy update step size of each agent according to the performance index; Specifically, during the training process, the performance indicators of the multi-agent system, such as task completion rate, resource utilization, time efficiency and collaboration efficiency, are regularly evaluated, and the weight coefficients w1, w2, w3 and w4 of the reward function are adjusted according to the evaluation results of the performance indicators. At the same time, the learning rate of the model and the step size of the strategy update are dynamically adjusted to further optimize the training effect.

[0060] S606, repeating the training steps (S602 to S604) of a single round and the step of obtaining performance indicators (S605) until a preset number of training rounds is reached or the performance indicators meet preset conditions (for example, the task completion rate exceeds a preset threshold, the resource utilization rate reaches a preset threshold, etc.).

[0061] Based on this, by combining the empirical data such as the state, action and reward of the multi-agent system during the execution of the plan as the knowledge to guide each round of training, the efficiency and accuracy of the multi-agent system training can be further improved.

[0062] In some possible embodiments, step S5, obtaining real-time available resource data and real-time external environment data, using the plan customization big model, dynamically adjusting the currently generated plan based on the task status data, the real-time available resource data and the real-time external environment data, may include: S501, real-time acquisition of task status time series data, available resource time series data and external environment time series data corresponding to the multi-agent system in the process of executing the currently generated plan; S502, inputting the task status time series data, the available resource time series data and the external environment time series data into a pre-trained time series prediction model to obtain a time series prediction result output by the time series prediction model; S503, generating plan adjustment auxiliary information corresponding to the time series prediction result; S504: Utilize the plan customization large model to dynamically adjust the currently generated plan based on task status data, real-time available resource data, real-time external environment data and plan adjustment auxiliary information.

[0063] It should be noted that in the process of dynamically adjusting the plan according to real-time data, the time series forecasting model can also be used to obtain auxiliary information for plan adjustment as one of the bases for guiding the adjustment of the plan by the large model of customized plan.

[0064] Specifically, firstly, the time series data of the multi-agent system in the process of executing the currently generated plan is collected, including task status time series data, available resource time series data, and external environment time series data. Then, these time series data are input into the trained time series prediction model to obtain the time series prediction results output by the time series prediction model. Among them, multiple time series data can be simultaneously input into the trained time series prediction model to obtain time series prediction results corresponding to multiple time series data; multiple time series data can also be input into the trained time series prediction model separately to obtain time series prediction results corresponding to various time series data.

[0065] Then, according to the preset rules, plan adjustment auxiliary information corresponding to the time series prediction results is generated. For example, if it is predicted that the resource demand will increase significantly, the resources can be deployed in advance; if it is predicted that the task progress will be delayed, the task allocation can be adjusted or the investment can be increased to speed up the progress.

[0066] Finally, the currently generated plan adjustment auxiliary information together with the task status data, real-time available resource data and real-time external environment data are input into the plan customization model, so that the plan customization model can dynamically adjust the current plan according to these data.

[0067] Based on this, by acquiring various time series data of the multi-agent system in the process of executing the plan in real time, and then using the time series prediction model to perform data prediction based on these time series data, and using the prediction results as one of the bases to guide the large model to adjust the plan in real time, the real-time and accuracy of task planning management can be further improved.

[0068] In some possible embodiments, the time series prediction model adopts a long short-term memory network model, and the training process of the time series prediction model may include: Preprocess the historical execution data to obtain time series data; the preprocessing includes time serialization and data cleaning; Divide the time series data into training, validation and test sets; Initialize and set the model parameters of the long short-term memory network model; the model parameters include the number of layers, the number of hidden units in each layer, the activation function, the optimizer, and the loss function; Iteratively train the LSTM network model using the training set, and update the model parameters of the LSTM network model based on the back propagation algorithm and the gradient descent method; The model performance evaluation index of the long short-term memory network model is obtained based on the validation set, and the model parameters and hyperparameters of the long short-term memory network model are adjusted according to the model performance evaluation index to obtain a trained time series prediction model.

[0069] Specifically, the time series prediction model can adopt the Long Short-Term Memory (LSTM) network model.

[0070] Exemplarily, in the process of training the time series prediction model, the data preparation and preprocessing steps are first performed, including: 1. Data collection: collect time series data of task status, resource availability and external environmental factors (such as weather conditions, market demand, policy changes, supply chain status, etc.) from systems such as production and manufacturing, logistics distribution, and project management. 2. Data cleaning: For missing values, a linear interpolation algorithm is used to fill them; specifically, for each missing value, two valid data points before and after it are found, and the missing value is calculated by linear interpolation; for outliers, a threshold judgment method is used, and a reasonable threshold range is set according to historical data and experience, and the data set is traversed. Data points that exceed the preset threshold range are regarded as outliers, and the outliers are replaced with the mean of their neighboring points (one or more before and after). 3. Data division: The preprocessed time series data is divided into a training set, a validation set, and a test set; the training set is used for model training, the validation set is used for model selection and parameter adjustment, and the test set is used to evaluate model performance.

[0071] Exemplarily, the network structure of the LSTM model includes an input layer, a hidden layer and an output layer, wherein the input layer is used to receive preprocessed time series data, including task status, resource availability and external environmental factors, etc.; the hidden layer is configured to include at least one layer, each hidden layer is composed of multiple LSTM units, each LSTM unit contains a forget gate, an input gate and an output gate, which are used to capture long-term dependencies in time series data; the output layer can be designed according to specific needs to output predicted task progress, resource requirements and external environmental changes, etc. Then, the parameters of the model are set, including the number of layers of the LSTM model, the number of hidden units in each layer, the activation function (such as ReLU, tanh, etc.), the optimizer (such as Adam, SGD, etc.) and the loss function (such as mean square error MSE, cross entropy, etc.).

[0072] Then, in the training phase, the LSTM model is iteratively trained using the training set, and the model parameters are optimized through the back-propagation algorithm and gradient descent method, so that the model can accurately capture the characteristics and patterns in the time series data. In the validation phase, the validation set is used to evaluate the performance of the model, such as prediction accuracy, loss value, etc., and the model parameters and hyperparameters (such as learning rate, batch size, etc.) are adjusted according to the evaluation results to improve the model performance. Finally, the trained LSTM model can be tested using the test set to evaluate its generalization ability and practical application effect.

[0073] Based on this, by using the prediction results and plan adjustment suggestions output by the trained LSTM model to provide to the big model, it can assist the big model in making more scientific and reasonable decisions.

[0074] Please refer to Figure 2 , Figure 2 FIG. 1 is a block diagram showing a complex task management device based on a large model provided by some embodiments of the present application. It should be understood that the complex task management device based on a large model is different from the above-mentioned Figure 1 Corresponding to the method embodiment, it is able to execute each step involved in the above method embodiment. The specific functions of the complex task management device based on the large model can be found in the description above. To avoid repetition, the detailed description is appropriately omitted here.

[0075] Figure 2 The complex task management device based on the large model includes at least one software function module that can be stored in a memory in the form of software or firmware or solidified in the complex task management device based on the large model, and the complex task management device based on the large model includes: The data collection module 210 is used to collect historical execution data corresponding to various task types; wherein the historical execution data includes task goal description, task execution steps, resource consumption records, time consumption information and external environmental factors; A feature extraction module 220 is used to preprocess the historical execution data and extract features from the preprocessed historical execution data to obtain a feature data set; The plan generation module 230 is used to obtain the basic data corresponding to the current task to be processed, and use the pre-trained plan customization model to generate a preliminary plan based on the feature data set and the basic data; wherein the basic data includes task objectives, task available resources and external environment data; The plan execution module 240 is used to use the pre-built multi-agent system to execute the currently generated plan in real time and obtain the task status data of the multi-agent system during the execution of the plan; The plan adjustment module 250 is used to obtain real-time available resource data and real-time external environment data, and use the plan customization model to dynamically adjust the currently generated plan based on task status data, real-time available resource data and real-time external environment data.

[0076] It can be understood that the above-mentioned device item embodiment corresponds to the method item embodiment of the present invention. The complex task management device based on a large model provided by the embodiment of the present invention can implement the complex task management method based on a large model provided by any method item embodiment of the present invention.

[0077] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method, and will not be described in detail here.

[0078] like Figure 3 As shown, some embodiments of the present application provide an electronic device 300, which includes: a memory 310, a processor 320, and a computer program stored in the memory 310 and executable on the processor 320, wherein the processor 320 reads the program from the memory 310 through a bus 330 and executes the program to implement a method of any embodiment of the complex task management method based on a large model as described above.

[0079] Processor 320 can process digital signals and can include various computing structures, such as complex instruction set computer structure, reduced instruction set computer structure, or a structure that implements a combination of multiple instruction sets. In some examples, processor 320 can be a microprocessor.

[0080] The memory 310 may be used to store instructions executed by the processor 320 or data related to the execution of instructions. These instructions and / or data may include codes for implementing some or all functions of one or more modules described in the embodiments of the present application. The processor 320 of the disclosed embodiment may be used to execute instructions in the memory 310 to implement the method shown above. The memory 310 includes a dynamic random access memory, a static random access memory, a flash memory, an optical memory, or other memory known to those skilled in the art.

[0081] Some embodiments of the present application further provide a computer-readable storage medium having a computer program stored thereon. The computer program is executed by a processor to execute the method described in the method embodiment.

[0082] Some embodiments of the present application further provide a computer program product, which, when executed on a computer, enables the computer to execute the method described in the method embodiment.

[0083] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0084] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0085] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0086] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.

[0087] The above description is only an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.

[0088] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0089] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

Claims

1. A complex task management method based on a large model, characterized in that: include: Collecting historical execution data corresponding to various task types; wherein the historical execution data includes task goal description, task execution steps, resource consumption records, time consumption information and external environmental factors; Preprocessing the historical execution data, and extracting features from the preprocessed historical execution data to obtain a feature data set; Obtain basic data corresponding to the current task to be processed, and use the pre-trained plan customization model to generate a preliminary plan based on the feature data set and the basic data; wherein the basic data includes task objectives, task available resources and external environment data; Utilize a pre-built multi-agent system to execute the currently generated plan in real time, and obtain task status data of the multi-agent system during the execution of the plan; Real-time available resource data and real-time external environment data are acquired, and the plan customization model is used to dynamically adjust the currently generated plan based on the task status data, the real-time available resource data and the real-time external environment data.

2. The complex task management method based on a large model according to claim 1 is characterized in that: The multi-agent system includes a task management agent, a resource management agent, an environment monitoring agent and a decision-making coordination agent; The task management agent is used to perform task decomposition, task allocation and task progress monitoring on the currently executed plan; The resource management agent is used to perform resource scheduling, resource allocation and resource status update on the currently executed plan; The environment monitoring agent is used to collect and analyze the current external environment data in real time; The decision-making coordination agent is used to integrate information of various agents and coordinate the behaviors among various agents.

3. The complex task management method based on a large model according to claim 1 is characterized in that: Also includes: Iteratively training the multi-agent system based on a preset reward function and a reinforcement learning algorithm to obtain an optimized and adjusted multi-agent system; The reward function is composed of multiple reward items and their corresponding weight coefficients, and the multiple reward items include a task completion reward item, a resource utilization efficiency reward item, a time cost reward item, and an inter-agent collaboration efficiency reward item; The task completion reward item is used to characterize the number of tasks completed by the agent, the quality of task completion and the importance of tasks; the resource utilization efficiency reward item is used to characterize the degree of resource conservation, efficiency improvement and resource sustainability; the time cost reward item is used to characterize the degree of shortening of the planned execution time and the degree of deviation between the actual execution time and the expected time; The inter-agent collaboration efficiency reward item is used to characterize the timeliness of information exchange, the effectiveness of information exchange, and the efficiency of collaborative problem solving between agents.

4. The complex task management method based on a large model according to claim 3 is characterized in that: The multi-agent system is iteratively trained based on a preset reward function and a reinforcement learning algorithm to obtain an optimized and adjusted multi-agent system, including: Initialize the parameters of the state and decision-making strategy of the pre-built multi-agent system; Acquire the experience data of the multi-agent system during the execution of the current plan, and store the experience data in the experience playback buffer; wherein the experience data includes the initial state, execution action, reward value and update state of each agent after the execution of the plan, and the reward value is calculated based on the reward function; Based on the experience data in the experience replay buffer, a preset reinforcement learning algorithm is used to update the parameters of the decision strategy of the multi-agent system; The steps of acquiring experience data and updating parameters are executed cyclically until a preset round end condition is reached to complete one round of training; Acquire the performance index of the multi-agent system, adjust the weight coefficient corresponding to each reward item in the reward function according to the performance index, and adjust the model learning rate and strategy update step size of each agent according to the performance index; Repeat the steps of performing a single round of training and obtaining the performance indicator until a preset number of training rounds is reached or the performance indicator meets a preset condition.

5. The complex task management method based on a large model according to claim 1 is characterized in that: The acquiring of real-time available resource data and real-time external environment data, using the plan customization large model, and dynamically adjusting the currently generated plan scheme based on the task status data, the real-time available resource data, and the real-time external environment data, includes: Acquire in real time the task status time series data, available resource time series data and external environment time series data corresponding to the multi-agent system in the process of executing the currently generated plan; Inputting the task status time series data, the available resource time series data and the external environment time series data into a pre-trained time series prediction model to obtain a time series prediction result output by the time series prediction model; Generate plan adjustment auxiliary information corresponding to the time series prediction result; The plan customization model is used to dynamically adjust the currently generated plan based on the task status data, the real-time available resource data, the real-time external environment data and the plan adjustment auxiliary information.

6. The complex task management method based on a large model according to claim 5 is characterized in that: The time series prediction model adopts a long short-term memory network model, and the training process of the time series prediction model includes: Preprocessing the historical execution data to obtain time series data; wherein the preprocessing includes time serialization and data cleaning; Dividing the time series data into a training set, a validation set, and a test set; Initialize and set the model parameters of the long short-term memory network model; wherein the model parameters include the number of layers, the number of hidden units in each layer, the activation function, the optimizer and the loss function; Iteratively training the long short-term memory network model using the training set, and updating the model parameters of the long short-term memory network model based on a back propagation algorithm and a gradient descent method; Based on the validation set, a model performance evaluation index of the long short-term memory network model is obtained, and the model parameters and hyperparameters of the long short-term memory network model are adjusted according to the model performance evaluation index to obtain a trained time series prediction model.

7. The complex task management method based on a large model according to any one of claims 1 to 6, characterized in that: The task types include at least one of production and manufacturing tasks, logistics and distribution tasks, project management tasks, resource scheduling tasks and emergency response tasks; the external environmental factors include at least one of weather conditions, market demand information, policy changes and supply chain status.

8. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor can implement the large model-based complex task management method as described in any one of claims 1 to 7 when executing the program.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the large model-based complex task management method as described in any one of claims 1 to 7 is executed.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method for managing complex tasks based on a large model as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Energy consumption prediction and management method based on large model agent

    CN119272938A

  • Inventory model training method and device based on deep reinforcement learning

    CN119398655A

  • Energy management multi-agent collaborative optimization method and system based on large model agent architecture

    CN119739071A

  • Method and platform for managing project / task intelligent objective on basis of super tree

    WO2021150039A1

Cited By

  • Action execution optimization method and device, equipment and storage medium

    CN120725054A