Adaptive scheduling method, system and device for multi-agent cooperative task and medium
By optimizing the multi-agent task decomposition and scheduling through a bidirectional task decomposition model and a real-time feedback mechanism, the problems of simple task decomposition and rigid scheduling in traditional methods are solved, thereby improving the execution efficiency and adaptability of multi-agent collaborative tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG POWER GRID CO LTD
- Filing Date
- 2026-04-14
- Publication Date
- 2026-07-10
Smart Images

Figure CN122367007A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent agent task execution technology, and in particular to an adaptive scheduling method, system, device and medium for multi-agent collaborative tasks. Background Technology
[0002] In current robot task execution scenarios, especially when dealing with complex macroscopic tasks, traditional task decomposition and scheduling methods have the following limitations: I. Limited Decomposition Mechanisms and Lack of Coordination: Existing methods typically employ a unidirectional (e.g., top-down or bottom-up) decomposition approach, making it difficult to balance the overall constraints of the macro-task with the individual capabilities of the agents. This results in unreasonable sub-tasks, failing to fully leverage the advantages of multiple agents and impacting task execution efficiency and quality.
[0003] Second, the scheduling process is rigid and lacks dynamic adaptation: During task execution, the state of the agent (such as unexpected situations, resource consumption, etc.) changes in real time. However, traditional scheduling methods cannot perceive these changes in real time and dynamically adjust the subtask division and allocation scheme, resulting in a decrease in the adaptability between the agent and the task, a decrease in execution efficiency, and may even lead to task failure. Summary of the Invention
[0004] This invention provides an adaptive scheduling method, system, device, and medium for multi-agent collaborative tasks, which can achieve efficient dynamic adaptation between target tasks and agent capabilities, thereby improving the execution efficiency and adaptability of multi-agent collaborative completion of complex tasks.
[0005] This invention provides an adaptive scheduling method for multi-agent cooperative tasks, comprising: The acquired tasks to be executed are collaboratively decomposed using a pre-defined bidirectional task decomposition model to generate an initial set of subtasks and an initial allocation scheme for agents corresponding to each initial subtask. According to the initial allocation scheme, the intelligent agent is controlled to execute the corresponding initial sub-task, and the behavioral feedback data generated during the execution process is collected in real time. The execution state and task decomposition adaptability of the intelligent agent are analyzed based on the behavioral feedback data to obtain analysis results. The initial subtask set and initial allocation scheme are then dynamically adjusted based on the analysis results to obtain an optimized subtask set and optimized allocation scheme. Based on the optimized subtask set and the optimized allocation scheme, the task execution order is optimized and scheduled, and a task execution instruction set is output to drive each intelligent agent to collaboratively complete the task to be executed according to the task execution instruction set.
[0006] This invention employs a pre-defined bidirectional task decomposition model to collaboratively decompose tasks, generating an initial set of subtasks and corresponding initial agent allocation schemes, laying a solid foundation for multi-agent collaborative execution. By analyzing the adaptability of agent execution states to task decomposition based on behavioral feedback data and dynamically adjusting the initial subtask set and allocation scheme according to the analysis results, real-time dynamic adaptation of task decomposition and agent capabilities is achieved, overcoming the limitations of traditional static scheduling methods in handling state changes during execution. Finally, based on the optimized subtask set and allocation scheme, the task execution order is optimized and scheduled, and a task execution instruction set is output to drive agents to complete tasks, thereby improving overall execution efficiency under multi-dimensional optimization objectives. Compared to existing technologies that lack dynamic adjustment mechanisms and struggle to achieve continuous task and capability adaptation, this invention effectively improves the execution efficiency and adaptability of multi-agent collaborative completion of complex tasks through a closed-loop mechanism of "decomposition-execution-feedback-adjustment-optimization."
[0007] Furthermore, the task to be executed includes macro-level task information and task association information; the task association information includes task target parameters, constraints, and agent capability attribute data; through a preset bidirectional task decomposition model, the acquired task to be executed is collaboratively decomposed to generate an initial set of sub-tasks and an initial allocation scheme for agents corresponding to each initial sub-task, specifically as follows: Using a pre-defined bidirectional task decomposition model, based on the task objective parameters and constraints, a hierarchical decomposition algorithm is employed to continuously decompose the task to be executed into multiple subtasks. The complexity of each subtask is calculated using a task complexity evaluation function. If the complexity of all subtasks does not exceed a preset complexity threshold, the decomposition stops and an initial set of subtasks is generated. Based on the initial subtask set and the agent's capability attribute data, an adaptation matrix is constructed, and task allocation is performed based on the adaptation matrix and an allocation optimization algorithm to generate an initial allocation scheme for the agent corresponding to each initial subtask.
[0008] This approach utilizes a pre-defined bidirectional task decomposition model. Based on task objective parameters and constraints, a hierarchical decomposition algorithm continuously decomposes the task to be executed into multiple subtasks. This allows for precise control of the decomposition granularity through a task complexity evaluation function, generating a highly adaptable initial set of subtasks. An adaptability matrix is constructed based on the agent's capability attribute data, and task allocation is performed using an allocation optimization algorithm. This generates an initial agent allocation scheme corresponding to each initial subtask, achieving a preliminary and accurate match between the task decomposition structure and the agent's capabilities. Through bidirectional modeling and optimized allocation, the decomposition rationality and resource adaptability of the multi-agent system in the initial task stage are effectively improved, thereby ensuring overall execution efficiency and the feasibility of load balancing.
[0009] Furthermore, the task complexity evaluation function is the product of the number of dimensions of the subtask objective function and the corresponding weight coefficient, plus the product of the subtask constraint difficulty coefficient and the corresponding weight coefficient; wherein, the weight coefficient is a pre-set parameter used to adjust the degree of influence of the number of dimensions of the objective function and the constraint difficulty coefficient on the complexity.
[0010] Furthermore, the elements in the adaptation matrix are calculated by summing the products of the agent's task processing type adaptation, execution efficiency parameter, and resource consumption coefficient with the adaptation coefficient, execution efficiency, and resource consumption of the corresponding subtask.
[0011] Furthermore, after task allocation based on the fitness matrix and the allocation optimization algorithm, the process also includes confidence evaluation of the initial allocation scheme through an intent reasoning mechanism, specifically: Using a pre-defined intent reasoning model, based on the initial allocation scheme and combining the agent's behavioral intent features and historical execution data, the intent confidence score for each agent's assigned task is output; the behavioral intent features are obtained through deep learning network analysis. When the confidence level of the intent is lower than a preset confidence threshold, the corresponding assignment task is marked as a low-confidence assignment pair.
[0012] By using a pre-defined intent reasoning model, combining the agent's historical execution data and its behavioral intent characteristics under the corresponding sub-task type to output intent confidence, and marking the allocation results based on the confidence threshold, it is possible to effectively identify the agent's true execution tendency and potential adaptability to sub-tasks in the early stages of task allocation. This avoids misallocation or inefficient allocation caused by relying solely on static capability attribute data, and further improves the stability and efficiency of overall task execution.
[0013] Furthermore, based on the behavioral feedback data, the execution state of the intelligent agent and the adaptability of the task decomposition are analyzed to obtain analysis results. Based on these results, the initial subtask set and initial allocation scheme are dynamically adjusted to obtain an optimized subtask set and optimized allocation scheme, specifically: Based on the behavioral feedback data and the agent's historical execution data, a dynamic intent deviation value is calculated using a preset intent reasoning model; When the dynamic intent deviation value exceeds the preset deviation threshold, Trigger an adaptation evaluation of the initial subtask set and initial allocation scheme to obtain task analysis results at the subtask decomposition granularity and scheme analysis results for agent scheme allocation; If the task analysis result is not suitable, the real-time complexity of the initial subtask is calculated, and the subtask is split or merged according to the real-time complexity to obtain the adjusted subtask set; otherwise, the initial subtask set is used as the adjusted subtask set. If the analysis result of the proposed scheme is deemed unsuitable, the suitability matrix is updated based on the behavioral feedback data, and the tasks are reassigned based on the updated suitability matrix and the allocation optimization algorithm to obtain an adjusted allocation scheme; otherwise, the initial allocation scheme is used as the adjusted allocation scheme. An optimized subtask set is obtained based on the adjusted subtask set, and an optimized allocation scheme is obtained based on the adjusted allocation scheme.
[0014] By triggering adaptability assessment through deviation thresholds, real-time perception and anomaly identification of the agent's execution state and task decomposition adaptability are achieved. This enables dynamic matching of task decomposition granularity and execution process, as well as real-time re-matching of agents and tasks, significantly enhancing the execution flexibility and collaborative efficiency of multi-agent systems in complex task scenarios.
[0015] Furthermore, based on the optimized subtask set and the optimized allocation scheme, the task execution order is optimized and scheduled, and a task execution instruction set is output, specifically as follows: The optimal execution order is obtained by optimizing the task execution order through a preset collaborative scheduling objective function. The collaborative scheduling objective function is the sum of the products of execution time, energy consumption, and error with their respective weight coefficients. The weight coefficients are used to adjust the importance of execution time, energy consumption, and error in the objective function. Based on the optimized subtask set and the optimized allocation scheme, and combined with the optimal execution order, a timing instruction containing execution time, subtask identifier and execution parameters is generated for each agent, and all timing instructions are integrated into a task execution instruction set for output.
[0016] This approach optimizes the task execution order by using a pre-defined collaborative scheduling objective function, achieving optimal execution order and global scheduling optimization across multiple performance metrics. It transforms the dynamically adjusted task decomposition structure and allocation scheme into precise, time-sequential executable instructions, thereby effectively reducing overall execution time and energy consumption, minimizing execution errors, and improving task completion quality.
[0017] Another embodiment of the present invention provides an adaptive scheduling system for multi-agent collaborative tasks, comprising: a task collaborative decomposition module, a task allocation optimization module, and a task scheduling execution module; The task collaborative decomposition module is used to collaboratively decompose the acquired tasks to be executed through a preset bidirectional task decomposition model, and generate an initial set of sub-tasks and an initial allocation scheme for the agent corresponding to each initial sub-task. The task allocation optimization module is used to control the agent to execute the corresponding initial sub-tasks according to the initial allocation scheme, and to collect the behavioral feedback data generated during the execution process in real time; to analyze the execution status of the agent and the adaptability of the task decomposition according to the behavioral feedback data, to obtain the analysis results, and to dynamically adjust the initial sub-task set and the initial allocation scheme according to the analysis results, so as to obtain the optimized sub-task set and the optimized allocation scheme. The task scheduling and execution module is used to optimize and schedule the task execution order according to the optimized subtask set and the optimized allocation scheme, and output a task execution instruction set to drive each intelligent agent to collaboratively complete the task to be executed according to the task execution instruction set.
[0018] This invention employs a pre-defined bidirectional task decomposition model to collaboratively decompose tasks, generating an initial set of subtasks and corresponding initial agent allocation schemes, laying a solid foundation for multi-agent collaborative execution. By analyzing the adaptability of agent execution states to task decomposition based on behavioral feedback data and dynamically adjusting the initial subtask set and allocation scheme according to the analysis results, real-time dynamic adaptation of task decomposition and agent capabilities is achieved, overcoming the limitations of traditional static scheduling methods in handling state changes during execution. Finally, based on the optimized subtask set and allocation scheme, the task execution order is optimized and scheduled, and a task execution instruction set is output to drive agents to complete tasks, thereby improving overall execution efficiency under multi-dimensional optimization objectives. Compared to existing technologies that lack dynamic adjustment mechanisms and struggle to achieve continuous task and capability adaptation, this invention effectively improves the execution efficiency and adaptability of multi-agent collaborative completion of complex tasks through a closed-loop mechanism of "decomposition-execution-feedback-adjustment-optimization."
[0019] Another embodiment of the present invention provides a terminal device, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements the steps of the adaptive scheduling method for multi-agent cooperative tasks of the present invention.
[0020] Another embodiment of the present invention provides a computer-readable storage medium item, including: a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located executes the adaptive scheduling method for multi-agent cooperative tasks of the present invention. Attached Figure Description
[0021] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating an embodiment of the adaptive scheduling method for multi-agent cooperative tasks provided by the present invention. Figure 2 This is a schematic diagram of another embodiment of the adaptive scheduling system for multi-agent collaborative tasks provided by the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0025] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0026] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0027] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.
[0028] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).
[0029] See Figure 1 To address the problem of task decomposition and scheduling in multi-agent cooperative processes in existing technologies, an embodiment of the present invention provides an adaptive scheduling method for multi-agent cooperative tasks, comprising steps S1 to S4, the specific steps of which are as follows: S1. Through a preset bidirectional task decomposition model, the acquired tasks to be executed are collaboratively decomposed to generate an initial set of subtasks and an initial allocation scheme for the agent corresponding to each initial subtask. The collaborative decomposition employs two task decomposition methods: top-down and bottom-up. Top-down decomposition breaks down the task to be executed into a preliminary sub-task hierarchy according to preset logical rules. Bottom-up decomposition integrates smaller tasks that the agent can complete upwards to form a set of sub-tasks that match the task to be executed. The initial allocation scheme is a specific scheme for allocating each initial sub-task to the corresponding agent based on the initial sub-task set and the ability attribute data of each agent. S2. According to the initial allocation scheme, control the intelligent agent to execute the corresponding initial sub-task, and collect the behavioral feedback data generated during the execution process in real time; S3. Analyze the execution state and task decomposition adaptability of the intelligent agent based on the behavioral feedback data, obtain the analysis results, and dynamically adjust the initial subtask set and initial allocation scheme based on the analysis results to obtain the optimized subtask set and optimized allocation scheme. Specifically, based on the behavioral feedback data, the current execution state of the agent is inferred, and it is analyzed whether this execution state is compatible with the initial task decomposition and initial allocation scheme; based on the analysis results, the initial subtask set and initial allocation scheme are dynamically adjusted, and the subtask division and agent allocation scheme are modified and optimized in real time to obtain an optimized subtask set and optimized allocation scheme. S4. Based on the optimized subtask set and the optimized allocation scheme, optimize and schedule the task execution order, and output the task execution instruction set to drive each intelligent agent to collaboratively complete the task to be executed according to the task execution instruction set.
[0030] Specifically, based on the optimized subtask set and the optimized allocation scheme, and taking into account factors such as the execution capabilities of each agent, the dependencies between tasks, and resource availability, and combined with a preset cooperative scheduling objective function, appropriate computing, communication, energy, and tool resources are allocated to the tasks to optimize the task execution order. The task execution instruction set is a set of instructions generated to drive each agent to execute tasks. These instructions clearly indicate when, where, and how each agent executes its assigned subtask, and serve as the basis for the robot to actually execute tasks.
[0031] This invention employs a pre-defined bidirectional task decomposition model to collaboratively decompose tasks, generating an initial set of subtasks and corresponding initial agent allocation schemes, laying a solid foundation for multi-agent collaborative execution. By analyzing the adaptability of agent execution states to task decomposition based on behavioral feedback data and dynamically adjusting the initial subtask set and allocation scheme according to the analysis results, real-time dynamic adaptation of task decomposition and agent capabilities is achieved, overcoming the limitations of traditional static scheduling methods in handling state changes during execution. Finally, based on the optimized subtask set and allocation scheme, the task execution order is optimized and scheduled, and a task execution instruction set is output to drive agents to complete tasks, thereby improving overall execution efficiency under multi-dimensional optimization objectives. Compared to existing technologies that lack dynamic adjustment mechanisms and struggle to achieve continuous task and capability adaptation, this invention effectively improves the execution efficiency and adaptability of multi-agent collaborative completion of complex tasks through a closed-loop mechanism of "decomposition-execution-feedback-adjustment-optimization."
[0032] In one embodiment, the task to be executed includes macro-level task information and task association information; the task association information includes task target parameters, constraints, and agent capability attribute data; through a preset bidirectional task decomposition model, the acquired task to be executed is collaboratively decomposed to generate an initial set of sub-tasks and an initial allocation scheme for agents corresponding to each initial sub-task, including steps S201 to S203, each step as follows: S201. Using a preset bidirectional task decomposition model, based on the task target parameters and the constraints, a hierarchical decomposition algorithm is used to continuously decompose the task to be executed into multiple subtasks. The bidirectional task decomposition model employs both top-down and bottom-up decomposition methods. Top-down decomposition starts from the overall macro-task and gradually breaks it down into smaller and more specific sub-tasks. Bottom-up decomposition starts from the agent's capabilities and executable operations, combining smaller tasks that the agent can complete to form a set of sub-tasks related to the macro-task. The two methods are combined for collaborative decomposition, breaking down the macro-task into a preliminary hierarchical structure of sub-tasks according to certain logic and rules. Simultaneously, considering the agent's capabilities and executable operations, starting from the smaller tasks that the agent can complete, it integrates upwards to form a set of sub-tasks that match the macro-task (bottom-up). Specifically, the objective function of the task T to be executed is defined as follows: in, The number of dimensions for the task objective; Indicates the first The target parameters are 3 dimensions, including task completion accuracy, execution time, and energy consumption threshold; Collecting multi-agent sets Capability attribute data for each agent The attribute vectors are as follows: in, Indicates the suitability of task processing type; Indicates execution efficiency parameters; Indicates the resource consumption coefficient; The constraints for the extraction task are as follows: in, Time constraints; Due to resource constraints; Constraints on subtask dependencies; Ultimately, a complete information model of the task to be executed is formed. The attribute vectors of all agents are represented by P(A) = P( express; Based on the objective function F(T) and constraints C(T) of the task to be executed, a hierarchical decomposition algorithm is used to break down T into n subtasks. The first layer of subtasks is set as follows. Each subtask Corresponding objective function and satisfy .
[0033] In practical applications, taking the task of inspecting all devices A today as an example: 1. Task Decomposition Layer: To complete the inspection of all devices A today, this is decomposed into inspection tasks A1, A2, and A3. 2. Operation Decomposition Layer: To complete the inspection of all devices A while meeting safety logic requirements, the A1 inspection task is further decomposed into C1 and C2 operations. 3. Robot Capability Requirements Layer: This layer defines the robot capabilities required to complete the inspection of all devices A. For the C1 operation task, robots R1, R2, and R3 can be selected. 4. Task Sequencing Layer: This layer requires meeting task priority requirements and provides the optimal task sequencing sequence. 5. Path Planning Layer: This layer further refines the process to the optimal path and generates movement tasks. Furthermore, the layers represent a hierarchical division of tasks. For example, the task of inspecting all devices A will be decomposed into inspection operation tasks and movement tasks. Different execution sequences will also add different operation tasks. The tasks in the last layer are the final execution tasks. The differences between the layers are shown in the table below. The hierarchical decomposition of the objective function corresponds to the macro-level objective function F(T)={f1,f2,…,fk} corresponding to the first-level subtasks. The objective function corresponding to each subtask is a subset of the macro-level objective function. The set of objective functions at all levels constitutes the entire macro-level objective. The hierarchical propagation of constraints corresponds to the time constraint c1, resource constraint c2, and dependency constraint c3, etc., of the macro-level task corresponding to the first-level subtasks. Subsequent levels refine certain constraints into more specific sub-constraints based on the actual decomposition.
[0034] S202. Calculate the complexity of each subtask using the task complexity evaluation function. If the complexity of all subtasks does not exceed the preset complexity threshold, stop the decomposition and generate an initial set of subtasks. The specific formula for the task complexity evaluation function is as follows: in, and These are the weighting coefficients; To constrain the difficulty coefficient, subtasks are represented. The difference values that satisfy the constraints; Let i be the complexity of the i-th subtask in the first-level subtask set; To satisfy subtasks The number of objective functions reaches the optimal number; If the complexity of all subtasks does not exceed the preset complexity threshold, that is... in, Set a preset complexity threshold; Then stop the decomposition and generate an initial set of subtasks. ; S203. Based on the initial subtask set and the agent's capability attribute data, construct an adaptation matrix, and perform task allocation based on the adaptation matrix and an allocation optimization algorithm to generate an initial allocation scheme for the agent corresponding to each initial subtask.
[0035] Among them, based on the agent's capability attribute vector Construct the fitness matrix The elements of the fitness matrix are as follows: in, The adaptation coefficient between subtask type and agent; For execution efficiency; For resource consumption; Based on the fitness matrix, the Hungarian algorithm is used for task allocation to generate an initial allocation scheme for agents corresponding to each initial subtask, as follows: This invention employs a pre-defined bidirectional task decomposition model. Based on task objective parameters and constraints, a hierarchical decomposition algorithm continuously decomposes the task to be executed into multiple subtasks. This allows for precise control of the decomposition granularity through a task complexity evaluation function, generating a highly adaptable initial set of subtasks. An adaptability matrix is constructed based on the agent's capability attribute data, and task allocation is performed using an allocation optimization algorithm. This generates an initial agent allocation scheme corresponding to each initial subtask, achieving a preliminary and accurate match between the task decomposition structure and the agent's capabilities. Through bidirectional modeling and optimized allocation, the decomposition rationality and resource adaptability of the multi-agent system in the initial task stage are effectively improved, thereby ensuring overall execution efficiency and the feasibility of load balancing.
[0036] In one embodiment, after task allocation based on the fitness matrix and allocation optimization algorithm, the method further includes evaluating the confidence of the initial allocation scheme through an intent reasoning mechanism, including steps S301 to S302, each step of which is as follows: S301. Using a preset intent reasoning model, based on the initial allocation scheme and combining the agent's behavioral intent features and historical execution data, output the intent confidence of each agent for the assigned task; the behavioral intent features are obtained through deep learning network analysis. Wherein, the intent reasoning model is The input to the intent reasoning model is the agent's historical execution data, as follows: in, Records of historical subtask execution; For intelligent agents Historical execution data; k is the number of historical task execution record entries; Subtask type; Each historical record It includes relevant execution information such as historical subtask types, execution duration, energy consumption, execution error, execution result (success / failure), and execution context (environment information); The historical execution data is processed by a deep learning network to extract temporal features, learn the agent’s behavior patterns under different types of sub-tasks, and obtain a behavior intention feature vector. Attention matching is performed between the features of the target sub-task and the intention feature vector to calculate the agent’s tendency to execute the target sub-task. Finally, the intention confidence is output through the Sigmoid function. S302. When the confidence level of the intent is lower than the preset confidence level threshold, the corresponding assignment task is marked as a low confidence assignment pair.
[0037] Specifically, the low-confidence allocation pairs and their corresponding confidence values are prioritized for subsequent adaptation assessment and dynamic adjustment. By lowering the dynamic intent deviation threshold corresponding to the low-confidence allocation pairs, they are more likely to trigger adaptation assessment, thus enabling earlier intervention and adjustment. Simultaneously, during the dynamic adjustment process, the sub-tasks and allocation schemes corresponding to the low-confidence allocation pairs will be preferentially included in the set to be adjusted and will participate in the reallocation first after the adaptation matrix is updated.
[0038] This invention uses a preset intent reasoning model to combine the agent's historical execution data and its behavioral intent features under the corresponding sub-task type to output intent confidence. Based on the confidence threshold, the allocation results are marked. This can effectively identify the agent's true execution tendency and potential adaptability to the sub-task in the early stage of task allocation, avoid misallocation or inefficient allocation caused by relying solely on static capability attribute data, and further improve the stability and efficiency of overall task execution.
[0039] In one embodiment, the execution state and task decomposition adaptability of the intelligent agent are analyzed based on the behavioral feedback data to obtain analysis results. The initial subtask set and initial allocation scheme are then dynamically adjusted based on the analysis results to obtain an optimized subtask set and optimized allocation scheme. This includes steps S401 to S405, as follows: S401. Based on the behavioral feedback data and the agent's historical execution data, calculate the dynamic intent deviation value using a preset intent reasoning model; Among them, the sampling period is set. Collect data from each intelligent agent Execute subtasks The real-time behavioral feedback data during the process is as follows: in, For task progress; Real-time energy consumption; For execution error; Based on the behavioral feedback data Compared with historical data The dynamic intent deviation value is calculated using the intent reasoning model, and the specific formula is as follows: in, The result of dynamic intent reasoning at time t includes the expected completion probability, expected execution time, expected energy consumption, and expected error. The forward propagation formula for the dynamic intent reasoning model is as follows: Inferdyn(aj, si, t) = LSTM_dynamic(Feed(aj,si,t), Hist(aj), θ_new) By fusing real-time feedback data Feed(aj,si,t)={fjt1,fjt2,fjt3} with historical data Hist(aj), a fused feature vector fused_feat is obtained. Based on the fused feature vector, the LSTM network parameters θ are updated using incremental learning or a sliding window method to obtain the updated parameters θ_new. The fused feature vector is then input into the updated LSTM network to calculate the dynamic intent inference result Inferdyn(aj,si,t) at time t. The specific calculation steps are as follows: I. Feature Coding feed_encoding = Encoder(Feed(aj,si,t)); hist_encoding = Encoder(Hist(aj)); II. Temporal Fusion: combined_input = [feed_encoding; hist_encoding]; hidden_state = LSTM(combined_input, θ_new); III. Task Matching: intent_feature = Attention(hidden_state, vec(si)); Four-output dynamic intent: Inferdyn(aj,si,t) = Sigmoid(W · intent_feature + b); S402. When the dynamic intention deviation value exceeds the preset deviation threshold, an adaptation evaluation of the initial subtask set and the initial allocation scheme is triggered to obtain the task analysis results at the subtask decomposition granularity and the scheme analysis results of the agent scheme allocation. Among them, when When this occurs, an adaptation assessment is triggered; A preset deviation threshold is set; if the dynamic intent deviation value does not exceed the preset deviation threshold, the current subtask set and allocation scheme remain unchanged, and execution continues while collecting behavioral feedback data for the next cycle. S403. If the task analysis result is not suitable, calculate the real-time complexity of the initial subtask, and split or merge the subtasks according to the real-time complexity to obtain the adjusted subtask set; otherwise, use the initial subtask set as the adjusted subtask set. The formula for calculating the real-time complexity is as follows; in, and For dynamic objectives and constraints at time step A; like If the subtask is too coarse-grained, it means the current subtask is too complex and the agent cannot complete it efficiently; by... Split into and Recalculate and assign the fitness matrix; like If the subtask is too finely decomposed, it is considered that the current subtask is too simple, which can easily lead to a lot of communication and switching overhead; by connecting adjacent subtasks... and merged into Update the subtask set S; where, This is the minimum complexity threshold; S404. If the analysis result of the proposed scheme is not suitable, the suitability matrix is updated according to the behavioral feedback data, and the task is reassigned according to the updated suitability matrix and the allocation optimization algorithm to obtain the adjusted allocation scheme; otherwise, the initial allocation scheme is used as the adjusted allocation scheme.
[0040] Among them, an unsuitable agent allocation means that the current agent is no longer suitable to execute the subtask due to a change in state (such as excessive energy consumption or hardware malfunction); according to the dynamic fitness matrix By combining updated real-time behavioral feedback data, the Hungarian algorithm is re-executed to generate an adjusted allocation scheme. Replace the original initial allocation scheme and update the intent reasoning model. ; S405. Obtain an optimized subtask set based on the adjusted subtask set, and obtain an optimized allocation scheme based on the adjusted allocation scheme.
[0041] This invention, through deviation threshold-triggered adaptability assessment, enables real-time perception and anomaly identification of the agent's execution state and task decomposition adaptability, achieving dynamic matching of task decomposition granularity and execution process, and real-time re-matching of agents and tasks, significantly enhancing the execution flexibility and collaborative efficiency of multiple agents in complex task scenarios.
[0042] In one embodiment, the task execution order is optimized and scheduled according to the optimized subtask set and the optimized allocation scheme, and a task execution instruction set is output. The method further includes steps S501 to S502, as follows: S501. Optimize the task execution order through a preset collaborative scheduling objective function to obtain the optimal execution order; the collaborative scheduling objective function is the sum of the products of execution time, energy consumption, and error with their respective weight coefficients; the weight coefficients are used to adjust the importance of execution time, energy consumption, and error in the objective function. The objective function for collaborative scheduling is constructed, and the specific formula is as follows: in, , and These are the weighting coefficients; , and These are execution time, energy consumption, and error, respectively. Based on subtask dependencies An improved genetic algorithm is used to optimize the execution order, with minimizing the total task completion time as the primary objective and balancing resource consumption as a secondary objective. The encoding method uses subtask numbers, the crossover operator uses two-point crossover, and the mutation operator uses adjacent swapping. The algorithm iterates until the convergence condition is met to obtain the optimal execution order. ; In practical applications, the optimal execution order adjusts the task execution order of the agent itself. For example, when agents A and B are working together, while agent A is waiting for agent B, other independent tasks can be executed first. S502. Based on the optimized subtask set and the optimized allocation scheme, and in conjunction with the optimal execution order, generate a timing instruction for each agent that includes the execution time, subtask identifier, and execution parameters, and integrate all timing instructions into a task execution instruction set for output.
[0043] Among them, according to and For each intelligent agent Generate timing instructions as follows: in, For the execution time, Parameters for subtask execution; Integrate all instructions into a final instruction set This drives each intelligent agent to collaboratively complete the task to be executed according to the task execution instruction set.
[0044] This invention optimizes the task execution order by using a preset collaborative scheduling objective function to obtain the optimal execution order, thereby achieving global scheduling optimization under multi-dimensional performance indicators. It transforms the dynamically adjusted task decomposition structure and allocation scheme into precise, time-sequential executable instructions, thereby effectively reducing the overall execution time and energy consumption, reducing execution errors, and improving the quality of task completion.
[0045] like Figure 2 As shown in the adaptive scheduling of multi-agent cooperative tasks, a corresponding system implementation is provided based on the above method implementation. This invention provides an adaptive scheduling system for multi-agent collaborative tasks, comprising: a task collaborative decomposition module 601, a task allocation optimization module 602, and a task scheduling execution module 603; The task collaborative decomposition module 601 is used to collaboratively decompose the acquired tasks to be executed through a preset bidirectional task decomposition model, and generate an initial set of sub-tasks and an initial allocation scheme for the intelligent agent corresponding to each initial sub-task. The task allocation optimization module 602 is used to control the agent to execute the corresponding initial sub-task according to the initial allocation scheme, and to collect the behavioral feedback data generated during the execution process in real time; to analyze the execution status of the agent and the adaptability of the task decomposition according to the behavioral feedback data, to obtain the analysis results, and to dynamically adjust the initial sub-task set and the initial allocation scheme according to the analysis results, so as to obtain the optimized sub-task set and the optimized allocation scheme. The task scheduling and execution module 603 is used to optimize and schedule the task execution order according to the optimized subtask set and the optimized allocation scheme, and output a task execution instruction set to drive each intelligent agent to collaboratively complete the task to be executed according to the task execution instruction set.
[0046] This invention employs a pre-defined bidirectional task decomposition model to collaboratively decompose tasks, generating an initial set of subtasks and corresponding initial agent allocation schemes, laying a solid foundation for multi-agent collaborative execution. By analyzing the adaptability of agent execution states to task decomposition based on behavioral feedback data and dynamically adjusting the initial subtask set and allocation scheme according to the analysis results, real-time dynamic adaptation of task decomposition and agent capabilities is achieved, overcoming the limitations of traditional static scheduling methods in handling state changes during execution. Finally, based on the optimized subtask set and allocation scheme, the task execution order is optimized and scheduled, and a task execution instruction set is output to drive agents to complete tasks, thereby improving overall execution efficiency under multi-dimensional optimization objectives. Compared to existing technologies that lack dynamic adjustment mechanisms and struggle to achieve continuous task and capability adaptation, this invention effectively improves the execution efficiency and adaptability of multi-agent collaborative completion of complex tasks through a closed-loop mechanism of "decomposition-execution-feedback-adjustment-optimization."
[0047] It is understood that the above system item embodiments correspond to the method item embodiments of the present invention, and can implement the adaptive scheduling method for multi-agent cooperative tasks provided by any of the above method item embodiments of the present invention.
[0048] It should be noted that the system embodiments described above are merely illustrative, and some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the system embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0049] For ease of description and brevity, the system embodiments of the present invention include all the implementation methods in the above-described adaptive scheduling method embodiments for multi-agent cooperative tasks, and will not be repeated here.
[0050] Based on the above embodiments of the adaptive scheduling method for multi-agent cooperative tasks, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the adaptive scheduling method for multi-agent cooperative tasks of any embodiment of the present invention.
[0051] For example, in this embodiment, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the terminal device.
[0052] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0053] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.
[0054] Based on the above-described method embodiments, another embodiment of the present invention provides a computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the adaptive scheduling method for multi-agent cooperative tasks described in any of the above-described method embodiments of the present invention.
[0055] The modules / units integrated in the device / terminal equipment, if implemented as software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0056] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. An adaptive scheduling method for multi-agent cooperative tasks, characterized in that, include: The acquired tasks to be executed are collaboratively decomposed using a pre-defined bidirectional task decomposition model to generate an initial set of subtasks and an initial allocation scheme for agents corresponding to each initial subtask. According to the initial allocation scheme, the intelligent agent is controlled to execute the corresponding initial sub-task, and the behavioral feedback data generated during the execution process is collected in real time. The execution state and task decomposition adaptability of the intelligent agent are analyzed based on the behavioral feedback data to obtain analysis results. The initial subtask set and initial allocation scheme are then dynamically adjusted based on the analysis results to obtain an optimized subtask set and optimized allocation scheme. Based on the optimized subtask set and the optimized allocation scheme, the task execution order is optimized and scheduled, and a task execution instruction set is output to drive each intelligent agent to collaboratively complete the task to be executed according to the task execution instruction set.
2. The adaptive scheduling method for multi-agent cooperative tasks as described in claim 1, characterized in that, The task to be executed includes macro-level task information and task-related information; the task-related information includes task objective parameters, constraints, and agent capability attribute data. The process involves using a pre-defined bidirectional task decomposition model to collaboratively decompose the acquired tasks to be executed, generating an initial set of subtasks and an initial allocation scheme for the agents corresponding to each initial subtask. Specifically: Using a pre-defined bidirectional task decomposition model, based on the task objective parameters and constraints, a hierarchical decomposition algorithm is employed to continuously decompose the task to be executed into multiple subtasks. The complexity of each subtask is calculated using a task complexity evaluation function. If the complexity of all subtasks does not exceed a preset complexity threshold, the decomposition stops and an initial set of subtasks is generated. Based on the initial subtask set and the agent's capability attribute data, an adaptation matrix is constructed, and task allocation is performed based on the adaptation matrix and an allocation optimization algorithm to generate an initial allocation scheme for the agent corresponding to each initial subtask.
3. The adaptive scheduling method for multi-agent cooperative tasks as described in claim 2, characterized in that, The task complexity evaluation function is the product of the number of dimensions of the subtask objective function and the corresponding weight coefficient, plus the product of the subtask constraint difficulty coefficient and the corresponding weight coefficient; wherein, the weight coefficient is a pre-set parameter used to adjust the degree of influence of the number of dimensions of the objective function and the constraint difficulty coefficient on the complexity.
4. The adaptive scheduling method for multi-agent cooperative tasks as described in claim 2, characterized in that, The elements in the fitness matrix are calculated by summing the products of the agent's task processing type fitness, execution efficiency parameter, and resource consumption coefficient with the fitness coefficient, execution efficiency, and resource consumption of the corresponding subtask.
5. The adaptive scheduling method for multi-agent cooperative tasks as described in claim 2, characterized in that, After the task allocation is performed based on the fitness matrix and the allocation optimization algorithm, the method further includes evaluating the confidence of the initial allocation scheme through an intent reasoning mechanism, specifically: Using a pre-defined intent reasoning model, based on the initial allocation scheme and combining the agent's behavioral intent features and historical execution data, the intent confidence score for each agent's assigned task is output; the behavioral intent features are obtained through deep learning network analysis. When the confidence level of the intent is lower than a preset confidence threshold, the corresponding assignment task is marked as a low-confidence assignment pair.
6. The adaptive scheduling method for multi-agent cooperative tasks as described in claim 2, characterized in that, The process involves analyzing the execution state and task decomposition adaptability of the intelligent agent based on the behavioral feedback data, obtaining analysis results, and dynamically adjusting the initial subtask set and initial allocation scheme based on the analysis results to obtain an optimized subtask set and optimized allocation scheme. Specifically: Based on the behavioral feedback data and the agent's historical execution data, a dynamic intent deviation value is calculated using a preset intent reasoning model; When the dynamic intent deviation value exceeds the preset deviation threshold, an adaptation evaluation of the initial subtask set and the initial allocation scheme is triggered to obtain the task analysis results at the subtask decomposition granularity and the scheme analysis results of the agent scheme allocation. If the task analysis result is not suitable, the real-time complexity of the initial subtask is calculated, and the subtask is split or merged according to the real-time complexity to obtain an adjusted set of subtasks. Otherwise, the initial set of subtasks is used as the adjusted set of subtasks; If the analysis result of the proposed scheme is not suitable, the suitability matrix is updated according to the behavioral feedback data, and the task is reassigned according to the updated suitability matrix and the allocation optimization algorithm to obtain the adjusted allocation scheme. Otherwise, the initial allocation scheme shall be used as the adjusted allocation scheme; An optimized subtask set is obtained based on the adjusted subtask set, and an optimized allocation scheme is obtained based on the adjusted allocation scheme.
7. The adaptive scheduling method for multi-agent cooperative tasks as described in claim 1, characterized in that, The step of optimizing the task execution order and outputting a task execution instruction set based on the optimized subtask set and the optimized allocation scheme is as follows: The optimal execution order is obtained by optimizing the task execution order through a preset collaborative scheduling objective function. The collaborative scheduling objective function is the sum of the products of execution time, energy consumption, and error with their respective weight coefficients. The weight coefficients are used to adjust the importance of execution time, energy consumption, and error in the objective function. Based on the optimized subtask set and the optimized allocation scheme, and combined with the optimal execution order, a timing instruction containing the execution time, subtask identifier, and execution parameters is generated for each agent, and all timing instructions are integrated into a task execution instruction set for output.
8. An adaptive scheduling system for multi-agent cooperative tasks, characterized in that, include: The module includes a task collaboration and decomposition module, a task allocation and optimization module, and a task scheduling and execution module. The task collaborative decomposition module is used to collaboratively decompose the acquired tasks to be executed through a preset bidirectional task decomposition model, and generate an initial set of sub-tasks and an initial allocation scheme for the agent corresponding to each initial sub-task. The task allocation optimization module is used to control the agent to execute the corresponding initial sub-tasks according to the initial allocation scheme, and to collect the behavioral feedback data generated during the execution process in real time; to analyze the execution status of the agent and the adaptability of the task decomposition according to the behavioral feedback data, to obtain the analysis results, and to dynamically adjust the initial sub-task set and the initial allocation scheme according to the analysis results, so as to obtain the optimized sub-task set and the optimized allocation scheme. The task scheduling and execution module is used to optimize and schedule the task execution order according to the optimized subtask set and the optimized allocation scheme, and output a task execution instruction set to drive each intelligent agent to collaboratively complete the task to be executed according to the task execution instruction set.
9. A terminal device, characterized in that, The system includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the adaptive scheduling method for multi-agent cooperative tasks as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, include: A stored computer program, wherein, when the computer program is executed, the device containing the computer-readable storage medium is controlled to perform the adaptive scheduling method for multi-agent cooperative tasks as described in any one of claims 1-7.