A multi-agent reinforcement learning-based multi-platform cooperative task allocation method
By constructing a multi-platform collaborative task model and using a multi-agent reinforcement learning algorithm, the problem of unreasonable task allocation in multi-platform collaborative task allocation is solved, achieving efficient and stable task allocation and resource utilization, and reducing the cost of simulation experiments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SYST OVERALL RES INST INST OF SYST ENG ACAD OF MILITARY SCI
- Filing Date
- 2026-04-24
- Publication Date
- 2026-07-10
AI Technical Summary
In multi-platform collaborative task allocation, there are problems such as unreasonable task allocation, low simulation experiment efficiency, and low resource utilization efficiency, resulting in high simulation system runtime costs, large sample requirements, and difficulty in optimizing task allocation and design schemes.
A multi-agent reinforcement learning-based approach is adopted to construct a multi-platform collaborative task model, which is transformed into a distributed partially observable Markov decision process. The multi-agent reinforcement learning algorithm is then used to determine the optimal decision and allocate tasks reasonably.
It improves the efficiency and stability of task completion strategies, enabling task objectives to be achieved with minimal resource consumption in various scenarios, thereby improving resource utilization efficiency and reducing the number of experiment runs.
Smart Images

Figure CN122367041A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of resource equipment task modeling, and in particular to a multi-platform collaborative task allocation method based on multi-agent reinforcement learning. Background Technology
[0002] Multi-platform collaborative task allocation technology is a key technology for multi-platform collaboration. It provides support for multi-platform collaborative task planning and improves the efficiency of task allocation simulation experiments through resource modeling and algorithm design.
[0003] The basic units of multi-platform collaborative task allocation are characterized by a large number of parameters, diverse types, and varied considerations. Therefore, research on multi-platform collaborative task allocation and simulation experiments faces a vast and uncertain data space. This leads to inefficient task allocation, a large sample size requirement for simulation experiment schemes related to task allocation, high simulation system runtime costs, and low experimental analysis efficiency. Therefore, it is necessary to solve the problem of multi-platform collaborative task allocation while minimizing the impact on task completion, improving resource utilization efficiency, and reducing the number of experimental runs, in order to support the optimization of task allocation and design schemes. Summary of the Invention
[0004] One objective of this application is to provide a multi-platform collaborative task allocation method based on multi-agent reinforcement learning, so as to alleviate, mitigate or eliminate the problems in related technologies.
[0005] An embodiment of the first aspect of this application provides a multi-platform collaborative task allocation method based on multi-agent reinforcement learning, comprising: constructing a multi-platform collaborative task model; determining an objective function for multi-platform collaborative task allocation; the multi-platform collaborative task model being used to allocate corresponding tasks to multiple platforms; establishing a distributed partially observable Markov decision process based on the multi-platform collaborative task model, wherein the reward function of the distributed partially observable Markov decision process is determined according to the objective function for multi-platform collaborative task allocation; using a multi-agent reinforcement learning algorithm, determining the optimal decision for multi-platform collaborative task allocation based on the reward function of the distributed partially observable Markov decision process; and allocating corresponding tasks to multiple platforms according to the optimal decision method for multi-platform collaborative task allocation.
[0006] In the technical solution of this application embodiment, a complete and realistic task model is established for the multi-platform collaborative task allocation problem. Based on this model, the multi-platform collaborative task allocation problem is transformed into a distributed partially observable Markov decision process. A multi-intelligence reinforcement learning algorithm is used to solve the multi-platform collaborative task allocation problem, determine the optimal decision, and allocate the tasks to be executed by each platform according to the optimal decision. This method exhibits good generalization performance in diverse scenarios and can effectively improve task completion results.
[0007] An embodiment of the second aspect of this application provides a multi-platform collaborative task allocation device based on multi-agent reinforcement learning, comprising: a first module configured to construct a multi-platform collaborative task model and determine an objective function for multi-platform collaborative task allocation, wherein the multi-platform collaborative task model is used to allocate corresponding tasks to multiple platforms; a second module configured to establish a distributed partially observable Markov decision process based on the multi-platform collaborative task model, wherein the reward function of the distributed partially observable Markov decision process is determined according to the objective function for multi-platform collaborative task allocation; a third module configured to use a multi-agent reinforcement learning algorithm to determine the optimal decision for multi-platform collaborative task allocation based on the reward function of the distributed partially observable Markov decision process; and a fourth module configured to allocate corresponding tasks to multiple platforms according to the optimal decision for multi-platform collaborative task allocation.
[0008] An embodiment of the third aspect of this application provides a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program that, when executed by the at least one processor, implements the method described above.
[0009] An embodiment of the fourth aspect of this application provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program implements the above-described method when executed by a processor.
[0010] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0011] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments disclosed in this application and should not be construed as limiting the scope of this application.
[0012] Figure 1 Here are flowcharts of multi-platform collaborative task allocation methods based on multi-agent reinforcement learning, as described in some embodiments of this application. Figure 2 A flowchart illustrating the optimal decision for multi-platform collaborative task allocation in some embodiments of this application; Figure 3 A schematic diagram illustrating the selection of multiple platform sequential communication targets in some embodiments of this application; Figure 4This is a schematic diagram illustrating the action delay during the task process in some embodiments of this application; Figure 5 This is a schematic diagram of the memory trajectory search and pairing process in some embodiments of this application; Figure 6 This is a schematic block diagram of a multi-platform collaborative task allocation device based on multi-agent reinforcement learning, according to some embodiments of this application. Figure 7 This is a block diagram of an exemplary computer device for some embodiments of this application. Detailed Implementation
[0013] The embodiments of the technical solution of this application will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of this application and are therefore merely examples, and should not be used to limit the scope of protection of this application.
[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0015] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly indicating the number, specific order, or primary and secondary relationship of the indicated technical features.
[0016] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0017] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0018] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), and similarly, "multiple groups" refers to two or more (including two), unless otherwise explicitly specified.
[0019] Multi-platform collaborative task allocation technology is a key technology for multi-platform collaboration. It provides support for multi-platform collaborative planning and improves the efficiency of simulation experiments through resource modeling and algorithm design.
[0020] The basic units of multi-platform collaborative task allocation are characterized by a large number of parameters, diverse types, and varied considerations. Therefore, research on multi-platform collaborative task allocation and simulation experiments faces a vast and uncertain data space. This results in a large sample size requirement for simulation experiment schemes, high simulation system runtime costs, and low experimental analysis efficiency. Therefore, it is necessary to solve the problem of multi-platform collaborative task allocation without affecting the task completion effect, improve resource utilization efficiency, and reduce the number of experiment runs, in order to support the optimization of task allocation and design schemes.
[0021] Multi-agent reinforcement learning (MAL) methods overcome the shortcomings of statistical methods for resource allocation. They are applicable regardless of sample size or the presence of obvious patterns, require minimal computation, and are convenient, typically avoiding discrepancies between quantitative and qualitative analysis results. A multi-platform collaborative task allocation method based on MML aims to minimize resource consumption. Using data samples generated from simulation experiments, MML is employed to achieve a reasonable allocation of resources as needed, resulting in appropriate task assignments for each platform.
[0022] Based on this, this disclosure proposes a multi-platform collaborative task allocation method based on multi-agent reinforcement learning. A model of multi-platform collaborative tasks is constructed based on the states of different platforms, which can reasonably and effectively simulate the dynamic process of task execution. According to this mathematical model, the multi-platform collaborative task allocation problem can be described as a distributed partially observable Markov decision process. A multi-agent reinforcement learning algorithm is used to solve this process to obtain reasonable task allocation results, improving the efficiency and stability of the task completion strategy. This allows for the execution of task objectives with minimal resource consumption in various scenarios.
[0023] Exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0024] Figure 1 This is a flowchart illustrating a multi-platform collaborative task allocation method 100 based on multi-agent reinforcement learning according to an exemplary embodiment. Figure 1 As shown, the multi-platform collaborative task allocation method 100 based on multi-agent reinforcement learning includes: Step 110: Construct a multi-platform collaborative task model and determine the objective function for multi-platform collaborative task allocation. The multi-platform collaborative task model is used to allocate corresponding tasks to multiple platforms. Step 120: Based on the multi-platform collaborative task model, establish a distributed partially observable Markov decision process, wherein the reward function of the distributed partially observable Markov decision process is determined according to the objective function of the multi-platform collaborative task allocation. Step 130: Using a multi-agent reinforcement learning algorithm, based on the reward function of a distributed partially observable Markov decision process, determine the optimal decision for multi-platform collaborative task allocation; Step 140: Assign corresponding tasks to multiple platforms according to the optimal decision-making method for multi-platform collaborative task allocation.
[0025] In the embodiments of this disclosure, multi-platform collaborative tasks require the collaboration of multiple platforms to jointly achieve a single task objective. During task execution, tasks are assigned to each platform according to the requirements of this collaborative task. For example, in the field of space protection, for a space protection task, multiple platforms could be multiple surface vessels or multiple aerial work devices. Depending on the state of an incoming abnormal target, these vessels or devices would execute corresponding response actions (such as early warning, expulsion, or obstruction) to complete the response task. In the field of rescue, for a rescue task, multiple platforms could be ambulances, fire trucks, rescue teams, etc. Depending on the state of the rescue target, these vehicles would perform corresponding rescue actions. In the field of scientific research, for a scientific expedition task, multiple platforms could be exploration equipment, research teams, etc. Depending on the state of the research target, these vehicles would perform corresponding research tasks. The multi-platform collaborative task allocation method 100 based on multi-agent reinforcement learning can be applied to different types of multi-platform collaborative tasks. Using the multi-agent reinforcement learning-based multi-platform collaborative task allocation method 100, tasks can be assigned to the multiple platforms involved in the multi-platform collaborative task.
[0026] According to embodiments of this disclosure, a model that better reflects real-world scenarios is constructed for multi-platform collaborative tasks, effectively simulating the dynamic process of task completion. Simultaneously, a multi-agent reinforcement learning algorithm is used to solve the task allocation problem in multi-platform collaborative tasks, resulting in more reasonable task allocation outcomes. This effectively improves the efficiency and stability of task execution strategies and can be effectively generalized to different task scenarios with various feature types.
[0027] The following section uses a multi-platform collaborative task, "protecting against anomalous aerial targets by multiple surface vessels," as an example to introduce the multi-agent reinforcement learning-based multi-platform collaborative task allocation method 100. In this task, multiple platforms refer to multiple surface vessels. A corresponding response strategy (such as early warning, expulsion, or obstruction) will be determined for each surface vessel, and the surface vessels will then perform the response task against the anomalous aerial targets according to the corresponding response strategy. It should be understood that this embodiment is merely an example. For tasks in other domains, the multi-agent reinforcement learning-based multi-platform collaborative task allocation method 100 can also be used to allocate corresponding tasks to the various platforms involved in the task. For example, for rescue missions, the multi-agent reinforcement learning-based multi-platform collaborative task allocation method 100 can determine rescue strategies for different platforms such as ambulances, fire trucks, and rescue teams, and each platform will execute the rescue task according to its respective rescue strategy.
[0028] Suppose there is a fleet of K surface work vessels of various types, and all K vessels can participate in protection missions. All vessels are equipped with protective equipment systems, and different types of vessels will undertake different tasks based on their respective protection responsibilities within the fleet. Multi-platform collaborative missions involve determining the appropriate strategy for each vessel and allocating tasks to them according to that strategy.
[0029] Each surface vessel is equipped with different types of protective equipment systems, each with different maximum and minimum capabilities, basic completion probabilities, and maximum quantity limits. The total number of different types of protective equipment systems on the surface vessels is W. Abnormal targets will launch unusual missions against different surface vessels from multiple angles and in a specific time sequence. These abnormal targets possess parameters such as flight speed and threat level, and after launch, they will fly to their targets along predetermined trajectories based on their approximate location.
[0030] The entire response process is discretized into S time points. At each time point, the coordinates of each anomalous target and each protective equipment will continuously change, while assuming that the coordinates of each platform, i.e., the surface work vessel, remain constant. When the relevant task constraints are met, the surface work vessel can choose to take a response action against a certain anomalous target or take no action. The anomalous target will then carry out different levels of anomalous tasks on different types of surface work vessels in the fleet. The entire protection process continues until the predetermined time ends.
[0031] The symbols used to construct the multi-platform collaborative task model are shown in the table below: Table 1: Symbol Declarations
[0032] In multi-platform collaborative task allocation problems, it is necessary to rationally allocate equipment resources to improve task completion effectiveness. In this embodiment of the protection task, it is necessary to protect our own surface vessels as much as possible, so that the damage to the surface vessels is as low as possible after the entire protection task is completed. Define decision variable x. ijk (h) indicates whether at time h the surface vessel k assigns protective equipment system i to the abnormal target j, x ijk (h) is a 0-1 variable, and its formula is:
[0033] Protective equipment system i has a certain probability p of mission completion against abnormal target j. ij , 0≤p ij ≤1, similarly, for an abnormal target j, the probability of mission completion for the waterborne operation vessel k is also q. jk , 0≤q jk ≤1. The probability ps that the surface vessel k will not be hit for a certain abnormal target j at time h. jk (h) satisfies the following calculation formula:
[0034] Define a waterborne operation vessel k with a waterborne operation vessel value coefficient v. k Therefore, the first sub-objective of the task is to maximize the total value of the remaining platforms after the task ends, which is also the total value of the remaining surface vessels in the fleet, D(X).
[0035] To reduce resource waste, rational task resource allocation can minimize the resource consumption required to complete a task. Define protective equipment system i with an equipment value coefficient c. i Therefore, the second sub-objective of the task is to minimize the resource consumption, i.e., the total resource consumption C(X) of all surface vessels in the fleet, after the task is completed.
[0036] In some embodiments, the objective function for multi-platform collaborative task allocation is the maximum value of the task completion effect, where the task completion effect indicates the difference between the total value of the remaining platforms among the multiple platforms after the task is completed and the resource consumption of the multiple platforms in completing the task. In the example protection task, the task completion effect may indicate the difference between the total value of the remaining surface vessels (i.e., the total value of the remaining platforms) and the total resource consumption (i.e., the resource consumption of the multiple platforms in completing the task) after the protection task is completed.
[0037] The objective function for the multi-platform collaborative task allocation problem is:
[0038] The mission completion effect E(X) is a positive indicator, which is to maximize the total value of the remaining platforms after the mission ends, that is, the total value of the remaining surface vessels of the fleet after the protection mission ends, D(X), while minimizing the resource consumption for completing the mission, that is, the total resource consumption of all surface vessels of the fleet during the protection mission, C(X).
[0039] Assume that at any given moment, the maximum execution capacity of a platform, such as a single surface vessel, is m actions that can be performed at most, which must satisfy the following constraints:
[0040] At any given time, for a given task object, it is typically executed by only one platform. For example, for a detected abnormal target, usually only one surface work vessel will perform the response task for that abnormal target. Multiple surface work vessels will not simultaneously perform response tasks for a single abnormal target. That is, the following constraint must be satisfied:
[0041] When a surface vessel is tasked with responding to an unusual target, it can only choose to use one type of protective equipment system, which must meet the following constraints:
[0042] Each platform has a maximum limit on the number of available resources, such as the number of protective equipment systems for each type of protective equipment on each surface work vessel. There is a maximum limit to the number of each type of protective equipment system. Once all the protective equipment in a certain type of protective equipment system has been used up, it will be impossible to use that type of protective equipment system to cope with a situation. That is, the following constraints must be met:
[0043] In the multi-platform collaborative task allocation problem, a platform will only execute a task if all task completion feasibility conditions are met. For example, a surface vessel will only use its protective equipment system to deal with an abnormal target if the task completion feasibility conditions are met. Let f be the task completion feasibility coefficient f of surface vessel k using protective equipment system i against an abnormal target at time h. ijk (h) is:
[0044] The feasibility conditions include: (1) At time h, the distance between the abnormal target j and the water work vessel k is between the maximum and minimum capability range of the protective equipment system i. When the abnormal target j is too far away or too close to the water work vessel k, the corresponding protective equipment system cannot be used to deal with the task. (2) At time h, the number of times the watercraft k is targeted is less than the upper limit H that the watercraft can withstand. After the number of times the watercraft k is targeted reaches the upper limit of the watercraft can withstand, it is severely damaged and cannot participate in any protection mission, that is, it cannot use any protection equipment system and cannot perform any response mission against any incoming abnormal target. (3) At time h, the abnormal target j is within the duty defense range of the water work vessel k. For different water work vessels, they will be located in different positions in the fleet and undertake different tasks, and have different defense ranges. (4) At time h, the abnormal target j is not targeted by the protective equipment system. There is a time window between the activation of the protective equipment system and the completion of the response task to the target, which is called the flight time window. The length of the time window is approximately the time required for the protective equipment to complete the response task to the target, and can be calculated from the respective movement speeds of the abnormal target and the protective equipment and the distance between the abnormal target j and the watercraft k. When the abnormal target j is in a certain flight time window, the outcome of the response task cannot be determined, and the watercraft cannot use the corresponding protective equipment system to respond to it.
[0045] Decision variable x ijk (h) It needs to meet the feasibility constraints for handling the task, that is, it needs to meet the following constraints:
[0046]
[0047] For step 120, in some embodiments, the distributed partially observable Markov decision process is composed of tuples.<S,U,P,r,Z,O,K,γ> The description is as follows: S is the state space, U is the action space, and P is the state transition function. Let Z be the reward function, O be the observation space, O be the observation function, and K be the number of platforms. This is the attenuation coefficient.
[0048] Based on the constructed multi-platform collaborative task model, the multi-platform collaborative task allocation problem can be transformed into a distributed partially observable Markov decision process (Dec-POMDP). This distributed partially observable Markov decision process can be represented by tuples.<S,U,P,r,Z,O,K,γ> Let S be the state space, U be the action space, P be the state transition function, r be the reward function, Z be the observation space, O be the observation function, and K be the number of platforms. This is the attenuation coefficient.
[0049] For the protection task in the example, the state s in the state space S can reflect the current situational information of the environment. In some embodiments, it can be divided into two parts according to the adversary and the player:
[0050] s en This indicates the status information of all currently abnormal targets, mainly including S p Lo represents the flight speed of each anomalous target. e La e Let S represent the longitude and latitude of each anomalous target, A represent the azimuth of each anomalous target, and T represent the damage coefficient of each anomalous target. al It includes the current status information of all platforms, i.e., the waterborne operation vessels, mainly including Lo a La a Indicates the longitude and latitude of each waterborne operation vessel, W a Indicates the remaining quantity of various types of protective equipment systems. This indicates the extent of damage to each vessel operating on the water.
[0051] Each platform's operational space The dimension is Taking a surface-to-water (STO) vessel as an example, t equals the total number of abnormal targets, w equals the number of types of protective equipment systems equipped on the STO, and an additional dimension is added to indicate that the STO does not perform any actions. The action space U uses one-hot encoding, with 1 representing an executable action and 0 representing an inactive action. At each time step, the surface-to-water vessel k... Will choose an action This combines them into a joint action vector. Under the influence of this joint action vector, the state will change according to the state transition function. Move to a new state.
[0052] In some embodiments, the reward function of the distributed partially observable Markov decision process includes a positive reward for the successful execution of the task objective, a negative reward for platform resource consumption, and a negative reward for platform damage. In decision-making processes across different domains, positive rewards can be given based on the successful completion of the task, such as when an abnormal target is successfully dealt with or a rescue target is successfully rescued. Negative rewards are given for resource consumption and platform damage, such as when protective equipment on a surface vessel is consumed, the surface vessel is damaged, or rescue equipment is consumed or damaged.
[0053] In distributed partially observable Markov decision-making processes, all platforms share a single reward function r(s,u). Based on the objective function of the multi-platform collaborative task model, the following reward function is defined:
[0054] At any given moment, when a platform's corresponding task is successfully executed, such as when an abnormal target is successfully dealt with, the environment will provide a positive reward. When the platform consumes resources or suffers damage, such as when a surface vessel launches protective equipment or is hit and destroyed by an abnormal target, the environment will provide a corresponding negative reward. Due to differences in platform and resource types, the environment will provide different negative rewards based on platform value coefficients (e.g., surface vessel value coefficients) and resource value coefficients (e.g., equipment value coefficients). When multiple events occur at a given moment, the rewards from all events will be accumulated as the reward for the current moment.
[0055] For each platform, the environment is partially observable; each platform has its own observation z, which is included in the observation space and defined by the observation function. Extract the observation z for each platform from state s. Taking a surface work vessel as an example, observation z includes information on abnormal targets detected by the surface work vessel and information on each surface work vessel in the fleet. The various attributes of the information in observation z are the same as those of the information in state s, but due to the partially observable nature, the amount of information in observation z is less than that in state s.
[0056] In some embodiments, such as Figure 2 As shown, step 130 includes: Step 210: Use a multi-agent reinforcement learning algorithm to obtain the memory trajectory of the multi-platform collaborative task allocation, wherein the memory trajectory of the multi-platform collaborative task allocation includes the action target trajectory, the new observation target trajectory, and the action effect trajectory. Step 220: Based on the memory trajectory of multi-platform collaborative task allocation, complete the reward decay and backtracking redistribution to obtain the updated reward function; Step 230: Based on the updated reward function, determine the optimal decision for multi-platform collaborative task allocation.
[0057] Due to the time-sensitive, probabilistic, and delayed nature of multi-platform collaborative task models, a multi-agent reinforcement learning algorithm (QMIX-dr) was used. Improvements were made to sequential communication for target selection, and to address reward decay and backtracking redistribution to adapt to the model's requirements. Each platform uses its own observation information and previous action information as input, returns its own action value, and sequentially determines the target for each platform's action, combining them into a joint action to change the environment. The multi-agent reinforcement learning algorithm maintains a memory trajectory at each round. At the end of the round, reward decay and backtracking redistribution are performed using this memory trajectory. The state, observations, actions, and updated rewards are stored in an experience pool, and samples are taken from the experience pool to determine the optimal decision for multi-platform collaborative task allocation.
[0058] Taking a platform as a surface-to-water work vessel as an example, since multiple surface-to-water work vessels are not allowed to perform tasks on the same target within the same phase, in the multi-agent reinforcement learning algorithm (QMIX-dr), each surface-to-water work vessel communicates its selected abnormal target sequentially. For example... Figure 3 As shown, at the same moment, surface work vessel 1 selects a target Tg1. Surface work vessel 2 will receive the target number selected by surface work vessel 1 and will not be able to select target Tg1. Similarly, subsequent surface work vessels will not be able to select targets previously determined by surface work vessels. This target communication structure design can both satisfy the constraint that surface work vessels cannot repeatedly select abnormal targets at the same time, and also allow surface work vessels with stronger protection capabilities to prioritize target selection, thereby improving the success rate of mission response. The action of a surface work vessel selecting a target at one moment will serve as the input for the surface work vessel at the next moment, and will evaluate whether to continue selecting that target or change the target.
[0059] In multi-platform collaborative task models, three unique environmental characteristics exist. First, there's the timeliness of task completion; for example, the goal is to quickly handle abnormal targets, and the longer a target exists, the closer it is to the fleet, and the greater its threat. Second, the effects of actions are highly probabilistic; for instance, whether a surface vessel's actions will produce an effect is determined probabilistically, leading to situations where the vessel performs an action but fails to alter the environment. Third, there's the delay in the effects of actions; for example, considering flight time windows, the effects of actions taken by surface vessels exhibit significant delays. Figure 4As shown, assuming the surface vessel performs a response to an abnormal target at time t1, the success of the response can only be determined several time intervals later. Furthermore, the delay in the response's effect is influenced by the distance between the intruding abnormal target and the surface vessel; the delay will vary depending on the distance of the abnormal target. In addition, the flight time for a second response to the same abnormal target will be shorter than the flight time for the first response because the target remains in flight, getting closer and closer to the convoy.
[0060] To adapt to the characteristics of the above models, the multi-agent reinforcement learning algorithm (QMIX-dr) employs a dynamically decaying reward system to facilitate rapid task processing by the platform. Regarding the probabilistic nature of action effects, QMIX-dr collects action effect information throughout the entire process to filter out actions that successfully produce effects. Regarding the delayed nature of action effects, QMIX-dr reconstructs rewards by searching for the corresponding action execution time. This functionality is achieved by acquiring memory trajectories and using a search-pairing approach.
[0061] In the multi-agent reinforcement learning algorithm (QMIX-dr), a memory trajectory is maintained for each round. The memory trajectory consists of the action target trajectory, the new observation target trajectory, and the action effect trajectory.
[0062] Action target trajectory T A It records the targets targeted by the platform's actions, such as the targets targeted by the actions performed by the surface work vessel at various times within a round. For example, at time h, the surface work vessel k performed an action. , for action Decoding allows us to determine the target of the action. If the action If no target is specified, treat it as empty. Therefore, the set of all motion targets of all vessels operating on the water at time h is:
[0063] Therefore, the set of action targets for the entire round (a total of S time points) can form the action target trajectory T. A :
[0064] New observed target trajectory T D It records the first appearance of mission objectives, such as the first appearance of each anomalous objective at various times within a round. For example, at time h, based on the current observation information o hWe can analyze the targets observed for the first time at time h, since undiscovered targets have all attributes empty. Assuming t targets are observed for the first time at time h, the set of newly observed targets at time h is:
[0065] Therefore, the set of new observed targets for the entire round (a total of S time points) can form the trajectory T of the new observed targets. D :
[0066] Motion effect trajectory T E It records information related to the action's effects, such as the action effects generated at various moments within a round. The action effect information *e* consists of three parts: the action initiator *as* (the agent), the action receiver *ar* (the target), and the action result *rs* (e.g., whether the response was successful).
[0067] The action result `rs` is a 0-1 variable, where 1 indicates the action produced an effect, and 0 indicates the action failed to produce an effect. Assuming there are `n` action effect records at time `h`, the set of action effects at time `h` is:
[0068] Therefore, the set of action effects for the entire round (a total of S time intervals) can form the action effect trajectory T. E :
[0069] The memory trajectory T of a round is composed of the action target trajectory T. A New observation target trajectory T D and motion effect trajectory T E The composition is as follows:
[0070] like Figure 5 As shown, the search process proceeds progressively backward from the final moment S. The pairing process uses the action receiver ar in each action effect e as a key to match between the action target set and the new observation target set. Each ar has a unique action target a and a new observation target d to which it is paired. A successful pairing of ar and a allows the acquisition of the action's execution time t. A A successful match between ar and d allows us to obtain the target's first observation time t. D .
[0071] Taking time t=4 as an example, from the action effect set E 4The process involves selecting action effect information one by one. If the action result rs is 0, the next action effect information is selected. If the action result rs is 1, the action receiver ar is used as the key, and the process proceeds sequentially backward from time t=4, matching is performed in the action target set A. For example, if the action receiver ar is in A... 2 If a matching action target 'a' is successfully found, the effective execution time 't' of the action is recorded. A =2, and by t A Starting from the previous time, proceeding sequentially, the action receiver ar is used as the key to match the new set of observed targets D. For example, ar is in D. 1 If a new observation target d is successfully found to match it, the time t of the first observation of the target is recorded. D =1.
[0072] For step 220, in some embodiments, reward decay includes decaying the positive reward obtained from the successful execution of the task objective according to a time decay factor, wherein the time decay factor increases over time.
[0073] At the time t of the first observation target D and the effective execution time t of the action A Subsequently, the positive reward for successfully completing the task objective will decay over time, according to the following formula:
[0074] r is the time decay coefficient. j The positive reward is earned for successfully executing a task objective, such as successfully handling an abnormal objective. The faster the platform executes tasks, the higher the positive reward will be. For example, the faster a surface vessel successfully handles an abnormal objective, the higher the reward will be, and the more time it will gain to handle other abnormal objectives in the future.
[0075] For step 220, in some embodiments, backtracking reallocation includes migrating the positive reward obtained from the successful execution of the diminished task objective to the effective execution time of the action in which the task objective was successfully executed. In one example, backtracking reallocation includes migrating the positive reward obtained from the successful handling of the diminished abnormal objective to the effective execution time of the action in which the abnormal objective was successfully handled.
[0076] Due to the interference of motion effect delays, the distribution of rewards at different times can become disordered. Therefore, as... Figure 5 As shown, after obtaining the decayed reward r' j Then, it needs to be migrated to the corresponding effective execution time t of the action. A Place.
[0077] Based on the first observation time t of the search target D The reward is decayed to meet the timeliness requirements of the execution process; for cases where the action effect is probabilistic, actions that successfully produce an effect are filtered out based on the action results rs of each action effect information; and the effective execution time t is searched. A The decayed reward is transferred to reduce the impact of the delay in action effects during execution.
[0078] The multi-agent reinforcement learning algorithm (QMIX-dr) initially explores random actions during training, gradually transitioning to the network selecting actions. It stores parameters such as state, observations, and actions for each round in an experience pool. After fixed rounds, a batch of data is randomly sampled from the experience pool and fed into the individual neural networks of each platform to obtain the action value for each platform. Using the action values and states of all platforms as input, it outputs a joint action value, and updates its parameters by setting up a completely identical target network.
[0079] In step 140, based on the optimal decision obtained, the corresponding tasks to be executed by each platform will be assigned.
[0080] Figure 6 This is a schematic block diagram illustrating a multi-platform collaborative task allocation device 600 based on multi-agent reinforcement learning according to an exemplary embodiment. Figure 6 As shown, the device 600 includes: a first module 610 configured to construct a multi-platform collaborative task model and determine the objective function for multi-platform collaborative task allocation, wherein the multi-platform collaborative task model is used to allocate corresponding tasks to multiple platforms; a second module 620 configured to establish a distributed partially observable Markov decision process based on the multi-platform collaborative task model, wherein the reward function of the distributed partially observable Markov decision process is determined according to the objective function for multi-platform collaborative task allocation; a third module 630 configured to use a multi-agent reinforcement learning algorithm to determine the optimal decision for multi-platform collaborative task allocation based on the reward function of the distributed partially observable Markov decision process; and a fourth module 640 configured to allocate corresponding tasks to multiple platforms according to the optimal decision for multi-platform collaborative task allocation.
[0081] It should be understood that Figure 6 The various modules of the device 600 shown can be connected to the reference. Figure 1 The steps in method 100 described correspond to each other. Therefore, the operations, features, and advantages described above for method 100 also apply to device 600 and its included modules. For the sake of brevity, some operations, features, and advantages will not be repeated here.
[0082] While specific functions have been discussed above with reference to specific modules, it should be noted that the functions of the modules discussed herein can be divided into multiple modules, and / or at least some functions of multiple modules can be combined into a single module. The specific actions performed by the modules discussed herein include the specific module itself performing the action, or alternatively, the specific module calling or otherwise accessing another component or module that performs the action (or performs the action in conjunction with the specific module). Therefore, a specific module performing an action can include the specific module performing the action itself and / or another module that performs the action, called or otherwise accessed by the specific module. For example, the first module 610, the second module 620, and the third module 630 described above can be combined into a single module in some embodiments.
[0083] It should also be understood that this article can describe various technologies in the general context of software and hardware components or program modules. The above regarding... Figure 6 The various modules described can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these modules can be implemented as computer program code / instructions configured to execute in one or more processors and stored in a computer-readable storage medium. Alternatively, these modules can be implemented as hardware logic / circuit. For example, in some embodiments, one or more of the first module 610, the second module 620, the third module 630, and the fourth module 640 can be implemented together in a System on Chip (SoC). The SoC may include an integrated circuit chip (which includes a processor (e.g., a Central Processing Unit (CPU), microcontroller, microprocessor, digital signal processor (DSP), etc.), memory, one or more communication interfaces, and / or one or more components of other circuitry) and may optionally execute received program code and / or include embedded firmware to perform functions.
[0084] According to one aspect of this disclosure, a computer device is provided, including a memory, a processor, and a computer program stored in the memory. The processor is configured to execute the computer program to implement the steps of any of the method embodiments described above.
[0085] According to one aspect of this disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the method embodiments described above.
[0086] In the following text, combined with Figure 7Illustrative examples describing such computer devices, non-transitory computer-readable storage media, and computer program products.
[0087] Figure 7 An example configuration of a computer device 700 that can be used to implement the methods described herein is shown.
[0088] Computer device 700 can be a variety of different types of devices. Examples of computer device 700 include, but are not limited to: desktop computers, server computers, laptop or netbook computers, mobile devices (e.g., tablet computers, cellular or other wireless phones (e.g., smartphones), notebook computers, mobile stations), wearable devices (e.g., glasses, watches), entertainment devices (e.g., entertainment appliances, set-top boxes communicatively coupled to a display device, game consoles), televisions or other display devices, automotive computers, and so on.
[0089] Computer device 700 may include at least one processor 702, memory 704, multiple communication interfaces 706, display device 708, other input / output (I / O) devices 710, and one or more mass storage devices 712 capable of communicating with each other, such as via system bus 714 or other suitable connections.
[0090] Processor 702 may be a single processing unit or multiple processing units, and all processing units may include single or multiple computing units or multiple cores. Processor 702 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operating instructions. Among other capabilities, processor 702 may be configured to acquire and execute computer-readable instructions stored in memory 704, mass storage device 712, or other computer-readable media, such as program code of operating system 716, program code of application program 718, program code of other program 720, etc.
[0091] Memory 704 and mass storage device 712 are examples of computer-readable storage media for storing instructions that are executed by processor 702 to perform the various functions described above. For example, memory 704 can generally include both volatile and non-volatile memory (e.g., RAM, ROM, etc.). Furthermore, mass storage device 712 can generally include hard disk drives, solid-state drives, removable media, including external and removable drives, memory cards, flash memory, floppy disks, optical disks (e.g., CDs, DVDs), storage arrays, network-attached storage, storage area networks, etc. Both memory 704 and mass storage device 712 can be collectively referred to herein as memory or computer-readable storage media, and can be non-transitory media capable of storing computer-readable, processor-executable program instructions as computer program code, which can be executed by processor 702 as a specific machine configured to perform the operations and functions described in the examples herein. Multiple programs can be stored on mass storage device 712. These programs include an operating system 716, one or more application programs 718, other programs 720, and program data 722, and they can be loaded into memory 704 for execution.
[0092] Although Figure 7 The modules 716, 718, 720, and 722, or portions thereof, are illustrated as being stored in memory 704 of computer device 700; however, modules 716, 718, 720, and 722 may be implemented using any form of computer-readable medium accessible by computer device 700. As used herein, “computer-readable medium” includes at least two types of computer-readable media: computer-readable storage media and communication media.
[0093] Computer-readable storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, DVD, or other optical storage devices, magnetic cassettes, magnetic tapes, disk storage devices or other magnetic storage devices, or any other non-transmission medium that can be used to store information for access by computer devices. In contrast, communication media can embody computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms. Computer-readable storage media as defined herein do not include communication media.
[0094] One or more communication interfaces 706 are used for exchanging data with other devices, such as via a network, direct connection, etc. Such communication interfaces can be one or more of the following: any type of network interface (e.g., a network interface card (NIC)), wired or wireless (such as IEEE 802.11 Wireless LAN (WLAN)) wireless interface, Wi-MAX interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth™ interface, Near Field Communication (NFC) interface, etc. Communication interface 706 can facilitate communication across a variety of network and protocol types, including wired networks (e.g., LAN, cable, etc.) and wireless networks (e.g., WLAN, cellular, satellite, etc.), the Internet, etc. Communication interface 706 can also provide communication with external storage devices (not shown), such as storage arrays, network-attached storage, storage area networks, etc.
[0095] In some examples, a display device 708, such as a monitor, may be included for displaying information and images to the user. Other I / O devices 710 may be devices that receive various inputs from the user and provide various outputs to the user, and may include touch input devices, gesture input devices, cameras, keyboards, remote controls, mice, printers, audio input / output devices, and so on.
[0096] The technologies described herein can be supported by these various configurations of computer device 700, and are not limited to specific examples of the technologies described herein. For example, the functionality can also be implemented wholly or partially on a “cloud” using a distributed system. A cloud includes and / or represents a platform for resources. The platform abstracts the underlying functionality of the cloud’s hardware (e.g., servers) and software resources. Resources may include applications and / or data that can be used when performing computational processing on a server remote from computer device 700. Resources may also include services provided via the Internet and / or via subscriber networks such as cellular or Wi-Fi networks. The platform can abstract resources and functionality to connect computer device 700 to other computer devices. Therefore, the implementation of the functionality described herein can be distributed throughout the cloud. For example, the functionality can be implemented partly on computer device 700 and partly through a platform that abstracts the functionality of the cloud.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and they should all be covered within the scope of the claims and specification of this application. In particular, as long as there is no structural conflict, the various technical features mentioned in the embodiments can be combined in any way. This application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A multi-platform collaborative task allocation method based on multi-agent reinforcement learning, characterized in that, include: A multi-platform collaborative task model is constructed, and an objective function for multi-platform collaborative task allocation is determined. The multi-platform collaborative task model is used to allocate corresponding tasks to multiple platforms. Based on the multi-platform collaborative task model, a distributed partially observable Markov decision process is established, wherein the reward function of the distributed partially observable Markov decision process is determined according to the objective function assigned by the multi-platform collaborative task. Using a multi-agent reinforcement learning algorithm, based on the reward function of the distributed partially observable Markov decision process, the optimal decision for multi-platform collaborative task allocation is determined; and The corresponding tasks are assigned to the multiple platforms according to the optimal decision-making method for multi-platform collaborative task allocation.
2. The method as described in claim 1, characterized in that, The objective function for multi-platform collaborative task allocation is the maximum value of the task completion effect, wherein the task completion effect indicates the difference between the total value of the remaining platforms among the multiple platforms after the task is completed and the resource consumption of the multiple platforms in completing the task.
3. The method as described in claim 1, characterized in that, The distributed partially observable Markov decision process consists of tuples<S,U,P,r,Z,O,N,γ> The description is as follows: S is the state space, U is the action space, and P is the state transition function. Let Z be the reward function, O be the observation space, O be the observation function, and N be the number of platforms. This is the attenuation coefficient.
4. The method according to any one of claims 1-3, characterized in that, The reward function of the distributed partially observable Markov decision process includes a positive reward for the successful execution of the task objective, a negative reward for the platform consuming resources, and a negative reward for the platform being damaged.
5. The method as described in claim 1, characterized in that, The method of using a multi-agent reinforcement learning algorithm to determine the optimal decision for multi-platform collaborative task allocation based on the reward function of the distributed partially observable Markov decision process includes: The multi-agent reinforcement learning algorithm is used to obtain the memory trajectory of multi-platform collaborative task allocation, wherein the memory trajectory of multi-platform collaborative task allocation includes the action target trajectory, the new observation target trajectory, and the action effect trajectory. Based on the memory trajectory of the multi-platform collaborative task allocation, the reward decay and backtracking redistribution are completed to obtain the updated reward function; Based on the updated reward function, the optimal decision for the multi-platform collaborative task allocation is determined.
6. The method as described in claim 5, characterized in that, The reward decay includes attenuating the positive reward obtained from the successful execution of the task objective according to a time decay coefficient, wherein the time decay coefficient increases as time goes on.
7. The method as described in claim 6, characterized in that, The backtracking redistribution includes migrating the diminished positive reward obtained from the successful execution of the task objective to the effective execution time of the action in which the task objective was successfully executed.
8. A multi-platform collaborative task allocation device based on multi-agent reinforcement learning, comprising: The first module is configured to construct a multi-platform collaborative task model and determine the objective function for multi-platform collaborative task allocation. The multi-platform collaborative task model is used to allocate corresponding tasks to multiple platforms. The second module is configured to establish a distributed partially observable Markov decision process based on the multi-platform collaborative task model, wherein the reward function of the distributed partially observable Markov decision process is determined according to the objective function assigned by the multi-platform collaborative task. The third module is configured to use a multi-agent reinforcement learning algorithm to determine the optimal decision for multi-platform collaborative task allocation based on the reward function of the distributed partially observable Markov decision process; and The fourth module is configured to assign corresponding tasks to the multiple platforms based on the optimal decision of the multi-platform collaborative task allocation.
9. A computer device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores a computer program that, when executed by the at least one processor, implements the method according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-7.