Communication perception task allocation method and device based on multi-objective asynchronous strategy
By optimizing the communication strategy of the drone swarm to an asynchronous constraint and dispersed part, Markov decision-making process can be observed, and the communication strategy is optimized using the multi-objective coupled PPO method, the problems of large communication overhead and difficult synchronization decision-making in the drone swarm are solved, and efficient task allocation and communication are achieved.
Patent Information
- Application Number
- CN202510512193.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The distributed task allocation method in the existing drone cluster is difficult to achieve efficient communication and synchronous decision-making under actual conditions, resulting in large and latency communication overhead, and traditional multi-agent reinforcement learning methods are difficult to optimize multiple goals.
The multi-objective asynchronous strategy is adopted to transform the communication strategy optimization problem of the multi-agent system into the asynchronous constraint dispersed part of the observable Markov decision-making process, build an asynchronous constraint multi-agent reinforcement learning environment, optimize the communication strategy using the multi-objective coupled PPO method, and centrally optimized under bandwidth constraints through the Lagrangian slack method to realize the asynchronous communication strategy.
While reducing bandwidth overhead and task conflicts, communication efficiency and allocation reliability are optimized, and multi-objective balance is achieved.
Smart Images

Figure CN120029326B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of task allocation, and particularly to a communication-aware task allocation method and device based on a multi-objective asynchronous strategy. Background Art
[0002] Distributed task allocation in a drone swarm is very sensitive to excessive bandwidth requirements and frequent inter-drone communication. Combining reinforcement learning with traditional distributed task allocation has shown great potential in improving algorithm performance and optimizing communication. However, existing research relies on ideal bandwidth assumptions and unrealistic time synchronization, making it impractical to train and validate under actual conditions.
[0003] With the breakthrough of artificial intelligence technology, combining reinforcement learning with traditional distributed task allocation methods has shown great potential in improving algorithm performance and optimizing communication. Existing research is difficult to implement on actual drones for the following reasons: 1) Communication strategy optimization methods based on multi-agent reinforcement learning assume that all drones execute decisions synchronously to obtain consistent rewards. However, in a real-world drone swarm, due to differences in performance, sensor configurations, and processing capabilities, drones cannot execute decisions synchronously. Frequent communication and synchronization times result in significant communication overhead and latency, making synchronous multi-agent reinforcement learning (MARL) training challenging. 2) In the distributed task allocation problem, the group must optimize multiple objectives simultaneously. It is necessary to reduce the task conflict rate between drones and minimize the communication overhead of the entire system as much as possible, which makes it difficult to apply traditional multi-agent reinforcement learning methods that only optimize a single objective. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a communication-aware task allocation method and device based on a multi-objective asynchronous strategy.
[0005] A communication-aware task allocation method based on a multi-objective asynchronous strategy, which is applicable to a drone cluster to implement communication-aware task allocation; each drone in the drone cluster deploys an agent, and a multi-agent system is formed by the drone cluster. The method includes:
[0006] According to the bandwidth constraint, transform the communication strategy optimization problem in the distributed task allocation algorithm of the multi-agent system into an asynchronous constraint decentralized partially observable Markov decision process.
[0007] According to the asynchronous constraint decentralized partially observable Markov decision process, construct an asynchronous constraint multi-agent reinforcement learning environment considering the communication process.
[0008] According to the asynchronous-constrained multi-agent reinforcement learning environment, the multi-objective asynchronous communication strategy in the distributed task allocation method is optimized using the multi-objective coupled PPO method to obtain the final multi-objective asynchronous communication strategy; the multi-objective coupled PPO method refers to the method of extending the Lagrangian relaxation method to the multi-agent proximal policy optimization algorithm, asynchronously collecting data in a distributed manner, and then using the Lagrangian function to centrally optimize the communication strategy under the given bandwidth demand constraint to achieve the simultaneous optimization of multiple objectives and obtain the final multi-objective asynchronous communication strategy.
[0009] Each drone adopts the final multi-objective asynchronous communication strategy to complete the communication perception task allocation of the drone cluster.
[0010] A communication perception task allocation device based on a multi-objective asynchronous strategy, which is applicable to a drone cluster to achieve communication perception task allocation; each drone in the drone cluster deploys an agent, and the drone cluster constitutes a multi-agent system. The device includes:
[0011] A problem modeling module for transforming the communication strategy optimization problem in the distributed task allocation algorithm of the multi-agent system into an asynchronous-constrained decentralized partially observable Markov decision process according to the bandwidth constraint.
[0012] An asynchronous-constrained multi-agent reinforcement learning environment construction module for constructing an asynchronous-constrained multi-agent reinforcement learning environment considering the communication process according to the asynchronous-constrained decentralized partially observable Markov decision process.
[0013] A multi-objective asynchronous communication strategy optimization module for optimizing the multi-objective asynchronous communication strategy in the distributed task allocation method using the multi-objective coupled PPO method according to the asynchronous-constrained multi-agent reinforcement learning environment to obtain the final multi-objective asynchronous communication strategy; the multi-objective coupled PPO method refers to the method of extending the Lagrangian relaxation method to the multi-agent proximal policy optimization algorithm, asynchronously collecting data in a distributed manner, and then using the Lagrangian function to centrally optimize the communication strategy under the given bandwidth demand constraint to achieve the simultaneous optimization of multiple objectives and obtain the final multi-objective asynchronous communication strategy.
[0014] A communication perception task allocation module based on the final multi-objective asynchronous communication strategy for each drone to complete the communication perception task allocation of the drone cluster.
[0015] The above communication-aware task allocation method and device based on a multi-objective asynchronous strategy. The method transforms the communication strategy optimization problem in the distributed task allocation algorithm of a multi-agent system into an asynchronous-constrained decentralized partially observable Markov decision process according to bandwidth constraints; constructs an asynchronous-constrained multi-agent reinforcement learning environment considering the communication process according to the asynchronous-constrained decentralized partially observable Markov decision process; optimizes the multi-objective asynchronous communication strategy in the distributed task allocation method using the multi-objective coupled PPO method to obtain the final multi-objective asynchronous communication strategy; and each unmanned aerial vehicle (UAV) uses the final multi-objective asynchronous communication strategy to complete the communication-aware task allocation for the UAV cluster. This method extends the concept of constrained reinforcement learning to the multi-agent proximal policy optimization algorithm using the multi-objective coupled PPO method, and asynchronously collects data in a distributed manner; uses the Lagrangian function to centrally optimize the communication strategy under the given bandwidth demand constraints, balances the communication demand and task conflicts, realizes the simultaneous optimization of multiple objectives, reduces the bandwidth overhead and minimizes the task conflicts, and achieves a better balance between communication efficiency and allocation reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic flow chart of the communication-aware task allocation method based on a multi-objective asynchronous strategy in one embodiment;
[0017] Figure 2 It is a schematic diagram of synchronous and asynchronous training data processing in another embodiment, where Figure 2 (a) is a schematic diagram of synchronous training data processing, Figure 2 (b) is a schematic diagram of asynchronous training data processing based on concatenation;
[0018] Figure 3 It is a schematic diagram of the verification experiment results in another embodiment, where Figure 3 (a) to Figure 3 (c) are schematic diagrams of the task conflict rate, bandwidth demand, and average number of communications in the scenario of 30 UAVs respectively, Figure 3 (d) to Figure 3 (f) are schematic diagrams of the task conflict rate, bandwidth demand, and average number of communications in the scenario of 40 UAVs respectively. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0020] In one embodiment, as Figure 1As shown, a communication-aware task allocation method based on a multi-objective asynchronous strategy is provided. This method is applicable to the implementation of communication-aware task allocation in a UAV swarm. Each UAV in the UAV swarm deploys an agent, and the UAV swarm constitutes a multi-agent system. The method includes the following steps:
[0021] Step 100: According to the bandwidth constraint, transform the communication strategy optimization problem in the distributed task allocation algorithm of the multi-agent system into an asynchronous constraint decentralized partially observable Markov decision process.
[0022] Specifically, a multi-objective asynchronous multi-agent reinforcement learning method is used to solve the communication-aware task allocation problem based on the multi-objective asynchronous strategy. Formalize the asynchronous allocation algorithm of integrated communication into an asynchronous partially observable Markov decision process, and extend the original learning objective to adapt to asynchronous requirements.
[0023] Take the asynchronous constraint decentralized partially observable Markov decision process (ACDEC-POMDP) as the model of this problem, which is defined as a tuple , where represents n the set of S agents, and represent the observation value and action set of all agents at time step t respectively. is the state transition function. R represents the reward function, which provides feedback from the environment according to the action. γ ∈[0,1] represents the discount factor. represents multiple auxiliary constraints, which are used to limit the update of the agent policy.
[0024] Bron-Kerbosch features (BK features) and value of message features (VoM features) are part of the observation, and at the same time, a normalization step is also included. On this basis, channel access features are proposed, which calculate the channel access priority and enable the agent to focus on the number of times it accesses the communication channel. The channel access priority can be calculated through the arctangent function, which shows a significant change at the maximum number of transmissions and can quickly form an ordered and continuous abandonment sequence for the agent. The specific calculation formula is as follows:
[0025] ;
[0026] where the adjustment factor represents the maximum number of transmissions in a competition period, and β is the number of transmissions of the agent in the channel.
[0027] The adaptive gating mechanism among agents is an action, and the shared reward reflects the change of the global task conflict.
[0028] Step 102: According to the asynchronous constrained partially observable Markov decision process, construct an asynchronous constrained multi-agent reinforcement learning environment considering the communication process.
[0029] Step 104: According to the asynchronous constrained multi-agent reinforcement learning environment, use the multi-objective coupled PPO method to optimize the multi-objective asynchronous communication strategy in the distributed task allocation method to obtain the final multi-objective asynchronous communication strategy; the multi-objective coupled PPO method refers to the method of extending the Lagrangian relaxation method to the multi-agent proximal policy optimization algorithm, asynchronously collecting data in a distributed manner, and then using the Lagrangian function to centrally optimize the communication strategy under the given bandwidth requirement constraint to achieve the simultaneous optimization of multiple objectives and obtain the final multi-objective asynchronous communication strategy.
[0030] Specifically, the performance of the distributed task allocation method may be affected by the actual communication process. Now introduce the key elements of ACDEC-POMDP to solve the communication strategy learning problem in the asynchronous algorithm. To enable agents to utilize the state information of other agents to achieve better policy collaboration, a standard learning paradigm widely used in multi-agent reinforcement learning is adopted, namely centralized training and decentralized execution (CTDE). This method is divided into two stages: the centralized training stage and the decentralized execution stage. In the centralized training stage, the centralized evaluation network evaluates the global state and processes the data at each time step. In the decentralized execution stage, each agent obtains its local observation results at different global times and independently executes its actions.
[0031] In the distributed task allocation algorithm, the optimization of the communication strategy can be regarded as a multi-objective optimization problem, aiming to maximize the global task reward of the agents while minimizing the communication overhead. To solve this problem, the concept of constrained reinforcement learning is extended to the multi-agent proximal policy optimization algorithm, and the multi-objective coupled PPO method is proposed. This method asynchronously collects data in a distributed manner; then, the Lagrangian function is used to centrally optimize the communication strategy under the given bandwidth requirement constraint to achieve the simultaneous optimization of multiple objectives.
[0032] Step 106: Each unmanned aerial vehicle adopts the final multi-objective asynchronous communication strategy to complete the communication sensing task allocation of the unmanned aerial vehicle cluster.
[0033] In the above communication-aware task allocation method based on the multi-objective asynchronous strategy, the method transforms the communication strategy optimization problem in the distributed task allocation algorithm of the multi-agent system into an asynchronous constrained decentralized partially observable Markov decision process according to the bandwidth constraint; constructs an asynchronous constrained multi-agent reinforcement learning environment considering the communication process according to the asynchronous constrained decentralized partially observable Markov decision process; optimizes the multi-objective asynchronous communication strategy in the distributed task allocation method by using the multi-objective coupled PPO method according to the asynchronous constrained multi-agent reinforcement learning environment to obtain the final multi-objective asynchronous communication strategy; and each unmanned aerial vehicle adopts the final multi-objective asynchronous communication strategy to complete the communication-aware task allocation of the unmanned aerial vehicle cluster. This method uses the multi-objective coupled PPO method to reduce the bandwidth overhead and minimize the task conflict simultaneously, achieving a better balance between communication efficiency and allocation reliability.
[0034] In one embodiment, the expression of the tuple of the asynchronous constrained decentralized partially observable Markov decision process in step 100 is:
[0035] ;
[0036] where represents the tuple of the asynchronous constrained decentralized partially observable Markov decision process, represents n the set of S agents, and respectively represent the observation value and the action set of all agents at time step t , represents the state transition function, represents the new global state, R represents the reward function, γ ∈[0,1] represents the discount factor, represents multiple auxiliary constraints for restricting the update of the agent policy, represents the th auxiliary constraint, .
[0037] The asynchronous constrained decentralized partially observable Markov decision process allows agents to make decisions asynchronously. The goal of each agent is to learn a policy while maximizing the expected discounted return and satisfying all auxiliary constraint conditions; the expression of the auxiliary constraint conditions is:
[0038] ;
[0039] where represents the auxiliary constraint function, represents the agentj At time step t the obtained reward, denotes the discount factor at time step t .
[0040] Specifically, on the one hand, the asynchronous constrained decentralized partially observable Markov decision process allows agents to make decisions asynchronously. At each global time step , a subset of agents needs to make decisions, while other agents do not need to make decisions due to environmental reasons. The set of observations is obtained through the function . Each agent in j selects its action according to its observation , where the action set is . Through the state transition function a new environmental state can be obtained. After the set of agents r executes its policy, the reward provided by the environment can be obtained through the reward function
[0041] On the other hand, the goal of each agent j is to learn a policy that aims to maximize the expected discounted return and satisfy all the auxiliary constraints as shown in the expressions of the above auxiliary constraints.
[0042] In one embodiment, in the asynchronous constrained decentralized partially observable Markov decision process in step 100: the action of agent j at its local time step is defined as a communication strategy based on a gating mechanism, enabling the agent to adaptively select the timing of data transmission, thereby coordinating communication to improve algorithm performance; the global shared reward at the local time step j of agent is defined as the sum of the changes in the average task conflict numbers of all agents; the observations include: Bron-Kerbosch features, message value features, and channel access features.
[0043] In one embodiment, step 104 includes: each agent independently and asynchronously collects its interaction data in an asynchronous constrained multi-agent reinforcement learning environment. During training, the interaction data from different agents is concatenated into trajectories according to the time relationship; the interaction data includes the global state, local observation, action, and reward; among them, the action is a communication strategy based on a gating mechanism; on the basis of the original MAPPO algorithm, in order to ensure that the gating mechanism strategy meets the constraint conditions, Lagrange multipliers are introduced and a Lagrangian function is defined; the expression of the Lagrangian function is:
[0044] ;
[0045] Among them, represents the Lagrangian function, represents the Lagrange multiplier, represents the Actor network parameters of agent j , represents the environmental reward at the current time step, represents the bandwidth penalty threshold, represents agent j at the t th time slot, the total amount of data to be sent, represents agent j 's action at the current time step, represents the time step t when the discount factor.
[0046] The Lagrangian function is used to centrally optimize the communication strategy under the given bandwidth demand constraint, realizing the simultaneous optimization of multiple objectives, and obtaining the final multi-objective asynchronous communication strategy.
[0047] Specifically, in the centralized training - decentralized execution framework, it is necessary to collect training data from each agent, including the interaction data between the agent and the environment, such as the global state, local observation, action, reward, etc. Then, these data are connected into trajectories in chronological order for centralized training. In the synchronous training scenario, all agents make decisions and collect interaction data simultaneously. However, in the asynchronous training scenario, some agents may have difficulty making decisions within the same global time step, which brings challenges to the processing and training of interaction data. In the asynchronous distributed task allocation algorithm, the communication strategy mainly affects the subsequent calculation process and the number of task conflicts in the next round of the algorithm. In other words, the decision made by an agent at a certain time step only affects the environmental reward before the next time step. Therefore, the concatenation method widely used in asynchronous multi-agent learning is adopted to process asynchronous data. Under the concatenation method, each agent independently and asynchronously collects its interaction data, and during training, the data from different agents is concatenated into trajectories according to the time relationship. As Figure 2As shown, data from Agent 1 and Agent 2 at different time steps are concatenated based on their time relationship, and their respective trajectories are as follows:
[0048] ;
[0049] .
[0050] On the one hand, since the communication strategies executed by agents in the distributed task allocation algorithm have little impact on the next calculation stage, this mechanism can correctly evaluate the impacts of different agent strategies. On the other hand, the concatenation method is more flexible and stable in engineering implementation, making it easier to extend the algorithm to real network environments for application. Synchronous training data processing is as shown in Figure 2 (a), and asynchronous training data processing based on concatenation is as shown in Figure 2 (b).
[0051] By splicing different training trajectory data to supplement the global state data, stable policy training is achieved under the CTDE framework.
[0052] In one embodiment, the Lagrangian function is used to centrally optimize the communication strategy under the given bandwidth demand constraint to achieve simultaneous optimization of multiple objectives, and the final multi-objective asynchronous communication strategy is obtained, including: defining the dual function of the original problem as:
[0053] ;
[0054] For the purpose of finding the optimal policy parameters that satisfy , when the Lagrange multiplier is a fixed value at time step t , each agent constructs a new reward to train its policy; the new reward is:
[0055] ;
[0056] where, represents the new reward of agent j at time step t, represents the reward of agent j at time step t, represents the Lagrange multiplier at time step t , represents the total amount of data that agent j needs to send in the t th time slot, represents the action of agent j at the current time step.
[0057] The training gradient is calculated by minimizing the mean squared error between the state value function and the new reward, and the Critic network parameters are updated according to this gradient. The update expression for the Critic network parameters is:
[0058] ;
[0059] where, are the Critic network parameters, is the state value function.
[0060] For the Actor network, the policy advantage value is defined, and the training gradient for maximizing the policy advantage value is calculated, and the Actor network parameters are updated according to this gradient. The update method for the Actor network parameters is:
[0061] ;
[0062] ;
[0063] ;
[0064] where, represents the Actor network parameters, represents the advantage of the state-action pair, represents the weighted advantage of the state-action pair, represents the weighting factor, represents the hyperparameter.
[0065] In one embodiment, the expression for the weighting factor is:
[0066] ;
[0067] where, represents the KL divergence, and respectively represent the policy distributions of the agent before and after the update.
[0068] Specifically, the action of the agent j at its local time step is defined as . is defined as a communication policy based on a gating mechanism, enabling the agent to adaptively select the timing of data transmission, thereby coordinating communication to improve the algorithm performance. At time step , the agent j inputs its local observation into the Actor network to obtain the communication policy , which determines whether to send a message. The agentj At time step t the global state is defined as , which is equivalent to the joint observations of all agents . In an asynchronous learning scenario, assume that at the global time step t , the assignment algorithm of agent j has been completed and a decision must be made, while the assignment algorithm of agent has not been completed and thus no decision can be made. For the global state , the observation of agent is assigned the zero vector, while is assigned the current observation vector, which means that only the observations of the agents that have made decisions are considered valid. The global state and the next global state are obtained through the state transition function . At the local time step j of agent , its reward is defined as . A cooperation relationship is introduced to coordinate communication and define the reward function. This can be achieved by constructing a common goal among agents based on a gated communication strategy. The common goal of the communication strategy should minimize the global task conflicts within a fixed period. In addition, the change in the task conflicts of a single agent is affected by the communication strategies of other agents and cannot fully reflect the effectiveness of communication. Therefore, the changes in the average task conflict numbers of all agents are added together to form a global shared reward:
[0069] ;
[0070] where represents the change in the average task conflict number of agent within the time step j , denoted as , represents the average task conflict number between agent and other agents at the beginning of the time step j , , where represents the task list of agent j , represents the task list of another agent k . The more task conflicts there are among the drones, the larger the value of
[0071] In one embodiment, the optimal policy parameters and the Lagrange multipliers The process includes: transforming the problem of finding the saddle point of the dual function and the optimal Lagrange multiplier into: ;
[0072] ;
[0073] wherein, represents the optimal Lagrange multiplier.
[0074] By taking the derivative of the Lagrangian function, update the Lagrange multiplier at t time; t The update expression of the Lagrange multiplier at
[0075] time is:
[0076] ;
[0077] wherein, represents t the Lagrange multiplier at +1 time, t represents the Lagrange multiplier at
[0078] time, and represents a hyperparameter.
[0079] ;
[0080] wherein, represents the penalty value network parameter, and represents the penalty value network.
[0081] Calculate the action penalty value of each agent at different time steps according to the penalty value network, and update the Lagrange multiplier according to the action penalty value; the update expression of the Lagrange multiplier is:
[0082] ;
[0083] wherein, represents the j action penalty value of agent t at time step
[0084] Specifically, based on the bandwidth penalty constraint, the communication strategy optimization problem in the distributed task allocation algorithm becomes the problem of solving an asynchronous, constrained Markov Decision Process. Its formulaic expression is as follows:
[0085] ;
[0086] Among them, represents the bandwidth penalty threshold.
[0087] To solve this problem, the Lagrangian relaxation method is extended to the multi-agent proximal policy optimization (MAPPO) algorithm, and a multi-objective coupled PPO method (abbreviation: MOC-PPO) is proposed for training multi-objective asynchronous policies. The original MAPPO trains the shared Actor network and value network in a centralized training and distributed execution manner. Each agent j utilizes its local observation to select an action , which is then used for the distributed execution of different policies. This method helps agents identify cooperative relationships and resolve potential conflicts between policies by sharing state and reward information among agents. On this basis, to ensure that the gated mechanism policy obtained by training satisfies the constraints, the Lagrange multiplier is introduced, and the Lagrangian function is defined as shown in the above Lagrangian function expression.
[0088] The dual problem of the original problem can be defined as: , assuming is fixed at time step t , define . To find the optimal policy parameters such that . Therefore, when is fixed, each agent trains its policy by constructing a new reward . Thus, the centralized Critic network is used to estimate the state value function . The training gradient is calculated by minimizing the mean squared error between the state value function and the reward , and this gradient is used to update the Critic network parameters , and the Critic network parameters are updated using the above update expression for the Critic network parameters.
[0089] For the Actor network, its parameters are updated by defining the advantage of the state-action pair , which is calculated based on the action value of the agent at each step:
[0090] ;
[0091] To avoid too large a difference in the policy distribution before and after parameter update, the KL divergence is used as the weighting factor ω for the advantage value. At the same time, ω is truncated to set an upper limit for the advantage value:
[0092] ;
[0093] wherein, represents the KL divergence between the old and new policy distributions, and respectively represent the policy distributions of the agent before and after the update . We calculate the training gradient for maximizing the policy advantage value , which is used to update the Actor network parameters , and the update of the Actor network parameters is as shown in the above Actor network parameter update expression.
[0094] To achieve the simultaneous optimization of multiple objectives, it is necessary to increase the weight of the data transmission bandwidth penalty while maximizing . To find the optimal parameters, a reinforcement learning method combined with Lagrangian dual optimization is adopted. It is necessary to find the saddle point of in the dual function and the optimal Lagrange multiplier , which is equivalent to solving:
[0095] ;
[0096] By taking the derivative of the Lagrangian function, the Lagrange multiplier at the above t moment can be updated using the update expression of the Lagrange multiplier at t moment.
[0097] Introduce a penalty value network to assist in solving the optimal Lagrange multiplier . This penalty value network is used to estimate the penalty value of the agent under different observations, denoted as: . The parameters of its penalty value network are updated by defining the temporal difference, and the specific update method is as shown in the above penalty value network parameter update expression.
[0098] By calculating the action penalty value of each agent at different time steps, the above Lagrange multiplier update expression can be used for update.
[0099] Based on this method, in the process of solving the primal-dual problem, the optimal policy parameters and the Lagrange multiplier can be obtained under the bandwidth demand constraint.
[0100] In one embodiment, the expression for the required bandwidth constraint based on time-division multiple access is:
[0101] ;
[0102] ;
[0103] ;
[0104] wherein, represents the upper limit of the probability that the agent sends a message at each step, represents the entropy of the Gaussian distribution, represents the average number of bits sent by the agent, represents the maximum transmission rate of TDMA, n represents the average number of symbols sent per second, represents the number of symbols included in each message, F represents the sampling frequency of the system, N represents the number of agents, m represents a message; represents the bandwidth penalty threshold, represents the agent j at the t th time slot needs to send the total amount of data, represents the agent j action at the current time step.
[0105] Specifically, in the case of the required bandwidth constraint based on time-division multiple access, the specific implementation of MOC-PPO includes: First, based on the required bandwidth constraint of time-division multiple access (TDMA): To minimize the communication overhead, the threshold for estimating the transmission probability is based on the upper limit. Then, the Lagrange method is used to relax this constraint.
[0106] For ease of analysis, assume that the distribution of message m follows a Gaussian distribution, and its entropy is represented by H ( m ). Under the TDMA-based network protocol, the maximum number of time slots is T, the occupancy time of each time slot is d milliseconds, and the maximum amount of data in each time slot is Z bytes. Therefore, the maximum transmission rate of TDMA is bits per second. According to the Shannon theorem, to ensure that no information is lost during the message transmission, the average number of bits sent by the agent must satisfy: .
[0107] , and H ( m ) The relationship between them is: , where n is the average number of symbols sent per second.
[0108] In addition, the mean of the message distribution can be used μ and the variance σ to calculate the entropy H(m) of the Gaussian distribution:
[0109] ;
[0110] Using the inequality relationship, it can be deduced that:
[0111] ;
[0112] Assume that under the gating mechanism, the probability that each agent sends a message at each time step is p , and each message contains L symbols. The sampling frequency of the system is set to F . In a system composed of N agents connected by a graph, the average number of symbols n can be calculated as: , where it is assumed that in the connected graph, each agent has an adjacency relationship with the other half of the agents in the communication channel. Next, the upper bound of the probability that each agent sends a message at each step can be deduced:
[0113] ;
[0114] In the TDMA communication protocol, only one agent can send a message within each time slot, and the sampling frequency can be calculated as: . According to the previous calculation, the maximum transmission rate is given by the following formula: .
[0115] Combining these, the upper bound of the probability that an agent sends data in the TDMA protocol can be deduced as:
[0116] ;
[0117] where represents the upper bound of the probability that an agent sends data in the TDMA protocol.
[0118] Based on the deduced upper bound of the data sending probability, the penalty for the bandwidth occupancy of the gating mechanism for an agent during a specific period can be described as:
[0119] ;
[0120] where represents the bandwidth penalty threshold, and the term represents that the agent j will be penalized when occupying the bandwidth in the t th time slot. represents the agentj In the t total amount of data to be sent in the
[0121] It should be understood that although Figure 1 each step in the flowchart of Figure 1 is shown sequentially according to the indication of the arrow, these steps are not necessarily executed sequentially in the order indicated by the arrow. Unless otherwise specified herein, there is no strict order restriction for the execution of these steps, and these steps may be executed in other orders. Moreover,
[0122] In a verification embodiment, to verify the performance of the multi-objective coupled PPO method (MOC-PPO), a hardware-in-the-loop (HIL) simulation platform was constructed to address the challenges of deploying the policy to the actual communication scenario. The hardware-in-the-loop simulation platform mainly consists of two parts: a virtual training platform and an actual verification platform. The virtual training platform consists of a self-organizing network simulation server platform (SONSP) and a training host equipped with a GPU. The training host will perform closed-loop asynchronous training using an asynchronous training method. The SONSP contains three communication nodes: a virtual node, a bridging node, and an actual node. The virtual node is a network device built into the SONSP and runs the same 5-layer communication protocol as the actual node. Each thread accesses the virtual node on the server through any available port on the training host. Inter-node communication on the server is used to simulate the self-organizing network communication between drones. The verification platform consists of a hybrid network communication terminal, a bridging communication terminal, and an airborne cooperative controller. The hybrid network communication terminal acts as the actual node in the SONSP. The bridging communication terminal acts as the bridging node in the SONSP and serves as a mediator between the actual node and the virtual node. The physical layer protocol runs on the actual terminal device, while the other four network protocol layers run on the SONSP. The cluster network of the system consists of N actual nodes and X virtual nodes, forming a hardware-in-the-loop (HIL) network simulation system to simulate the cluster network communication. This N+X method provides good scalability for communication nodes. The airborne cooperative controller is designed to execute a task scheduler when the drone is running. To verify the effectiveness of MOC-PPO, the policy was trained in task allocation scenarios with 30 communication nodes and 40 communication nodes. MACPO was used as a comparison method and tested in multiple scenarios. MACPO is another advanced constrained multi-agent reinforcement learning method, where the constraint is also designed as a bandwidth demand constraint. Under the same conditions, the ablation version of MOC-PPO was used to optimize the communication constraint and degenerated into the MAPPO algorithm under asynchronous training. The evaluation metrics include task conflict, required bandwidth, and average data transfer times. The experimental results are as Figure 3 shown and compared with other methods. MOC-PPO trained an effective policy that can balance communication overhead and task conflict rate. As Figure 3 shown in the communication scenarios of 30 nodes and 40 nodes, compared with the original ACBBA, the average task conflict rate of ACBBA optimized by MOC-PPO decreased by 1% and 62.5% respectively, while the communication requirements decreased by 84% and 68% respectively, and the average communication frequency decreased by 82% and 72% respectively. The results show that MOC-PPO significantly reduces the communication requirements and frequency, while reducing the task conflict rate. However, the performance of MACPO is not ideal. In the scenario of 30 drones, although the communication requirements decreased significantly, the task conflict rate increased to 15% ( Figure 3(a) to Figure 3 (c)). In the scenario of 40 UAVs, the task conflict rate decreased, but the communication demand increased by 8% ( Figure 3 (d) to Figure 3 (f)). MACPO lacks a penalty network to approximate the time-division multiple access bandwidth demand constraint, making it difficult to balance communication penalties and task conflicts during training. Similarly, ablation of the bandwidth constraint optimization in MOC-PPO also makes it difficult to train an effective communication strategy, and the average task conflict rates increase to 15% and 13% respectively. Such a high increase in the task conflict rate will hinder the multi-UAV system from completing the specified tasks, thus highlighting the importance of bandwidth constraint optimization. Overall, MOC-PPO successfully balances communication demand and task conflicts, achieving simultaneous optimization of multiple objectives under constraints.
[0123] As mentioned above, these observations have local characteristics and are not affected by the number of UAVs. By constructing a strategy that only depends on the local characteristics of adjacent UAVs and shares the same network parameters, a general communication strategy applicable to UAVs of different scales can be trained. The optimized communication strategy was tested using the N+X mode in a hardware-in-the-loop (HIL) experimental scenario, and the communication strategy trained in the 30-UAV scenario was applied to the 40-UAV and 50-UAV scenarios. The experimental results show that this communication strategy reduces the task conflict rate of ACBBA and significantly reduces the required communication bandwidth in all scenarios. In the scenarios of 30, 40, and 50 UAVs, compared with ACBBA, the bandwidth demands are reduced by 84%, 62.5%, and 44.4% respectively. These results indicate that the communication strategy trained in a small-scale environment can also improve the performance of task allocation algorithms in large-scale and even real environments.
[0124] A hardware-in-the-loop (HIL) platform was constructed based on the Self-Organizing Network Simulation Platform Server (SONSP) for training and verification to solve the problem of deploying the strategy trained in the simulation environment to the actual communication scenario. In addition, the performances of MOC-PPO and MACPO under different UAV scale configurations were carefully evaluated. The experimental results show that MOC-PPO can reduce the bandwidth demand by up to 84% and the task conflict rate by up to 62.5%. In addition, the communication strategy trained under small-scale conditions can be extended to a larger-scale environment. To address the gap between simulation and reality, a hardware-in-the-loop (HIL) platform was used to verify the trained strategy, demonstrating the scalability of the UAV swarm from 30 to 50. A large number of experimental comparisons with the state-of-the-art methods (such as MACPO) show that MOC-PPO achieves a better balance between communication efficiency and allocation reliability.
[0125] In one embodiment, a communication-aware task allocation device based on a multi-objective asynchronous strategy is provided. This device is applicable to the implementation of communication-aware task allocation in a UAV swarm. Each UAV in the UAV swarm deploys an agent, and the UAV swarm constitutes a multi-agent system. The device includes:
[0126] A problem modeling module, configured to transform the communication strategy optimization problem in the distributed task allocation algorithm of the multi-agent system into an asynchronous-constrained decentralized partially observable Markov decision process according to bandwidth constraints.
[0127] An asynchronous-constrained multi-agent reinforcement learning environment construction module, configured to construct an asynchronous-constrained multi-agent reinforcement learning environment considering the communication process according to the asynchronous-constrained decentralized partially observable Markov decision process.
[0128] A multi-objective asynchronous communication strategy optimization module, configured to optimize the multi-objective asynchronous communication strategy in the distributed task allocation method by using the multi-objective coupled PPO method according to the asynchronous-constrained multi-agent reinforcement learning environment, and obtain the final multi-objective asynchronous communication strategy. The multi-objective coupled PPO method refers to a method that extends the Lagrangian relaxation method to the multi-agent proximal policy optimization algorithm, asynchronously collects data in a distributed manner, and then uses the Lagrangian function to centrally optimize the communication strategy under the given bandwidth demand constraints to achieve the simultaneous optimization of multiple objectives and obtain the final multi-objective asynchronous communication strategy.
[0129] A communication-aware task allocation module based on the above, configured to enable each UAV to adopt the final multi-objective asynchronous communication strategy to complete the communication-aware task allocation of the UAV swarm.
[0130] In one of the embodiments, the tuple of the asynchronous-constrained decentralized partially observable Markov decision process in the problem modeling module is:
[0131] ;
[0132] Wherein, represents the tuple of the asynchronous-constrained decentralized partially observable Markov decision process, represents n the set of agents, S represents the global state, and respectively represent the observation value and action set of all agents at time step t , represents the state transition function, represents the new global state, R represents the reward function, γ ∈[0,1] represents the discount factor, represents multiple auxiliary constraints for restricting the update of the agent policy, Indicates the th auxiliary constraint, .
[0133] The asynchronous constraint decentralized partially observable Markov decision process allows agents to make decisions asynchronously. The goal of each agent is to learn a measurement while maximizing the expected discounted return , and satisfying all auxiliary constraint conditions; the auxiliary constraint conditions are as shown in the above auxiliary constraint condition expressions.
[0134] In one embodiment, in the asynchronous constraint decentralized partially observable Markov decision process in the problem modeling module: the action of the agent j at its local time step is defined as a communication strategy based on a gating mechanism, enabling the agent to adaptively select the timing of data transmission, thereby coordinating communication to improve algorithm performance; at the local time step j of the agent , the global shared reward is defined as the sum of the changes in the average task conflict numbers of all agents; the observations include: Bron-Kerbosch features, message value features, and channel access features.
[0135] In one embodiment, the multi-objective asynchronous communication strategy optimization module is further configured to independently and asynchronously collect its interaction data for each agent in the asynchronous constraint multi-agent reinforcement learning environment. During training, the interaction data from different agents is concatenated into a trajectory according to the time relationship; the interaction data includes the global state, local observations, actions, and rewards; among them, the action is a communication strategy based on a gating mechanism; on the basis of the original MAPPO algorithm, in order to ensure that the gating mechanism strategy satisfies the constraint conditions, Lagrange multipliers are introduced, and the Lagrangian function is defined as shown in the above Lagrangian function expression; the Lagrangian function is used to centrally optimize the communication strategy under the given bandwidth demand constraint, achieving the simultaneous optimization of multiple objectives to obtain the final multi-objective asynchronous communication strategy.
[0136] In one embodiment, the Lagrangian function is used to centrally optimize the communication strategy under the given bandwidth demand constraint, achieving the simultaneous optimization of multiple objectives to obtain the final multi-objective asynchronous communication strategy, including: defining the dual function of the original problem as:
[0137] ;
[0138] For the purpose of finding the optimal policy parameters that satisfy , with the Lagrange multipliers at time step tWhen it is a fixed value, each agent constructs a new reward to train its policy; the new reward is as shown in the above new reward expression; the training gradient is calculated by minimizing the mean square error between the state value function and the new reward, and the Critic network parameters are updated according to this gradient; the update method of the Critic network parameters is as shown in the above update expression of the Critic network parameters; for the Actor network, the policy advantage value is defined, and the training gradient for maximizing the policy advantage value is calculated, and the Actor network parameters are updated according to this gradient; the update method of the Actor network parameters is as shown in the above update expression of the Actor network parameters.
[0139] In one embodiment, the weighting factor in the multi-objective asynchronous communication policy optimization module is as shown in the above weighting factor expression.
[0140] In one embodiment, the multi-objective asynchronous communication policy optimization module obtains the optimal policy parameters under the bandwidth demand constraint and the Lagrange multiplier The process of finding the saddle point in the dual function and the optimal Lagrange multiplier is transformed into:
[0141] ;
[0142] where represents the optimal Lagrange multiplier.
[0143] By taking the derivative of the Lagrangian function, the Lagrange multiplier at t time is updated, and the update method of the Lagrange multiplier at t time is as shown in the above value update expression. The penalty value network is used to estimate the penalty value of the agent under different observations, and the penalty value network parameters are updated by temporal difference. The penalty value network parameter update method is as shown in the above penalty value network parameter update expression; the action penalty value of each agent at different time steps is calculated according to the penalty value network, and the Lagrange multiplier is updated according to the action penalty value; the Lagrange multiplier update method is as described in the above Lagrange multiplier update expression.
[0144] In one embodiment, the required bandwidth constraint based on time division multiple access in the multi-objective asynchronous communication policy optimization module is as shown in the above required bandwidth constraint expression based on time division multiple access.
[0145] For the specific limitations of the communication awareness task allocation device based on the multi-objective asynchronous strategy, reference may be made to the limitations of the communication awareness task allocation method based on the multi-objective asynchronous strategy in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned communication awareness task allocation device based on the multi-objective asynchronous strategy can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0146] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0147] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A communication-aware task allocation method based on a multi-objective asynchronous strategy, characterized in that The method is applicable to the communication perception task allocation of an unmanned aerial vehicle (UAV) cluster. Each UAV in the UAV cluster deploys an agent, and the UAV cluster constitutes a multi-agent system. The method includes: According to the bandwidth constraint, transform the communication strategy optimization problem in the distributed task allocation algorithm of the multi-agent system into an asynchronous constraint decentralized partially observable Markov decision process; Construct an asynchronous constraint multi-agent reinforcement learning environment considering the communication process according to the asynchronous constraint decentralized partially observable Markov decision process; According to the asynchronous constraint multi-agent reinforcement learning environment, use the multi-objective coupled proximal policy optimization (PPO) method to optimize the multi-objective asynchronous communication strategy in the distributed task allocation method, and obtain the final multi-objective asynchronous communication strategy. The multi-objective coupled PPO method refers to the method of extending the Lagrangian relaxation method to the multi-agent proximal policy optimization algorithm, asynchronously collecting data in a distributed manner, and then using the Lagrangian function to centrally optimize the communication strategy under the given bandwidth demand constraint to achieve the simultaneous optimization of multiple objectives and obtain the final multi-objective asynchronous communication strategy; Each UAV adopts the final multi-objective asynchronous communication strategy to complete the communication perception task allocation of the UAV cluster; Among them, the tuple of the asynchronous constraint decentralized partially observable Markov decision process is: ; Among them, represents a tuple of an asynchronous constrained partially observable Markov decision process, represents n a set of S agents, and respectively represent the observations and action sets of all agents at time step t , represents the state transition function, represents the new global state, R represents the reward function, γ ∈[0,1] represents the discount factor, represents multiple auxiliary constraints for restricting the update of the agent policy, represents the i th auxiliary constraint, i = 1, 2, ……, n ; The asynchronous constrained decentralized partially observable Markov decision process allows agents to make decisions asynchronously, and the goal of each agent is to learn a measurement while maximizing the expected discounted return , and to satisfy all auxiliary constraints; the auxiliary constraints are: ; Among them, represents the auxiliary constraint function, represents the agent j at time step t the reward obtained, represents the time step t the discount factor at that time, represents the agent j the action at the current time step, is the state at time step t at that time.
2. The communication awareness task allocation method based on a multi-objective asynchronous strategy according to claim 1, characterized in that, In the asynchronous constraint decentralized partially observable Markov decision process: Define the action of the agent j at its local time step as a communication strategy based on a gating mechanism, enabling the agent to adaptively select the timing of data transmission, thereby coordinating communication to improve algorithm performance; At the local time step of the agent j the global shared reward is defined as the sum of the changes in the average number of task conflicts among all agents; at The observations include: Bron-Kerbosch features, message value features, and channel access features.
3. The communication-aware task allocation method based on the multi-objective asynchronous strategy according to claim 1, wherein According to the asynchronous constraint multi-agent reinforcement learning environment, use the multi-objective coupled PPO method to optimize the multi-objective asynchronous communication strategy in the distributed task allocation method, and obtain the final multi-objective asynchronous communication strategy. The multi-objective coupled PPO method refers to the method of extending the Lagrangian relaxation method to the multi-agent proximal policy optimization algorithm, asynchronously collecting data in a distributed manner, and then using the Lagrangian function to centrally optimize the communication strategy under the given bandwidth demand constraint to achieve the simultaneous optimization of multiple objectives and obtain the final multi-objective asynchronous communication strategy, including: In the asynchronous constraint multi-agent reinforcement learning environment, each agent independently and asynchronously collects its interaction data. During training, the interaction data from different agents is concatenated into a trajectory according to the time relationship. The interaction data includes the global state, local observation, action, and reward. Among them, the action is a communication strategy based on a gating mechanism; On the basis of the original MAPPO algorithm, in order to ensure that the gating mechanism strategy meets the constraint conditions, introduce Lagrange multipliers and define the Lagrangian function as: ; Among them, represents the Lagrangian function, represents the Lagrange multiplier, represents the agent j 's Actor network parameters, represents the agent j at time step t obtained reward, represents the bandwidth penalty threshold, represents the agent j at the t th time slot, the total amount of data to be sent, represents the agent j 's action at the current time step, represents the time step t when the discount factor; Use the Lagrangian function to centrally optimize the communication strategy under the given bandwidth demand constraint to achieve the simultaneous optimization of multiple objectives and obtain the final multi-objective asynchronous communication strategy.
4. The communication-aware task allocation method based on the multi-objective asynchronous strategy according to claim 3, characterized in that, Use the Lagrangian function to centrally optimize the communication strategy under the given bandwidth demand constraint to achieve the simultaneous optimization of multiple objectives and obtain the final multi-objective asynchronous communication strategy, including: Define the dual function of the original problem as: ; For the purpose of finding the optimal policy parameters that satisfy When the Lagrange multiplier is a fixed value at time step t each agent constructs a new reward to train its policy; the new reward is: ; Among them, represents the new reward of the agent j at time step t, represents the reward of the agent j at time step t ; represents the Lagrange multiplier at time step t ; represents the total amount of data that the agent j needs to send in the t th time slot; represents the action of the agent j at the current time step; The training gradient is calculated by minimizing the mean square error between the state value function and the new reward, and the Critic network parameters are updated according to this gradient; the update method of the Critic network parameters is as follows: ; Among them, represents the Critic network parameters, represents the state value function; For the Actor network, the policy advantage value is defined, and the training gradient for maximizing the policy advantage value is calculated, and the Actor network parameters are updated according to this gradient; the update method of the Actor network parameters is as follows: ; ; ; Among them, represents the Actor network parameters, represents the advantage of the state-action pair, represents the advantage of the state-action pair after weighted processing, represents the weighting factor, represents the hyperparameter, is the discount factor.
5. The communication-aware task allocation method based on a multi-objective asynchronous strategy according to claim 4, wherein The weighting factor is: ; Among them, represents the KL divergence, and respectively represent the policy distributions of the agent before and after updating the policy.
6. The communication-aware task allocation method based on a multi-objective asynchronous strategy according to claim 4, characterized in that Obtaining the optimal policy parameters and the Lagrange multipliers in the process of including: The problem of finding the saddle point and the optimal Lagrange multiplier in the dual function is transformed into: ; Among them, represents the optimal Lagrange multiplier; By taking the derivative of the Lagrangian function, the Lagrange multiplier at t the moment is updated; t The update method of the Lagrange multiplier at the moment is as follows: ; ; Among them, represents t+ the Lagrange multiplier at time 1, represents t the Lagrange multiplier at time, represents the hyperparameter; Adopt a penalty value network Estimate the penalty value of the agent under different observations, and update the parameters of the penalty value network through temporal difference. The method for updating the parameters of the penalty value network is as follows: ; Among them, represents the penalty value network parameters, represents the penalty value network; The action penalty value of each agent at different time steps is calculated according to the penalty value network, and the Lagrange multiplier is updated according to the action penalty value; the update method of the Lagrange multiplier is: ; Among them, represents the action penalty value of the agent j at time step t.
7. The communication-aware task allocation method based on the multi-objective asynchronous strategy according to claim 1, wherein The required bandwidth constraint based on time division multiple access includes: ; ; ; Among them, represents the upper limit of the probability that the agent sends a message at each step, represents the entropy of the Gaussian distribution, represents the average number of bits sent by the agent, represents the maximum transmission rate of TDMA, represents the number of symbols included in each message, F represents the sampling frequency of the system, N represents the number of agents, and m represents the message; represents the bandwidth penalty threshold, represents the agent j at the t th time slot needs to send the total amount of data, represents the agent j action at the current time step, represents the time step t when the discount factor, n is the number of symbols sent per second.
8. A communication-aware task allocation device based on a multi-objective asynchronous strategy, characterized in that The device is applicable to the communication-aware task allocation of an unmanned aerial vehicle (UAV) cluster. Each UAV in the UAV cluster deploys an agent, and a multi-agent system is formed by the UAV cluster. The device includes: A problem modeling module, configured to convert the communication policy optimization problem in the distributed task allocation algorithm of the multi-agent system into an asynchronous constraint decentralized partially observable Markov decision process according to the bandwidth constraint; the tuple of the asynchronous constraint decentralized partially observable Markov decision process is: ; Among them, represents a tuple of an asynchronous constrained partially observable Markov decision process, represents n a set of S intelligent agents, and respectively represent the observations and action sets of all intelligent agents at time step t , represents the state transition function, represents the new global state, R represents the reward function, γ ∈[0,1] represents the discount factor, represents multiple auxiliary constraints for restricting the update of the agent policy, represents the i th auxiliary constraint, i = 1, 2, ……, n ; The asynchronous constrained decentralized partially observable Markov decision process allows agents to make decisions asynchronously. The goal of each agent is to learn a measurement while maximizing the expected discounted return and satisfying all the auxiliary constraints; the auxiliary constraints are as follows: ; Among them, represents the auxiliary constraint function, represents the agent j at the time step t the reward obtained, represents the time step t the discount factor at, represents the agent j the action at the current time step, is the state at the time step t ; An asynchronous constraint multi-agent reinforcement learning environment construction module, configured to construct an asynchronous constraint multi-agent reinforcement learning environment considering the communication process according to the asynchronous constraint decentralized partially observable Markov decision process; A multi-objective asynchronous communication policy optimization module, configured to optimize the multi-objective asynchronous communication policy in the distributed task allocation method by using the multi-objective coupled proximal policy optimization (PPO) method according to the asynchronous constraint multi-agent reinforcement learning environment, and obtain the final multi-objective asynchronous communication policy; the multi-objective coupled PPO method refers to a method of extending the Lagrangian relaxation method to the multi-agent proximal policy optimization algorithm, collecting data asynchronously in a distributed manner, and then using the Lagrangian function to centrally optimize the communication policy under the given bandwidth requirement constraint to simultaneously optimize multiple objectives and obtain the final multi-objective asynchronous communication policy; A communication-aware task allocation module based on the final multi-objective asynchronous communication policy is used for each UAV to complete the communication-aware task allocation of the UAV cluster.
Citation Information
Patent Citations
Active power distribution network distributed optimization method, system and device and storage medium
CN116629461A
Virtual power plant distributed resource collaborative optimization scheduling method and system
CN117610813A
Cited By
A Bee Product Scheduling Control Method Based on Multi-Agent Policy Optimization
CN122569234A