A cooperative charging planning method based on multi-agent deep reinforcement learning

CN115907377BActive Publication Date: 2026-09-22KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211462417.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2026-09-22
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

但导致如下问题:决策智能体必须花费更多的时间才能从训练样本中找到有用的信息

Benefits of technology

[0072]本发明使用多个MC可以尽可能避免在大规模或密集型无线可充电传感器网络中由于对传感器节点充电不及时造成的传感器节点能量耗尽。本发明基于多智能体深度强化学习优化无线可充电传感器网络中多MC协作调度问题,即提出一种多MC异步执行充电任务的框架,并使用协作通信单元,其中使用注意力机制为每个决策智能体提取其他智能体信息,使多MC可以更好的协作,避免相互抢占工作区域,从而在保证最小死亡节点数的前提下,使各个MC的移动路径长度最短,最大化多MC的充电效用。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115907377B_ABST
    Figure CN115907377B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-agent deep reinforcement learning's cooperation charging planning method, belong to wireless rechargeable sensor network (WRSN) field.In large-scale or intensive WRSN using multiple mobile chargers (MC) can avoid the node energy depletion caused by charging sensor node not in time.Due to unevenly distributed distribution and charging time of node, multiple MC needs asynchronous charging, so that the multi-agent deep reinforcement learning algorithm of the past is difficult to be used in the scene of multiple MC collaborative charging.Based on multi-agent deep reinforcement learning, the application optimizes the problem of multiple MC cooperative scheduling in WRSN, that is, a framework for multiple MC asynchronous charging is proposed, and a cooperative communication unit is used to dynamically extract information of other agents for each decision-making agent.The application aims to make multiple MCs better collaborate, so as to minimize the length of the movement path of each MC under the premise of ensuring the minimum number of dead nodes, and maximize the charging utility of multiple MCs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless rechargeable sensor networks, and more specifically to a cooperative charging planning method (MACC) based on multi-agent deep reinforcement learning. Background Technology

[0002] Energy constraints have always been a significant factor limiting the development of wireless sensor networks. Wireless Rechargeable Sensor Networks (WRSNs) are wireless sensor networks that deploy mobile chargers (MCs) to charge sensors within an energy-constrained network. WRSNs are now widely used in various fields such as military, agricultural production, forest fire prevention, and ecological monitoring. Effectively planning the charging path of the MCs to extend the WRSN's lifetime has become a key research problem.

[0003] Using multiple MCs (Multiple Agents) in large-scale or dense wireless rechargeable sensor networks can minimize the energy depletion of sensor nodes due to untimely charging. Single-agent deep reinforcement learning methods alone are insufficient to meet the requirements of multi-MC collaboration. Furthermore, because each sensor node is unevenly distributed and requires different charging times, multiple MCs need to perform asynchronous charging, making the structure of previous multi-agent deep reinforcement learning algorithms unsuitable for scenarios involving collaborative charging of multiple MCs. To address these issues, this paper studies a multi-agent deep reinforcement learning collaborative charging planning method that can satisfy asynchronous collaboration among multiple MCs, which can significantly improve the charging efficiency of multiple MCs.

[0004] In 2020, Meiyi Yang et al. published "Dynamic Charging Scheme Problem With Actor-Critic Reinforcement Learning" in the IEEE Internet of Things Journal, proposing a novel dynamic charging scheme (DCS) for WRSN based on the actor-critic reinforcement learning (ACRL) algorithm. The method utilizes single-agent deep reinforcement learning to dynamically select charging nodes for a single MC (Mechanical Controller), outperforming traditional heuristic methods in terms of node average lifetime and MC movement path length. However, a single MC cannot meet the charging needs of large-scale or dense wireless rechargeable sensor networks.

[0005] In 2021, Yuxin Chen and other scholars presented "VarLenMARL: A Framework of Variable-Length Time-Step Multi-Agent Reinforcement Learning for Cooperative Charging in Sensor Networks" at the IEEE International Conference on Sensing, Communication, and Networking. This paper introduced a novel multi-agent deep reinforcement learning framework called VarLenMARL. The training sample collection mechanism in VarLenMARL uses old information from other agent cores (MCs) to make action decisions for the decision MC, allowing each MC to complete an action in a variable-length time step before estimating the reward. This solves the problem of asynchronous charging for multiple MCs. However, this leads to the following issues: the decision agent must spend more time finding useful information from the training samples. The normal training process of the agent becomes unstable due to interference from this old data.

[0006] In 2022, Yongheng Liang et al. presented an algorithm called Asynchronous and Scalable Multi-Agent PPO for Cooperative Charging at the International Conference on Autonomous Agents and Multiagent Systems. This algorithm allows agents to perform asynchronous learning and decision-making within an asynchronous Proximal Policy Optimization (POMDP). This addresses the problem of asynchronous charging required for multiple agent-matrix (MC) systems. The method restricts the observation region of an MC to the few sensor nodes closest to the decision MC and uses a distributed execution-centralized training multi-agent deep reinforcement learning framework. The distributed execution-centralized training framework aims to address the problem of unstable environments caused by the influence of other agents' decisions during simultaneous multi-agent decision-making. However, in asynchronous scenarios, other agents are often in the process of executing actions while the decision-making agent is making its decision. Therefore, the distributed execution-centralized training framework often fails to learn useful cooperative policies and is not suitable for asynchronous multi-agent problems, leading to a lack of cooperation among agents.

[0007] The ATOC algorithm proposed by Jiechuan Jiang et al. in their 2018 paper "Learning Attentional Communication for Multi-Agent Cooperation" at the 32nd Conference on Neural Information Processing Systems aims to solve the problem of multi-agent communication and cooperation. However, the ATOC algorithm is designed for the problem of synchronous decision-making among multiple agents and has been tested in a "multi-agent particle environment" scenario.

[0008] Existing literature only considers the problem of asynchronous multi-agent communication, without considering the problem of asynchronous multi-agent communication assistance. Summary of the Invention

[0009] This invention proposes a collaborative charging planning method based on multi-agent deep reinforcement learning to solve the above problems.

[0010] The technical solution adopted in this invention is: a collaborative charging planning method based on multi-agent deep reinforcement learning, comprising the following steps:

[0011] 1. Scenario for building a multi-mobile charger (MC) wireless rechargeable sensor network (WRSN):

[0012] The WRSN is deployed within a defined two-dimensional monitoring area, comprising: n identical sensor nodes; 1 base station (BS); 1 service station (SS); and m identical monitoring stations (MCs). The time from the start of WRSN deployment until all nodes die is called its lifecycle. The WRSN lifecycle is divided into several identical time slots t, where time slot t is short and indivisible.

[0013] Sensor nodes are randomly distributed within the monitoring area, but their locations are fixed. i} represents a set of nodes, where i represents the node index, 1≤i≤n; s i The two-dimensional coordinates of the node; Es represents the total energy of the node; Represents node s i The remaining energy in time slot t; p i (t) represents the instantaneous energy consumption rate of the node in time slot t. This represents the average energy consumption rate of a node from the start of its network lifetime to the current time slot t. From the initial moment, all nodes collect data and transmit it to the base station via multi-hop forwarding. Due to the unpredictability of events occurring around sensor nodes and the bursty data streams from sensor nodes, the energy consumption rate of a sensor node changes dynamically. The remaining energy of the sensor is updated once every time slot t, as shown in the following formula:

[0014]

[0015] The average energy consumption rate of the sensor is updated once per time slot t, as shown in the following formula:

[0016]

[0017] h s This represents the threshold for a node. All nodes in each time slot t... The node will send its status information to the base station via multi-hop forwarding. As shown in the following formula:

[0018]

[0019] When a node runs out of energy, it will go into hibernation and will be unable to provide any service to the network.

[0020] MCs are autonomous mobile devices that can move freely within the WRSN deployment area. MCs can obtain their own real-time location. j Let} denote the set of MCs, where j represents the index of the MC, and 1≤j≤m; MC j The two-dimensional coordinates of MC. The total energy capacity of MC is E. m MC's movement speed is v, and its movement energy consumption is q. m Charging power is expressed as q. c Charging efficiency is denoted as η. MCs are divided into idle MCs and occupied MCs. In each time slot t, an idle MC receives a charging target node from the base station via long-distance real-time communication and proceeds to perform one-to-one charging; while an occupied MC continues to complete its charging task. Each charging task of an MC takes several time slots, which are defined as MCs. j A time step, denoted as Where t represents the sequence number of the time step, i.e. The start time slot t; j represents the id of the MC; the MC at time slot t jThe charging task is received and execution begins. Because the sensor nodes are unevenly distributed and the charging time varies, multiple MCs need to perform asynchronous charging. This means that in most cases, the time steps of different MCs contain a different number of time slots, and the time steps begin and end in different time slots. The MC can obtain the most recently updated state information of the node that sent the charging request in time slot t from the base station. In each time slot t, all MCs send their own status information to the base station, represented as:

[0021]

[0022] in Indicates time slot tMC j Location, MC j The location of the node being traveled to or charging, Δt represents the position of the node in MC. j The estimated remaining time to complete the current charging task. If MC is idle, then... And Δt = 0.

[0023] The service station has sufficient power to wirelessly charge the MC. m This represents the threshold of MC. If MC is not set after each charging task is completed... i The energy is less than h m If the MC needs to return to the service station to replenish its energy, it cannot perform charging tasks during this period.

[0024] The base station maintains the state information of low-energy nodes and all MCs. (1≤i≤n, 1≤j≤m). In each time slot t, the energy is lower than h. s Nodes are inserted into a request queue of length |A| according to a first-come, first-served principle. If the number of requests in time slot t is greater than |A|, requests exceeding the queue length will be discarded. Nodes in the request queue cannot be duplicated, and dead nodes will be removed. If the number of nodes in the request queue is less than |A|, empty slots in the queue are filled with zeros, and empty values ​​in the queue are considered invalid actions. If the request queue is not empty and there are idle MCs, the base station uses the MACC algorithm to select target nodes for charging for the idle MCs sequentially based on the valid actions in the request queue, and sends them to the corresponding MCs. Nodes selected as having action values ​​will be removed from the request queue.

[0025] 2. To maximize energy utilization and minimize the number of dead nodes, the following optimization problem is established:

[0026] The charging solution problem is abstracted into a multi-objective optimization problem, where the primary objective is to maximize the energy utilization of all MCs, and the secondary objective is to minimize the number of dead nodes. Maximizing the energy utilization of all MCs is equivalent to minimizing the movement distance of MCs. The objective of the optimization problem is defined as:

[0027]

[0028] d j MC j The distance traveled. i Indicates the liveness status of a node:

[0029]

[0030] 3. Construct a collaborative charging planning algorithm based on multi-agent deep reinforcement learning. Step three includes the following steps:

[0031] 3.1 The asynchronous partially observable Markov process (POMDP) ​​is defined as follows:

[0032] Asynchronous multi-agent POMDP is defined as a tuple (N, S, O, A, R, P, γ). Here, N = {1, 2, ..., M} is the set of M agents. S is the set of states describing the environment, O is the set of observations of the agents, A is the set of actions of the agents, R is the reward function, P is the state transition function, and γ is the reward discount factor. In time slot t, each agent j in N receives local observations related to state S. In time slot t, only one set of agents is available. A decision needs to be made. Let represent the random policy of agent j. Agent j in N′ is based on... Obtain the action a of agent j j The reward value r that agent j receives after completing an action. j R is a function of state and action, i.e. In the next time slot t+1, the next state is generated according to the state transition function P, i.e., P:

[0033] 3.2 Define the observation space, action space, and reward function for all agents:

[0034] Indicates MC at time slot t j For node s i The observation information tuple is defined as follows:

[0035]

[0036] in MC j To node si The relative position vector. That is, for an idle MC, MC j To node s i The relative position vector for occupying MC, MC j At the current time step The target node to node s i The relative position vector.

[0037] The observation space of agent j in time slot t j (t) is defined as follows:

[0038]

[0039] in This represents a tuple of observation information for all nodes in the request pool at time slot t.

[0040] In the action space a of agent j in time slot t j (t) is defined as follows:

[0041] a j (t)={s i (t)},(1≤i≤|A|) (9)

[0042] Where {s i The table (t) represents the set of nodes in the request queue at decision time slot t.

[0043] Agent j in Reward function at the end of time slot t The definition is as follows:

[0044]

[0045] in express MC j Energy used to charge nodes express MC j Energy used for movement In order to be in The node i that dies in the middle, Let α represent the two-dimensional coordinates of node i, and α be the penalty coefficient.

[0046] 3.3 The framework definition of the collaborative charging planning algorithm based on multi-agent deep reinforcement learning is as follows:

[0047] In this invention, each MC corresponds to an agent, and idle MCs make action decisions through their corresponding agents. Agents in the action decision-making state are called decision agents. General multi-agent frameworks aim to solve the problem of environmental instability caused by all agents making decisions simultaneously in the same time slot. However, in the scenario of multiple MCs asynchronously executing charging tasks, other agents often have already made decisions and are in the process of executing actions when the decision agent makes its decision. Therefore, compared to general multi-agent frameworks, the problem of multiple MCs asynchronously executing charging tasks is more suitable to be decomposed into an independent learning problem where a series of agents make decisions sequentially. Simultaneously, to avoid independent learning by different agents, resulting in a lack of collaboration and causing a large number of inefficient movements by the MCs, this invention provides the decision agent with the state information of other agents through a cooperative communication unit, forming a communication group. By sharing the encoding of local observations and action intentions within the dynamically formed communication group, the decision agent can establish a relatively more global environmental perception, infer the intentions of other agents, and cooperate in decision-making. The cooperative communication unit also plays a role in reducing environmental instability.

[0048] The independent learning agents of this invention use an actor-critic structure. Multiple agents share a single actor network and a single critic network.

[0049] The ActorNet takes as input the observations of all agents and outputs an action value. The ActorNet consists of a policy network, an attention unit, and a communication channel. The policy network is divided into two parts: ActorNet(I) and ActorNet(II). The policy network and attention unit are composed of an MLP. The communication channel is composed of a GRU. The parameters of the CriticNet, ActorNet, Communication Channel, and Attention Unit are θ. Q θ μ θ g and θ p The decision-making process of the actor network is as follows: At decision time slot t, ActorNet(I) uses Observation, that is, the local observations of each agent. j (t) is used as input to encode the local observations of agent j as the action intention of agent j, represented as: and The attention unit acts as a binary classifier, influencing the decisions of other agents outside the main agent. As input, it determines whether these agents j participate in the communication. If they participate, the decision-making agent will form a communication group with the other participating agents. The communication channel then connects the individual agents in the communication group. Merging as input and outputting the integrated action intent is represented as After being fed into the rest of the policy network, ActorNet(II), the output action is... Then, illegal actions are blocked by using a mask.

[0050] The critic network is parameterized by fully connected layers, taking the observations of all agents as input and outputting the action state Q-value.

[0051] 3.4 During the training process of the collaborative charging planning algorithm based on multi-agent deep reinforcement learning, the neural network update process is as follows:

[0052] This algorithm is an extension of the DDPG algorithm. The experience replay pool is a structure used to store training data. Training data (o, a, r, o′) is sampled from the experience replay pool to update the commentator network and the actor network.

[0053] Update the critic network according to the following formula:

[0054]

[0055] The policy gradient is represented as:

[0056]

[0057] Using the chain rule, the gradient of the action intent integrated in the communication channel is represented as:

[0058]

[0059] The gradient is backpropagated to the policy network and communication channels to update the parameters.

[0060] Then the target Q network soft update is:

[0061] θ′=τθ+(1-τ)θ′ (14)

[0062] The attention unit is trained as a binary classifier for communication. For each decision agent j and its communication group G... j Calculate the difference in average Q-values ​​between communication-based actions and non-communication-based actions (expressed as...). ):

[0063]

[0064] (ΔQ) j ,h j ) is stored in queue D, where ΔQ j The weights represent the performance enhancements resulting from communication. At the end of a training cycle, ΔQ in D is... jPerforming min-max normalization yields... The labels are used as the binary classifier, and θ is updated using log loss. p .

[0065]

[0066] 4. During the training process of the cooperative charging planning algorithm based on multi-agent deep reinforcement learning, historical data from WRSN is used to train the algorithm offline, resulting in a well-trained deep reinforcement learning model for solving the multi-MC cooperative scheduling problem.

[0067] First, initialize the parameters of the actor network and the critic network. Then, initialize the experience replay pool.

[0068] During the execution phase, in each time slot t, the base station acquires If the request queue is not empty and there is an idle MC j Then MC j The corresponding agent j becomes the decision agent in turn to make action decisions, and is MC. j Establish time step Next training data During decision-making, the actor network inputs the observed values ​​o(t) of all agents = (o1(t),…,o…). j (t)), output action a j (t), to obtain training data =(o(t),a j (t)). Then MC j The charging process will take several time slots to complete. At time slot t', MC... j Earn reward value r j (t′) and the observed values ​​of all agents o(t′)=(o1(t′),…,o j (t′)), supplementing training data Each agent has its own experience replay pool, The data is stored in the corresponding agent's experience replay pool. This phase is repeated until the number of samples in the experience pool reaches a set threshold, after which the training phase begins. During training, the actor network and critic network are trained using the experience replay pools of all agents. A certain amount of training data is extracted from the experience pool. The parameters of the actor network and critic network are updated using the training data until the neural network parameters converge.

[0069] 5. During the execution of the collaborative charging planning algorithm based on multi-agent deep reinforcement learning, the WRSN state information is input into the trained deep reinforcement learning model, and the charging action value of MC is calculated by the model:

[0070] In each time slot t, the base station acquires If the request queue is not empty and there is an idle MC j Then MC j The corresponding agents j become decision-making agents in sequence to make action decisions. During decision-making, the trained agent network is input with the observed values ​​o(t) = (o1(t), ..., o2(t)). j (t)), output action a j (t). Then MC j The charging process will take several time slots to complete.

[0071] The beneficial effects of this invention are:

[0072] This invention utilizes multiple MCs to minimize energy depletion of sensor nodes in large-scale or dense wireless rechargeable sensor networks due to untimely charging. Based on multi-agent deep reinforcement learning, this invention optimizes the multi-MC cooperative scheduling problem in wireless rechargeable sensor networks. Specifically, it proposes a framework for asynchronous execution of charging tasks by multiple MCs and employs a cooperative communication unit. An attention mechanism is used to extract information from other agents for each decision-making agent, enabling better cooperation among the MCs and preventing them from competing for work areas. This minimizes the number of dead nodes while shortening the movement path length of each MC, maximizing the charging efficiency of the multiple MCs. Attached Figure Description

[0073] Figure 1 This is a flowchart of the collaborative charging planning algorithm based on multi-agent deep reinforcement learning of the present invention;

[0074] Figure 2 This is a model diagram of the wireless rechargeable sensor network of the present invention;

[0075] Figure 3 This is a schematic diagram illustrating the asynchronous execution of charging tasks by multiple MCs according to the present invention;

[0076] Figure 4 This is a schematic diagram of the actor network and commentator network of the present invention. Detailed Implementation

[0077] To provide a more detailed description of the present invention and to facilitate understanding by those skilled in the art, the present invention will be further described below in conjunction with the accompanying drawings and embodiments. The embodiments in this section are for illustrative purposes and are not intended to limit the present invention.

[0078] Example 1: As Figure 1-4 As shown, a collaborative charging planning method based on multi-agent deep reinforcement learning includes the following steps:

[0079] 1. Scenario for building a multi-mobile charger (MC) wireless rechargeable sensor network (WRSN):

[0080] The WRSN is deployed within a defined two-dimensional monitoring area, comprising: n identical sensor nodes; 1 base station (BS); 1 service station (SS); and m identical monitoring stations (MCs). The time from the start of WRSN deployment until all nodes die is called its lifecycle. The WRSN lifecycle is divided into several identical time slots t, where time slot t is short and indivisible.

[0081] In this embodiment, the parameters are set as follows: Figure 2 As shown, the sensor node is located within the unit square [0,1]. 2 The energy is randomly and uniformly generated within the sensor, with the initial residual energy of each node randomly generated between 10 and 20 J. The velocity of the MCs is 0.1 m / s. Additionally, the energy consumption rate of the sensor... (1≤i≤n) is a random number, ranging from 0.1 to 0.5 J / s.

[0082] Sensor nodes are randomly distributed within the monitoring area, but their locations are fixed. i} represents a set of nodes, where i represents the node index, 1≤i≤n; s i The two-dimensional coordinates of the node; Es represents the total energy of the node; Represents node s i The remaining energy in time slot t; p i (t) represents the instantaneous energy consumption rate of the node in time slot t. This represents the average energy consumption rate of a node from the start of its network lifetime to the current time slot t. From the initial moment, all nodes collect data and transmit it to the base station via multi-hop forwarding. Due to the unpredictability of events occurring around sensor nodes and the bursty data streams from sensor nodes, the energy consumption rate of a sensor node changes dynamically. The remaining energy of the sensor is updated once every time slot t, as shown in the following formula:

[0083]

[0084] The average energy consumption rate of the sensor is updated once per time slot t, as shown in the following formula:

[0085]

[0086] h s This represents the threshold for a node. All nodes in each time slot t... The node will send its status information to the base station via multi-hop forwarding. As shown in the following formula:

[0087]

[0088] When a node runs out of energy, it will go into hibernation and will be unable to provide any service to the network.

[0089] MCs are autonomous mobile devices that can move freely within the WRSN deployment area. MCs can obtain their own real-time location. j Let} denote the set of MCs, where j represents the index of the MC, and 1≤j≤m; MC j The two-dimensional coordinates of MC. The total energy capacity of MC is E. m MC's movement speed is v, and its movement energy consumption is q. m Charging power is expressed as q. c Charging efficiency is denoted as η. MCs are divided into idle MCs and occupied MCs. In each time slot t, an idle MC receives a charging target node from the base station via long-distance real-time communication and proceeds to perform one-to-one charging; while an occupied MC continues to complete its charging task. Each charging task of an MC takes several time slots, which are defined as MCs. j A time step, denoted as Where t represents the sequence number of the time step, i.e. The start time slot t; j represents the id of the MC; the MC at time slot t j The charging task is received and execution begins. Because the sensor nodes are unevenly distributed and the charging time varies, multiple MCs need to perform asynchronous charging. This means that in most cases, the time steps of different MCs contain a different number of time slots, and the time steps begin and end in different time slots. The MC can obtain the most recently updated state information of the node that sent the charging request in time slot t from the base station. In each time slot t, all MCs send their own status information to the base station, represented as:

[0090]

[0091] in Indicates time slot tMC j Location, MC j The location of the node being traveled to or charging, Δt represents the position of the node in MC. j The estimated remaining time to complete the current charging task. If MC is idle, then... And Δt = 0.

[0092] The service station has sufficient power to wirelessly charge the MC. m This represents the threshold of MC. If MC is not set after each charging task is completed... i The energy is less than h mIf the MC needs to return to the service station to replenish its energy, it cannot perform charging tasks during this period.

[0093] The base station maintains the state information of low-energy nodes and all MCs. (1≤i≤n, 1≤j≤m). In each time slot t, the energy is lower than h. s Nodes are inserted into a request queue of length |A| according to a first-come, first-served principle. If the number of requests in time slot t is greater than |A|, requests exceeding the queue length will be discarded. Nodes in the request queue cannot be duplicated, and dead nodes will be removed. If the number of nodes in the request queue is less than |A|, empty slots in the queue are filled with zeros, and empty values ​​in the queue are considered invalid actions. If the request queue is not empty and there are idle MCs, the base station uses the MACC algorithm to select target nodes for charging for the idle MCs sequentially based on the valid actions in the request queue, and sends them to the corresponding MCs. Nodes selected as having action values ​​will be removed from the request queue.

[0094] 2. To maximize energy utilization and minimize the number of dead nodes, the following optimization problem is established:

[0095] The charging solution problem is abstracted into a multi-objective optimization problem, where the primary objective is to maximize the energy utilization of all MCs, and the secondary objective is to minimize the number of dead nodes. Maximizing the energy utilization of all MCs is equivalent to minimizing the movement distance of MCs. The objective of the optimization problem is defined as:

[0096]

[0097] d j MC j The distance traveled. i Indicates the liveness status of a node:

[0098]

[0099] 3. Construct a collaborative charging planning algorithm based on multi-agent deep reinforcement learning. Step three includes the following steps:

[0100] 3.1 The asynchronous partially observable Markov process (POMDP) ​​is defined as follows:

[0101] Asynchronous multi-agent POMDP is defined as a tuple (N, S, O, A, R, P, γ). Here, N = {1, 2, ..., M} is the set of M agents. S is the set of states describing the environment, O is the set of observations of the agents, A is the set of actions of the agents, R is the reward function, P is the state transition function, and γ is the reward discount factor. In time slot t, each agent j in N receives local observations related to state S. In time slot t, only one set of agents is available. A decision needs to be made. Let represent the random policy of agent j. Agent j in N′ is based on... Obtain the action a of agent j j The reward value r that agent j receives after completing an action. j R is a function of state and action, i.e. In the next time slot t+1, the next state is generated according to the state transition function P, that is...

[0102] 3.2 Define the observation space, action space, and reward function for all agents:

[0103] Indicates MC at time slot t j For node s i The observation information tuple is defined as follows:

[0104]

[0105] in MC j To node s i The relative position vector. That is, for an idle MC, MC j To node s i The relative position vector for occupying MC, MC j At the current time step The target node to node s i The relative position vector.

[0106] The observation space of agent j in time slot t j (t) is defined as follows:

[0107]

[0108] in This represents a tuple of observation information for all nodes in the request pool at time slot t.

[0109] In the action space a of agent j in time slot t j (t) is defined as follows:

[0110] a j (t)={s i (t)},(1≤i≤|A|) (9)

[0111] Where {s i The table (t) represents the set of nodes in the request queue at decision time slot t.

[0112] Agent j in Reward function at the end of time slot t The definition is as follows:

[0113]

[0114] in express MC j Energy used to charge nodes express MC j Energy used for movement In order to be in The node i that dies in the middle, Let α represent the two-dimensional coordinates of node i, and α be the penalty coefficient.

[0115] 3.3 The framework definition of the collaborative charging planning algorithm based on multi-agent deep reinforcement learning is as follows:

[0116] In this invention, each MC corresponds to an agent, and idle MCs make action decisions through their corresponding agents. Agents in the action decision-making state are called decision agents. General multi-agent frameworks aim to solve the problem of environmental instability caused by all agents making decisions simultaneously in the same time slot. However, in the scenario of multiple MCs asynchronously executing charging tasks, other agents often have already made decisions and are in the process of executing actions when the decision agent makes its decision. Therefore, compared to general multi-agent frameworks, the problem of multiple MCs asynchronously executing charging tasks is more suitable to be decomposed into an independent learning problem where a series of agents make decisions sequentially. Simultaneously, to avoid independent learning by different agents, resulting in a lack of collaboration and causing a large number of inefficient movements by the MCs, this invention provides the decision agent with the state information of other agents through a cooperative communication unit, forming a communication group. By sharing the encoding of local observations and action intentions within the dynamically formed communication group, the decision agent can establish a relatively more global environmental perception, infer the intentions of other agents, and cooperate in decision-making. The cooperative communication unit also plays a role in reducing environmental instability.

[0117] The independent learning agents of this invention use an actor-critic structure. Multiple agents share a single actor network and a single critic network.

[0118] The ActorNet takes as input the observations of all agents and outputs an action value. The ActorNet consists of a policy network, an attention unit, and a communication channel. The policy network is divided into two parts: ActorNet(I) and ActorNet(II). The policy network and attention unit are composed of an MLP. The communication channel is composed of a GRU. The parameters of the CriticNet, ActorNet, Communication Channel, and Attention Unit are θ. Q θ μ θ g and θ p The decision-making process of the actor network is as follows: At decision time slot t, ActorNet(I) uses Observation, that is, the local observations of each agent. j (t) is used as input to encode the local observations of agent j as the action intention of agent j, represented as: and The attention unit acts as a binary classifier, influencing the decisions of other agents outside the main agent. As input, it determines whether these agents j participate in the communication. If they participate, the decision-making agent will form a communication group with the other participating agents. The communication channel then connects the individual agents in the communication group. Merging as input and outputting the integrated action intent is represented as After being fed into the rest of the policy network, ActorNet(II), the output action is... Then, illegal actions are blocked by using a mask.

[0119] The critic network is parameterized by fully connected layers, taking the observations of all agents as input and outputting the action state Q-value.

[0120] 3.4 During the training process of the collaborative charging planning algorithm based on multi-agent deep reinforcement learning, the neural network update process is as follows:

[0121] This algorithm is an extension of the DDPG algorithm. The experience replay pool is a structure used to store training data. Training data (o, a, r, o′) is sampled from the experience replay pool to update the commentator network and the actor network.

[0122] Update the critic network according to the following formula:

[0123]

[0124] The policy gradient is represented as:

[0125]

[0126] Using the chain rule, the gradient of the action intent integrated in the communication channel is represented as:

[0127]

[0128] The gradient is backpropagated to the policy network and communication channels to update the parameters.

[0129] Then the target Q network soft update is:

[0130] θ′=τθ+(1-τ)θ′ (14)

[0131] The attention unit is trained as a binary classifier for communication. For each decision agent j and its communication group G... j Calculate the difference in average Q-values ​​between communication-based actions and non-communication-based actions (expressed as...). ):

[0132]

[0133] (ΔQ) j ,h j ) is stored in queue D, where ΔQ j The weights represent the performance enhancements resulting from communication. At the end of a training cycle, ΔQ in D is... j Performing min-max normalization yields... The labels are used as the binary classifier, and θ is updated using log loss. p .

[0134]

[0135] 4. During the training process of the cooperative charging planning algorithm based on multi-agent deep reinforcement learning, historical data from WRSN is used to train the algorithm offline, resulting in a well-trained deep reinforcement learning model for solving the multi-MC cooperative scheduling problem.

[0136] First, initialize the parameters of the actor network and the critic network. Then, initialize the experience replay pool.

[0137] During the execution phase, in each time slot t, the base station acquires If the request queue is not empty and there is an idle MC j Then MC j The corresponding agent j becomes the decision agent in turn to make action decisions, and is MC. j Establish time step Next training data

[0138] like Figure 3 As shown, during decision-making, the actor network inputs the observed values ​​of all agents, o(t) = (o1(t),...,o...). j (t)), output action a j (t), to obtain training data a j (t)). Then MC j The charging process will take several time slots to complete. At time slot t', MC... j Earn reward value r j (t′) and the observed values ​​of all agents o(t′)=(o1(t′),…,o j (t′)), supplementing training data a j (t),r j (t′),o(t′)). Each agent has its own experience replay pool, which will The experience replay samples are stored in the corresponding agent's experience replay pool. This phase is repeated until the number of samples in the experience pool reaches a set threshold, after which the training phase begins. During the training phase, the actor network and critic network are trained using the experience replay pools of all agents. The structures of the actor network and critic network are as follows: Figure 4 As shown. A certain amount of training data is drawn from the experience pool. The parameters of the actor network and the critic network are updated using the training data until the neural network parameters converge.

[0139] 5. During the execution of the collaborative charging planning algorithm based on multi-agent deep reinforcement learning, the WRSN state information is input into the trained deep reinforcement learning model, and the charging action value of MC is calculated by the model:

[0140] In each time slot t, the base station acquires If the request queue is not empty and there is an idle MC j Then MC j The corresponding agents j become decision-making agents in sequence to make action decisions. During decision-making, the trained agent network is input with the observed values ​​o(t) = (o1(t), ..., o2(t)). j (t)), output action a j (t). Then MC j The charging process will take several time slots to complete.

[0141] The above description is merely a specific idea of ​​the present invention to facilitate understanding by researchers in the field. However, the implementation of the present invention is not limited to the above description. Those skilled in the art can make improvements or modifications based on the present invention, and all improvements or modifications utilizing the concept of the present invention are considered to be within the scope of protection of the present invention.

Claims

1. A collaborative charging planning method based on multi-agent deep reinforcement learning, characterized in that, Includes the following steps: Step 1: Construct a scenario for a multi-mobile charger MC wireless rechargeable sensor network (WRSN); Step 2: Establish an optimization problem with the goal of maximizing energy utilization and minimizing the number of dead nodes; Step 3: Construct a collaborative charging planning algorithm based on multi-agent deep reinforcement learning; In step 3, constructing a collaborative charging planning algorithm based on multi-agent deep reinforcement learning specifically includes: Step 3.1: The asynchronous part of the observable Markov process POMDP is defined as follows: Asynchronous POMDP for multiple agents is defined as a tuple. ,in yes A collection of intelligent agents It is a set of states describing the environment. It is the set of observations of the intelligent agent. It is the set of actions of the intelligent agent. It is a reward function. It is a state transition function. It is a reward discount factor, in the time slot , Each intelligent agent Get and State Related local observations In the time slot There is only one set of available agents. A decision is needed. Represents intelligent agents random strategy, intelligent agents in based on Get intelligent agent action intelligent agent Reward value obtained after completing the action It is a function of state and action. ,Right now The next time slot According to the state transition function To generate the next state, that is... : ; Step 3.2: Define the observation space, action space, and reward function for all agents: Indicates time slot MC j For nodes The observation information tuple is defined as follows: (7) in MC j To the node The relative position vector, that is, for an idle MC, MC j To the node The relative position vector for occupying MC, MC j At the current time step target node to node The relative position vector; intelligent agent In the time slot observation space The definition is as follows: (8) in Indicates time slot Request a tuple of observation information for all nodes in the pool; intelligent agent In the time slot Action space The definition is as follows: (9) in Table in decision-making time slots Request the set of nodes in the queue; intelligent agent exist End of time slot reward function The definition is as follows: (10) in express MC j Energy used to charge nodes express MC j Energy used for movement In order to be in The node of death , Represents a node Two-dimensional coordinates, This is the penalty coefficient; Step 3.3: The framework of the cooperative charging planning algorithm based on multi-agent deep reinforcement learning is defined as follows: Each MC corresponds to an agent. Idle MCs make action decisions through their corresponding agents. Agents in the action decision state are called decision agents. Independent learning agents use an actor-critic structure. Multiple agents share an actor network and a critic network. The Actor Network (ANN) takes as input the observations of all agents and outputs an action value. The ANN consists of a policy network, an attention channel, and a communication channel. The policy network is divided into two parts, ActorNet(I) and ActorNet(II). The policy network and attention channel are composed of an MLP, and the communication channel is composed of a GRU. The parameters of the CriticNet, ActorNet, Communication Channel, and AttentionUnit are as follows: , , and The decision-making process of an actor network is as follows: Decision slots At that time, ActorNet(I) uses Observation, that is, the local observation of each agent. As input, for the intelligent agent Encoding local observations as intelligent agents The intention of the action is expressed as ,and The attention unit acts as a binary classifier, influencing the decisions of other agents outside the decision-making agent. As input, and to determine these agents Whether to participate in communication, and if so, the decision-making agent will form a communication group with other participating agents. The communication channel will then connect the agents within the communication group. Merging as input and outputting the integrated action intent is represented as , After being fed into the rest of the policy network, ActorNet(II), the output action is... Then, illegal actions are blocked by using a mask. The critic network is parameterized by fully connected layers, taking the observations of all agents as input and outputting the action state Q-value; Step 3.4: During the training of the collaborative charging planning algorithm based on multi-agent deep reinforcement learning, the neural network update process is as follows: The cooperative charging planning algorithm based on multi-agent deep reinforcement learning is an extension of the DDPG algorithm. The experience replay pool is a structure used to store training data, and training data is sampled from the experience replay pool. Update the critic network and the activist network; Update the critic network according to the following formula: (11) The policy gradient is represented as: (12) Using the chain rule, the gradient of the action intent integrated in the communication channel is represented as: (13) The gradient is backpropagated to the policy network and communication channel to update the parameters; Then the target Q network soft update is: (14) The attention unit is trained as a binary classifier for communication, for each decision agent j and its communication group G. j Calculate the difference in average Q-values ​​between communication-based actions and non-communication-based actions, expressed as: : ( (15) Will Stored in queue D, where The weights representing the performance enhancements resulting from communication, when a training cycle ends, represent the weights in D. Performing min-max normalization yields , The labels are used as the binary classifier labels, and the log loss is used to update them. , (16); Step 4: During the training process of the cooperative charging planning algorithm based on multi-agent deep reinforcement learning, the algorithm is trained offline using WRSN historical data to obtain a well-trained deep reinforcement learning model for solving the multi-MC cooperative scheduling problem. Step 5: During the execution of the collaborative charging planning algorithm based on multi-agent deep reinforcement learning, the WRSN state information is input into the trained deep reinforcement learning model, and the charging action value of MC is obtained through model calculation.

2. The cooperative charging planning method based on multi-agent deep reinforcement learning according to claim 1, characterized in that, In step 1, the specific steps for constructing a model of a multi-mobile charger wireless rechargeable sensor network include: The WRSN is deployed within a defined two-dimensional monitoring area, including: One identical sensor node; one base station (BS); one service station (SS); For each identical MC, the time from deployment to the death of all nodes in a WRSN is called its lifecycle, which is divided into several identical time slots. time slot Its duration is short and indivisible; Sensor nodes are randomly distributed within the monitoring area, with fixed locations. Represents a set of nodes, where Indicates the node sequence number. ; express Two-dimensional coordinates; This represents the total energy of the node; Represents a node In the time slot The remaining energy; Indicates the node in the time slot Instantaneous energy consumption rate, This represents the node's time slot from the start of its network lifecycle to the current time slot. The average energy consumption rate of the sensor nodes is dynamically adjusted starting from the initial moment. All nodes collect data and transmit it to the base station via multi-hop forwarding. Due to the unpredictability of events occurring around the sensor nodes and the bursty data streams from the sensor nodes, the energy consumption rate of the sensor nodes changes dynamically in each time slot. The sensor's remaining energy is updated once, using the following formula: (1) Each time slot The average energy consumption rate of the sensor is updated once, as shown in the following formula: (2) The threshold representing the node, per time slot. all The node will send its status information to the base station via multi-hop forwarding. As shown in the following formula: (3) When a node runs out of energy, it will enter a dormant state and will be unable to provide any service to the network. MC is a device with autonomous mobility that can move freely within the WRSN deployment area. MC can obtain its own real-time location. Denotes the set consisting of MC, where Indicates the sequence number of MC. ; express In two-dimensional coordinates, the total energy capacity of MC is MC's movement speed is Mobile energy consumption is Charging power is expressed as Charging efficiency is expressed as MCs are divided into idle MCs and occupied MCs, in each time slot. Idle MCs receive charging target nodes from base stations via long-distance real-time communication and proceed to perform one-to-one charging; while occupied MCs continue to complete their charging tasks. Each charging task of an MC takes several time slots, which are defined as MCs. j A time step, denoted as ;in Indicates the sequence number of the time step, i.e. The start time slot ;j represents the MC ID; time slot MC j Upon receiving the charging task and initiating its execution, due to the uneven distribution of sensor nodes and the varying charging times required, multiple MCs need to perform asynchronous charging. This means that different MCs have different numbers of time slots in their time steps, and their time steps begin and end in different time slots. Each MC can obtain the most recent update from the node that sent the charging request via the base station. Time slot status information In each time slot Each MC sends its own status information to the base station, represented as: (4) in Indicates time slot MC j Location, (t) represents MC j Location of a node that is en route to or currently charging. In MC j The estimated remaining time to complete the current charging task, if MC is in an idle state. and ; The service station has enough power to wirelessly charge the MC. This represents the threshold of MC. If MC is not set after each charging task is completed... i Energy less than If the MC needs to return to the service station to replenish its energy, it cannot perform charging tasks during this period. The base station maintains the state information of low-energy nodes and all MCs. In each time slot Energy lower than The nodes are inserted according to the first-come, first-served principle, with a length of [length missing]. Request queue, if time slot The number of requests is greater than Requests exceeding the request queue length will be discarded. Nodes in the request queue cannot be duplicated, and dead nodes will be removed from the request queue. If the number of nodes in the request queue is less than... If the request queue is empty, the empty slots will be filled with zeros. An empty value in the request queue is an illegal action. If the request queue is not empty and there is an idle MC, the base station will use the MACC algorithm to select the target node for charging for the idle MC in turn according to the legal actions in the request queue, and send it to the corresponding MC. The node selected as the action value will be deleted from the request queue.

3. The cooperative charging planning method based on multi-agent deep reinforcement learning according to claim 2, characterized in that, Step 2, specifically establishing the optimization problem, includes: The charging solution problem is abstracted into a multi-objective optimization problem, where the primary objective is to maximize the energy utilization of all MCs, and the secondary objective is to minimize the number of dead nodes. Maximizing the energy utilization of all MCs is equivalent to minimizing the movement distance of MCs. The objective of the optimization problem is defined as follows: (5) MC j The distance traveled Indicates the liveness status of a node: (6)。 4. The cooperative charging planning method based on multi-agent deep reinforcement learning according to claim 3, characterized in that, In step 4, during the training process of the cooperative charging planning algorithm based on multi-agent deep reinforcement learning, historical data from WRSN is used to train the algorithm offline, obtaining a well-trained deep reinforcement learning model for solving the multi-MC cooperative scheduling problem. Specifically, this includes: First, initialize the parameters of the actor network and the critic network, and initialize the experience replay pool; During the execution phase, in each time slot Base station acquisition If the request queue is not empty and there is an idle MC j Then MC j Corresponding intelligent agent In turn, they become decision-making agents to make action decisions and serve as MCs. j Establish time step Next training data During decision-making, the actor network inputs the observations of all agents. Output action Obtain training data =( ), then MC j The charging process will take several time slots to complete. MC j Earn reward points and the observations of all agents Supplement training data = ( Each agent has its own experience replay pool, which will... The data is stored in the corresponding agent's experience replay pool. The execution phase is repeated until the number of samples in the experience pool reaches a set threshold. Then, the training phase begins. During the training phase, the actor network and the critic network are trained using the experience replay pools of all agents. A certain amount of training data is extracted from the experience pool, and the parameters of the actor network and the critic network are updated using the training data until the neural network parameters converge.

5. The cooperative charging planning method based on multi-agent deep reinforcement learning according to claim 4, characterized in that, In step 5, during the execution of the cooperative charging planning algorithm based on multi-agent deep reinforcement learning, the WRSN state information is input into the trained deep reinforcement learning model, and the solution to the multi-MC cooperative scheduling problem is obtained through model calculation. Specifically, this includes: In each time slot Base station acquisition If the request queue is not empty and there is an idle MC j Then MC j Corresponding intelligent agent In sequence, the agents become decision-making agents and make action decisions. During the decision-making process, the trained agent network inputs the observations of all agents. Output action After that, MC j The charging process will take several time slots to complete.

Citation Information

Patent Citations

  • WSN node intelligent clustering and mobile charging equipment path planning method

    CN110061538A

  • WRSN multi-mobile charger optimal scheduling method based on reinforcement learning

    CN112738752A