A multi-charger carrier separation charging method based on deep reinforcement learning
By using a multi-charger carrier separate charging method based on deep reinforcement learning, the charging-recycling strategy of MCV is optimized, which solves the problems of high sensor node mortality and low energy utilization, and enables long-term operation of sensor networks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- KUNMING UNIV OF SCI & TECH
- Filing Date
- 2022-08-24
- Publication Date
- 2026-05-15
AI Technical Summary
In the prior art, wireless rechargeable sensor networks with multiple wireless chargers carried by MCVs have not been effectively optimized in the charging-recycling strategy, resulting in high sensor node mortality and low energy utilization, and the MCVs have to travel too far during the charging process.
A multi-charger carrier separate charging method based on deep reinforcement learning is adopted. By constructing a target optimization model, the charging-recovery scheduling of MCVs is optimized using deep reinforcement learning algorithms, and charging and recovery actions are selected in a reasonable manner to reduce the mortality rate of sensor nodes and the travel distance of MCVs.
It effectively reduced the mortality rate of sensor nodes, improved energy utilization, and extended the lifespan of sensor networks.
Smart Images

Figure CN115589367B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless power transfer technology for extending the life cycle of wireless sensor networks, and specifically relates to a multi-charger carrier separate charging method based on deep reinforcement learning. Background Technology
[0002] Energy constraints have long been a significant factor limiting the development of Wireless Sensor Networks (WSNs). Because sensors have limited energy capacity, WSNs cannot operate for extended periods, hindering their development and limiting their applications. With the advancement of wireless charging technology, deploying Mobile Charging Vehicles (MCVs) to charge sensors in energy-constrained WSNs has effectively extended the network's lifetime, leading to the emergence of Wireless Rechargeable Sensor Networks (WRSNs). WRSNs are now widely used in military, agricultural production, forest fire prevention, and ecological monitoring, among other fields. However, effectively planning the charging path of the MCV to further extend the WRSN's lifetime remains a crucial research challenge.
[0003] The Wireless Rechargeable Sensor Network (WRSN) incorporates an MCV (Multi-Channel Vehicle) on top of the WSN. The MCV carries multiple portable wireless chargers, while the sensor nodes are equipped with radio frequency (RF) circuitry, enabling them to receive energy from the wireless chargers. During the charging-recycling algorithm's scheduling process, the MCV moves to the sensor node requesting charging to place the wireless charger and begin charging, or moves to the wireless charger node that has completed its charging task to retrieve the charger, allowing the WRSN to operate sustainably.
[0004] In their 2017 paper "Improving Charging Capacity for Wireless Sensor Networks by Deploying One Vehicle with Multiple Removable Chargers" published in *Ad Hoc Networks*, Tao Zou et al. proposed a wireless rechargeable sensor network (WRSN) with multiple removable wireless chargers. In each charging cycle, the MCV (Multi-Vehicle Vehicle) carries several chargers from the base station and sequentially places them at the locations of requesting sensor nodes along a pre-planned path. After each charger completes its charging task, it retrieves all chargers along the same path. Compared to traditional WRSNs, where the MCV needs to stop to charge the sensors, potentially missing important charging opportunities and causing sensor nodes to fail, this invention solves this problem by using an MCV carrying multiple wireless chargers.
[0005] The energy of sensors in the network changes dynamically in real time. The MCV charges the sensors through a pre-planned path, but there may be situations where sensors with high energy consumption rates cannot be charged in time. Moreover, this charging mode of charging first and then recovering energy significantly increases the driving range of the MCV and reduces energy utilization efficiency.
[0006] Current technologies do not utilize deep reinforcement learning to optimize dynamic charging-recovery scheduling strategies for a single MCV carrying multiple detachable chargers in a WRSN. Most existing dynamic charging-recovery strategies only consider separating the charging and recovery paths of the MCV; that is, in a single charging round, the sensor is charged first, then the charger is recovered. They do not consider the combined and overlapping of charging and recovery actions during the charging process, meaning that charging and recovery actions can occur concurrently without a specific order. This strategy could effectively reduce the mortality rate of sensor nodes and the travel distance of the MCV. Summary of the Invention
[0007] To address the aforementioned technical problems, this invention provides a multi-charger carrier separate charging method based on deep reinforcement learning, which can improve energy utilization while effectively reducing the mortality rate of sensor nodes in the network and increasing the lifespan of the sensor network.
[0008] To achieve the above-mentioned technical objectives, the present invention is implemented through the following technical solution:
[0009] A multi-charger carrier separate charging method based on deep reinforcement learning includes the following steps:
[0010] S1: Build a wireless rechargeable sensor model on a two-dimensional plane. The model includes: a base station (BS) with long-distance communication capabilities; a mobile charger vehicle (MCV) carrying m portable wireless chargers, which has loading-unloading, communication, and computing functions; a service station (SS) that can replenish energy for the MCV and wireless chargers; and n homogeneous sensor nodes.
[0011] S2: Establish an objective optimization model with the goal of minimizing the number of dead nodes and the moving distance of the MCV, and design a multi-charger carrier separation charging scheme based on deep reinforcement learning (CSMCSDRL).
[0012] S3: The MCV selects, through the CSMCSDRL algorithm, to move to a sensor node in the target network to place a wireless charger for charging, or to move to a wireless charger node that has completed the charging task to retrieve the charger.
[0013] S4: When the remaining energy of the MCV is not sufficient to support its next action, it retrieves all wireless chargers and returns to the service station, or when the remaining energy of all wireless chargers reaches the threshold, the MCV retrieves all wireless chargers and returns to the service station for energy replenishment to prepare for the next round of charging scheduling; otherwise, execute step S3.
[0014] Preferably, in S1, the base station is deployed in the middle of the two-dimensional plane.
[0015] Preferably, in S1, the maximum battery capacity of the sensor node is E s , and the energy consumption model of node i at time slot t is:
[0016]
[0017] In the formula, f k,i (f i,j )(1 < i < n) kbps is the data stream transmitted from sensor node k(i) to sensor node i(j); f i,B kbps is the data stream transmitted from sensor node i to the base station; ρ is the energy consumption for receiving 1 kb / s of data; c i,j (c i,B ) is the power consumption when sensor node i transmits 1 kb / s of data to sensor node j (BS), and where β is the energy consumption factor, d i,j is the distance between sensor nodes i and j; γ is the signal attenuation factor, and its value is related to many factors. For example, when the sensor nodes are deployed close to the ground, there are many obstacles and interference, and the value of γ is larger. Here, the value of γ is taken as 3; e p and e sLet be the energy consumption of the processor module and the sensor module in the sensor node, respectively; then, in time slot t, the remaining energy of sensor node i is:
[0018]
[0019] Then, in time slot t, the energy requirement of sensor node i is:
[0020]
[0021] In the above formula, λ (0 < λ ≤ 1) is the charging coefficient, and λE s The upper limit for charging the sensor node;
[0022] When the remaining energy of the sensor node is less than the threshold SE th At that time, it sends a charging request to the base station. Where x i and y i These are the x and y coordinates of the sensor node, respectively. It stores its remaining energy; when the sensor node's energy is fully charged, it sends a recharge request to the base station. Where x i and y i The x and y axes represent the horizontal and vertical coordinates of the wireless charger that charges the sensor. This represents the remaining energy of the charger; when the remaining energy of the sensor node is zero, it is marked as dead and will no longer accept energy replenishment from the charger.
[0023] Preferably, the initial energy of the MCV in S1 is E. m J has a speed of vm / s and energy consumption of q during its movement. m J / m, its initial position starts from the service station, and at step k, the remaining energy of MCV is:
[0024]
[0025] In the formula d k Let MCV be the distance traveled at step k.
[0026] Preferably, the m identical wireless chargers carried by the MCV in S1 have an initial energy of E. c J, charging power is q c W, at time slot t, the remaining energy of the wireless charger j is:
[0027]
[0028] The above formula is the model of the remaining energy of wireless charger j charging sensor i; if the wireless charger does not charge any sensor, its remaining energy remains unchanged.
[0029] When the remaining energy of the MCV after performing the next action and recovering all wireless chargers is insufficient to support its return to the service station, or when the remaining energy of all wireless chargers reaches the threshold CE... th At that time, MCV retrieves all wireless chargers and returns them to the service station for recharging, preparing them for the next charging cycle;
[0030] Preferably, the coordinates of the S1 service station are (X... s ,Y s )for:
[0031]
[0032]
[0033] In the above formula, w i This represents the charging frequency of sensor node i, because w i It is positively correlated with the power consumption of the node, therefore it is expressed as Where, p i The average power consumption of the sensor node over one cycle is expressed as: Where T is the length of the charging sequence within the network's lifetime. Let be the power consumption of sensor node i in time slot t; then
[0034] Preferably, the S2 multi-charger carrier separation charging scheme based on deep reinforcement learning is as follows: when the remaining energy of the sensor node reaches the threshold or the wireless charger finishes charging, the sensor will send a charging request or recycling request to the base station in a multi-hop manner. The base station will then send the request information to the MCV. The MCV will add the received request to the request queue, and then select the charging or recycling action according to the node information in the request queue based on deep reinforcement learning, and finally execute the charging or recycling task.
[0035] Preferably, the objective optimization model of S2 is:
[0036]
[0037] Where, N dead Let d be the number of dead sensor nodes in each step, d be the moving distance of MCV in each step, and K be the length of the charging sequence; with minimizing the number of dead sensors as the primary objective.
[0038] Preferably, the specific steps of the CSMCSDRL algorithm in S3 are as follows:
[0039] The agent MCV is in the current state s t Select the optimal action based on the current network with probability ε. Or randomly select an action a with probability 1-ε. t Where A is the action space and a t ∈A; MCV receives a reward r during its interaction with the environment. t And reach a new state s t+1 Then, the experience tuples (s) obtained during the interaction between MCV and the environment are... t ,a t ,r t ,s t+1 ) Store the data in the experience replay pool; when the experience replay pool is full, randomly select a batch of samples from it to train the current Q network, and in each C step, assign the parameter θ of the current Q network to the parameter θ' of the target Q network;
[0040] In the CSMCSDRL algorithm, both sensors and wireless chargers are considered as selectable nodes; the action of MCV can be to transport the wireless charger to the sensor node that requests charging to charge it, or to retrieve the wireless charger node in the network that has completed the charging task.
[0041] Preferably, the reward function is defined as:
[0042] r = e -αd +βN dead (9)
[0043] In the above formula, α is the distance coefficient; β is the penalty factor for sensor node death, which is a negative number;
[0044] Therefore, the total reward for one round of charging scheduling is:
[0045]
[0046] MCVs continuously learn within the network to obtain higher reward values R. total In turn, they learn better strategies to approach the optimal solution.
[0047] The beneficial effects of this invention are:
[0048] This invention optimizes the charging-recycling scheduling of MCVs (Multi-Channel Vehicles) using a deep reinforcement learning algorithm. This algorithm treats both sensors and wireless chargers as selectable nodes, allowing the MCV to consider both current and future rewards, and rationally select charging or recycling actions. This not only prevents sensor nodes from dying due to excessive waiting time but also effectively reduces charging costs. This method improves energy utilization while significantly reducing the mortality rate of sensor nodes in the network, thus extending the network's lifespan. Attached Figure Description
[0049] Figure 1: Schematic diagram of a multi-charger carrier separate charging method based on deep reinforcement learning;
[0050] Figure 2 : A diagram of a wireless rechargeable sensor network model;
[0051] Figure 3 : CSSMSCDRL network structure diagram;
[0052] Figure 4 : Schematic diagram of the charging-recycling scheduling scheme of the present invention. Detailed Implementation
[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] Example 1
[0055] like Figure 1 The diagram shows the algorithm flowchart of the method of this invention. The specific steps of this invention's multi-charger carrier separate charging scheme based on deep reinforcement learning are as follows:
[0056] S1: As Figure 2 The model depicts a wireless rechargeable sensor system: a base station is deployed in the center of a two-dimensional planar region; n homogeneous sensor nodes are randomly distributed in the network, each with a maximum battery capacity of E. s When the remaining energy of the sensor node is less than the threshold SE th When the sensor is fully charged, it sends a charging request to the base station; when the sensor node's energy is fully charged, it sends a recycling request to the base station; when the sensor node's remaining energy is zero, it is marked as dead and no longer accepts energy replenishment from the charger; the initial energy of MCV is E. m J has a speed of vm / s and energy consumption of q during its movement. m J / m, its initial position starts from the service station; the MCV carries m identical wireless chargers, and its initial energy is E. c J, charging power is q c W; Service stations are deployed in areas of the network with high average energy consumption;
[0057] In a wireless rechargeable sensor network, each sensor node and the base station node jointly form a sensor network through a self-organizing method, which is mainly responsible for data collection, forwarding, storage, and processing within the monitoring area; the base station is responsible for receiving data from sensor nodes and further analyzing and processing the data. In addition, the base station is also responsible for communicating with the MCV over a long distance to issue charging tasks; the service station, MCV, and wireless charger form a charging system, which is mainly responsible for providing energy for the sensor network. Among them, the MCV provides energy for sensor nodes by transporting wireless chargers, and the service station provides energy replenishment for the MCV and wireless chargers.
[0058] S2: Aiming to minimize the number of dead nodes and the moving distance of the MCV, an objective optimization model is established, and a carrier separation with multi-charger charging scheme based on deep reinforcement learning (CSMCSDRL) is designed. The working process of this scheme is as follows: When the remaining energy of a sensor node reaches the threshold or the wireless charger completes charging, the sensor sends a charging request or a recycling request to the base station through multi-hop, and the base station then sends the request information to the MCV. The MCV adds the received request to the request queue, and then based on deep reinforcement learning, selects a charging or recycling action according to the node information in the request queue, and finally executes the charging or recycling task;
[0059] S3: The MCV selects to move to the sensor node in the target network to place a wireless charger for charging, or selects to move to the wireless charger node that has completed the charging task to recycle the charger; When the energy of all wireless chargers reaches the threshold or the remaining energy of the MCV is not enough to support its next action and then it recovers all wireless chargers and returns to the service station, the MCV recovers all wireless chargers and returns to the service station to replenish energy, preparing for the next charging scheduling;
[0060] Specifically, the energy consumption model of the sensor node at time slot t is:
[0061]
[0062] In the formula, f k,i (f i,j )(1 < i < n) kbps is the data stream transmitted from sensor node k(i) to sensor node i(j); f i,B kbps is the data stream transmitted from sensor node i to the base station; ρ is the energy consumption of receiving 1 kb / s data; c i,j (c i,B) represents the power consumption when sensor node i transmits 1kb / s of data to sensor node j (BS), and Where β is the energy consumption factor, d i,j denoted as γ, the distance between sensor nodes i and j; γ is the signal attenuation factor, the value of which depends on many factors. For example, when sensor nodes are deployed close to the ground, there are more obstacles and greater interference, so the value of γ will be larger. Here, the value of γ is taken as 3; e p and e s Let be the energy consumption of the processor module and the sensor module in the sensor node, respectively; then, in time slot t, the remaining energy of sensor node i is:
[0063]
[0064] Then, in time slot t, the energy requirement of sensor node i is:
[0065]
[0066] In the above formula, λ (0 < λ ≤ 1) is the charging coefficient, and λE s The upper limit for charging the sensor node.
[0067] When the remaining energy of the sensor node is less than the threshold SE th At that time, it sends a charging request to the base station. Where x i and y i These are the x and y coordinates of the sensor node, respectively. It stores its remaining energy; when the sensor node's energy is fully charged, it sends a recharge request to the base station. Where x i and y i The x and y axes represent the horizontal and vertical coordinates of the wireless charger that charges the sensor. This represents the remaining energy of the charger; when the remaining energy of the sensor node is zero, it is marked as dead and will no longer accept energy replenishment from the charger.
[0068] Specifically, at step k, the remaining energy of MCV is:
[0069]
[0070] In the formula d k Let MCV be the distance traveled at step k.
[0071] At time slot t, the remaining energy of wireless charger j is:
[0072]
[0073] The above formula is the model of the remaining energy of the wireless charger j charging the sensor i;
[0074] When all wireless chargers reach their energy threshold or the MCV's remaining energy is insufficient to support its next action, the MCV will recycle all wireless chargers and return them to the service station to replenish their energy, preparing for the next charging schedule.
[0075] Specifically, the coordinates of the service station are (X... s ,Y s )for:
[0076]
[0077]
[0078] In the above formula, w i This represents the charging frequency of sensor node i, because w i It is positively correlated with the power consumption of the node, therefore it is expressed as Where, p i The average power consumption of the sensor node over one cycle is expressed as: Where T is the length of the charging sequence within the network's lifetime. Let be the power consumption of sensor node i in time slot t; then
[0079] Specifically, the target optimization model of this invention is as follows:
[0080]
[0081] Where, N dead Let d be the number of dead sensor nodes in each step, d be the moving distance of MCV in each step, and K be the length of the charging sequence; with minimizing the number of dead sensors as the primary objective.
[0082] like Figure 3 As shown, the network model in CSMCSDRL consists of two neural networks: a Q-network with parameters θ and a target Q-network with parameters θ'. The agent MCV in the current state s... t Select the optimal action based on the current network with probability ε. Or randomly select an action a with probability 1-ε. t Where A is the action space and a t ∈A; MCV receives a reward r during its interaction with the environment. t And reach a new state s t+1 Then, the experience tuples (s) obtained during the interaction between MCV and the environment are... t ,a t ,r t ,s t+1The samples are stored in the experience replay pool. When the experience replay pool is full, a batch of samples is randomly selected from it to train the current Q network. In each C step, the parameters θ of the current Q network are assigned to the parameters θ' of the target Q network.
[0083] In the CSMCSDRL algorithm, both sensors and wireless chargers are considered as selectable nodes; the action of MCV can be to transport the wireless charger to the sensor node that requests charging to charge it, or to retrieve the wireless charger node in the network that has completed the charging task.
[0084] To enable the agent to learn a better policy, the reward function is defined as:
[0085] r = e -αd +βN dead (9)
[0086] In the above formula, α is the distance coefficient; β is the penalty factor for sensor node death, which is a negative number; d is the movement distance of the MCV in each step; N dead The number of death sensors per step;
[0087] Therefore, the total reward for one round of charging scheduling is:
[0088]
[0089] MCVs continuously learn within the network to obtain higher reward values R. total In turn, they learn better strategies to approach the optimal solution.
[0090] S4: When the remaining energy of the MCV is insufficient to support its next action, it will recycle all wireless chargers and return them to the service station, or when the remaining energy of all wireless chargers reaches the threshold, the MCV will recycle all wireless chargers and return them to the service station for energy replenishment, thus completing one round of charging scheduling.
[0091] Specifically, at step k+1, when or At that time, MCV retrieved all wireless chargers and returned them to the service station for recharging, including L. Hamilton The shortest Hamiltonian path length for MCV to return all wireless chargers to the service station after recycling.
[0092] Example 2
[0093] like Figure 4 As shown, for example, for a period of time, two sensor nodes n2 and n3 have their remaining energy below the threshold SE. th Send charging requests at different times Two sensor nodes, n1 and n4, send recycling requests after their batteries are fully charged. After the base station sends these requests to the MCV and forms a set D(t1) = (s2, s3, c1, c4), the CSMCSDRL algorithm plans a charging sequence based on the location information and remaining energy information in the request set, generating the sequence c1→n2→n3→c4. That is, the MCV first retrieves the wireless charger c1, then moves to sensor nodes n2 and n3 respectively to place the wireless chargers to charge them, and finally retrieves charger c4. This sequence ensures that the MCV's movement path is shortest without sensor nodes n2 and n3 failing.
Claims
1. A multi-charger carrier separate charging method based on deep reinforcement learning, characterized in that, Includes the following steps: S1: Construct a wireless rechargeable sensor model on a two-dimensional plane; the model includes: a base station (BS) with long-range communication capabilities; an MCV carrying m portable wireless chargers, with loading-unloading, communication, and computing capabilities; a service station (SS) that can replenish energy for the MCV and wireless chargers; and n homogeneous sensor nodes. S2: To minimize the number of dead nodes and the moving distance of the MCV, an objective optimization model is established, and a multi-charger carrier separate charging scheme (CSMCSDRL) based on deep reinforcement learning is designed. S3: MCV uses the CSMCSDRL algorithm to select to move to a sensor node in the target network to place a wireless charger to charge it, or to select to move to a wireless charger node that has completed the charging task to retrieve the charger. S4: When the remaining energy of the MCV is insufficient to support its next action, it will recycle all wireless chargers and return them to the service station, or when the remaining energy of all wireless chargers reaches the threshold, the MCV will recycle all wireless chargers and return them to the service station for energy replenishment, in preparation for the next round of charging scheduling; otherwise, proceed to step S3. The maximum battery capacity of the sensor node in S1 is Es, and the energy consumption model of node i in time slot t is: where, fk,i(fi,j) (1 < i < n) kbps is the data stream transmitted from sensor node k(i) to sensor node i(j); fi,B kbps is the data stream transmitted from sensor node i to the base station; p is the energy consumption for receiving 1 kb / s of data; ci,j(ci,B) is the power consumption when sensor node i transmits 1 kb / s of data to sensor node j (BS), and ci,j = βd ,j, where β is the energy consumption factor and di,j is the distance between sensor nodes i and j; Y is the signal attenuation factor, and its value is related to many factors. When the sensor nodes are deployed close to the ground, there are many obstacles and large interference, and the value of Y is larger. Here, the value of Y is taken as 3; ep and es are the energy consumptions of the processor module and the sensor module in the sensor node respectively. Then, at time slot t, the remaining energy of sensor node i is: p (t) = p (t_1) _ pi (t) (2) Then, in time slot t, the energy requirement of sensor node i is: p (t ) = λEs _ p (t ) (3) In the above formula, λ (0 < λ ≤ 1) is the charging coefficient, and λEs is the upper limit of the charging of the sensor node; When the remaining energy of a sensor node is less than the threshold SEth, it sends a charging request si = (xi, yi, p) to the base station. ), where xi and yi are the x and y coordinates of the sensor node, respectively, and p Its remaining energy; when the sensor node's energy is fully charged, it sends a recharge request ci = (xi, yi, p) to the base station. i), where xi and yi are the x and y coordinates of the wireless charger that charges the sensor, respectively, and p i represents the remaining energy of the charger; when the remaining energy of the sensor node is zero, it is marked as dead and will no longer accept energy replenishment from the charger. The S2 multi-charger carrier separation charging scheme based on deep reinforcement learning is as follows: when the remaining energy of the sensor node reaches the threshold or the wireless charger finishes charging, the sensor will send a charging request or recycling request to the base station through a multi-hop method. The base station will then send the request information to the MCV. The MCV will add the received request to the request queue, and then select the charging or recycling action according to the node information in the request queue based on deep reinforcement learning, and finally execute the charging or recycling task. The network model in the S3 CSMCSDRL consists of two neural networks: one is a Q network with parameter θ, and the other is a target Q network with parameter θ'. The specific steps of the CSMCSDRL algorithm in S3 are as follows: In the current state st, the agent MCV selects the optimal action at = m with probability ε based on the current network state. x Q(st, A;θ), or randomly select an action at with probability 1 - ε, where A is the action space and at ∈ A; the MCV obtains a reward rt during its interaction with the environment and reaches a new state st+1; then the experience tuple (st, at, rt, st+1) obtained by the MCV during its interaction with the environment is stored in the experience replay pool; when the experience replay pool is full, a batch of samples is randomly selected from it to train the current Q network, and at each C step, the parameter θ of the current Q network is assigned to the parameter θ' of the target Q network; In the CSMCSDRL algorithm, both sensors and wireless chargers are considered selectable nodes; the action of MCV is to transport the wireless charger to the sensor node that requests charging to charge it, or to retrieve the wireless charger node in the network that has completed the charging task.
2. The multi-charger carrier separate charging method based on deep reinforcement learning according to claim 1, characterized in that, The base station in S1 is deployed in the middle of a two-dimensional plane.
3. The multi-charger carrier separate charging method based on deep reinforcement learning according to claim 1, characterized in that, In S1, the initial energy of the MCV is EmJ, its speed is vm / s, and its energy consumption during movement is qmJ / m. Its initial position starts from the service station. At step k, the remaining energy of the MCV is: p (k ) = p (k _1) _ dk qm (4) In the formula, dk is the distance traveled by the MCV at the k-th step.
4. The multi-charger carrier separate charging method based on deep reinforcement learning according to claim 1, characterized in that, In S1, the MCV carries m identical wireless chargers with an initial energy of EcJ and a charging power of qcW. At time slot t, the remaining energy of wireless charger j is: p j (t ) = p j (t _1) _ p (t ) (5) The above formula represents the remaining energy model of the wireless charger j charging sensor i; if the wireless charger does not charge any sensor, its remaining energy remains unchanged: When the remaining energy of the MCV after performing the next action and recovering all wireless chargers is insufficient to support its return to the service station, or when the remaining energy of all wireless chargers reaches the threshold CEth, the MCV recovers all wireless chargers and returns them to the service station for energy replenishment, preparing for the next round of charging.
5. The multi-charger carrier separate charging method based on deep reinforcement learning according to claim 1, characterized in that, The coordinates of the S1 service station are (XS, YS): In the above formula, "i" represents the charging frequency of sensor node i. Since "i" is positively correlated with the power consumption of the node, it is expressed as: Where pi is the average power consumption of the sensor node over one cycle, expressed as pi Where T represents the charging time during the network's lifespan. The length of the sequence, p Let be the power consumption of sensor node i in time slot t; then wi 6. The multi-charger carrier separate charging method based on deep reinforcement learning according to claim 1, characterized in that, The objective optimization model for S2 is as follows: Where Ndead is the number of dead sensor nodes in each step, d is the moving distance of MCV in each step, and K is the length of the charging sequence; and the main objective is to minimize the number of dead sensors.