A DRL-based method for optimizing the data upload path of wirelessly rechargeable drones.

By equipping drones with a mobile charger featuring a high-gain RF antenna and an improved Double DQN algorithm, the drone path is optimized, solving the problem of insufficient drone battery life and enabling efficient uploading of IoT communication data.

CN117062182BActive Publication Date: 2026-05-26ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ANHUI UNIV
Filing Date
2023-09-04
Publication Date
2026-05-26

Smart Images

  • Figure CN117062182B_ABST
    Figure CN117062182B_ABST
Patent Text Reader

Abstract

This invention relates to a DRL-based method for optimizing the data upload path of a wirelessly rechargeable drone, comprising: constructing an IoT communication and wireless charging scenario system; establishing a first mathematical model for the onboard battery consumption of the mission drone; establishing a second mathematical model for the data upload channel; establishing a third mathematical model for energy replenishment during the wireless charging process; establishing a fourth mathematical model for the path optimization objective; and determining the state set S, the action set A, and the reward function r. t The optimal path strategy π is obtained through offline learning using an improved Double DQN algorithm. * This invention provides a more convenient charging service for task units that assist base stations in performing data uploads in IoT communication systems. When data real-time requirements are high, the more convenient charging method significantly improves the endurance of the task units, achieving high data upload efficiency in IoT communication systems assisted by drones.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence, power and communication technologies, and in particular to a method for optimizing the data upload path of a wireless rechargeable drone based on DRL. Background Technology

[0002] Thanks to their highly flexible deployment capabilities, drones have been widely used in emerging fields such as the Internet of Things (IoT) in recent years. They aim to compensate for the shortcomings of existing base station communications, acting as mobile communication access points to connect ground users and provide emergency data services, ensuring the quality of communication services for mobile users in large-scale scenarios. However, limited by the energy constraints of onboard batteries, drones face a severe energy shortage problem when providing services.

[0003] The inefficiency caused by the energy constraints of drones is mainly manifested in the fact that as time and missions continue, the individual drone's energy is constantly consumed, requiring it to return to a charging station at some point during mission execution, which cannot meet the high real-time data upload requirements. With the emergence of various new energy supply technologies, battery replenishment methods have made significant progress. Among them, wireless power transmission technology can decouple the location of energy sources from the location of sensors, transferring energy from energy-rich areas to energy-poor areas, allowing drones to effectively harvest energy while performing data transmission tasks.

[0004] Equipping drones with high-gain radio frequency antennas as mobile chargers to provide charging services for drones performing data upload tasks can be an effective energy solution. In the communication context of mission-assisted IoT systems, which involve both user data upload and on-demand charging, an effective drone path planning strategy is needed to ensure the efficient execution of individual drone data upload tasks.

[0005] While researchers have conducted extensive studies on optimizing data upload paths for drones, employing algorithms such as ant colony optimization, genetic algorithms, and reinforcement learning to find optimal paths, most of these studies focus on optimizing flight paths for missions supported by onboard power. They primarily consider the energy utilization rate of a single flight without recharging, neglecting the need for energy replenishment technologies to further meet the requirements of efficient mission execution. Therefore, there is an urgent need to develop a data upload path optimization method for wirelessly rechargeable drones. This method should meet the real-time data upload needs of IoT devices while also considering wireless energy replenishment to improve the drone's single-flight lifespan. This approach has significant research and application value for optimizing flight paths for mission drones. Summary of the Invention

[0006] To address the inefficiency caused by drone energy shortages in existing mission-oriented drones (UAVs) for communication in IoT systems assisted by mission drones, this invention aims to provide a DRL-based wireless rechargeable UAV data upload path optimization method that significantly improves the endurance of mission drones and achieves high data upload efficiency for UAV-assisted IoT communication systems under conditions of high data real-time requirements.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: a method for optimizing the data upload path of a wireless rechargeable drone based on DRL, the method comprising the following sequential steps:

[0008] (1) Construct an IoT communication and wireless charging scenario system, which includes a mission machine, a mobile charger and M mobile IoT devices;

[0009] (2) Establish a first mathematical model for the consumption of the mission aircraft's onboard battery;

[0010] (3) Establish a second mathematical model for the data upload channel;

[0011] (4) Establish a mathematical model for energy replenishment during the wireless charging process, i.e., the third mathematical model;

[0012] (5) Based on the IoT communication and wireless charging scenario system, the first mathematical model, the second mathematical model and the third mathematical model, a fourth mathematical model is established for the path optimization objective;

[0013] (6) Based on the IoT communication and wireless charging scenario system, the first mathematical model, the second mathematical model, the third mathematical model, and the fourth mathematical model, determine the state set S, the action set A, and the reward function. ;

[0014] (7) Based on the state set S, action set A, and reward function The optimal path strategy is obtained by using an improved Double DQN algorithm for offline learning. .

[0015] Step (1) specifically refers to: The IoT communication and wireless charging scenario system is denoted as the system. The system space is divided into an N*N grid, where each grid is a square unit with a side length of c, and N is the number of grids in the horizontal and vertical directions. A task machine is deployed within the scenario to perform the upload task, a mobile charger provides wireless charging service, and M mobile IoT devices each have a value... The amount of data is waiting to be uploaded;

[0016] Let the current time slot of the system be t, where t=0. ,2 …T, Let T be the length of a single time slot, and T be the moment when the system terminates. Then, the position of mobile IoT device m at time slot t is: It is represented as, where m∈{1,2…M}, This represents the x-coordinate of the mobile IoT device m in the grid space. This represents the ordinate of the mobile IoT device m in the grid space. This represents the fixed height of the mobile IoT device m; assuming the remaining amount of data to be uploaded by the mobile IoT device m in time slot t is represented as... ;

[0017] The mobile charger, acting as an energy supplier, departs from its docking point at the start of the mission, following a pre-determined flight path and speed. Move and provide power to the mission machine, real-time location information.

[0018] express, This represents the x-coordinate of the mobile charger in the grid space. This represents the ordinate of the mobile charger in the grid space. A fixed flight altitude for the mobile charger; at the start of the mission, the mission aircraft departs from the docking point and flies at a constant speed. Flight, real-time location It means that, among them, This represents the x-coordinate of the mission machine in the grid space. This represents the ordinate of the mission machine in the grid space. This refers to the fixed flight altitude of the mission aircraft.

[0019] Step (2) specifically refers to: the maximum battery capacity of the mission machine is , This represents the remaining charge in time slot t. The mission aircraft consumes a constant amount of energy for each flight maneuver it performs. Without considering charging, a mathematical model is established for the battery level of the mission machine in time slot t+1, denoted as the first mathematical model, with the following expression:

[0020]

[0021] In the formula, This indicates the remaining battery power in time slot t+1.

[0022] Step (3) specifically refers to: during the process of establishing a communication connection between the mission machine and the mobile IoT device and starting data upload, the mission machine flies at a sufficient altitude. At this time, the line-of-sight wireless transmission communication between the mobile IoT device and the mission machine is guaranteed. The expression for the channel gain between the two is as follows:

[0023]

[0024] in, This represents the channel gain at a reference distance of 1m. The Euclidean distance between the mission machine and the mobile IoT device m in time slot t is expressed as follows:

[0025]

[0026] in, This represents the Euclidean distance between the mission machine and the mobile IoT device m on the horizontal axis. This represents the Euclidean distance between the mission machine and the mobile IoT device m on the vertical axis; The mission aircraft is at a fixed flight altitude; a mathematical model for data transmission at time slot t is established, denoted as the second mathematical model:

[0027]

[0028] in, W is the transmission rate of the data transmission link established between the task machine in time slot t and the mobile IoT device m, where W is the signal bandwidth. It refers to the transmit power of IoT devices. It is noise power.

[0029] Step (4) specifically refers to: the mobile charger is equipped with a high-gain radio frequency antenna for transmitting wireless power, providing energy supply services to the mission machine via a fixed deployment trajectory; during mission execution, the distance between the mission machine and the mobile charger varies in different time slots, and the power of wireless charging will drop sharply when the distance between them increases. Assuming full energy conversion efficiency in the wireless transmission link, the power obtained at the mission machine is calculated using the Friis free space propagation model. Expressed as follows:

[0030]

[0031] in, It is the transmitter power. and It refers to the antenna gain at the transmitting and receiving ends. It is the transmission wavelength. It is the Euclidean distance between the transmitter and receiver, expressed as follows:

[0032]

[0033] in, This represents the Euclidean distance between the mission unit and the mobile charger in the horizontal direction within time slot t. This represents the Euclidean distance between the mission unit and the mobile charger in the longitudinal direction along time slot t. The constant height difference between the mission device and the mobile charger is represented; a mathematical model for energy replenishment during the wireless charging process is established and denoted as the third mathematical model, the expression of which is as follows:

[0034]

[0035] In the formula, This represents the energy received by the mission machine in a single time slot. The length of a single time slot.

[0036] Step (5) specifically refers to the following: In the dynamic scenario of optimizing the mission vehicle path planning problem, the variables are the location information of the mobile IoT device, the mobile charger, and the mission vehicle, the device data upload queue, and the mission vehicle battery status information; the optimization goal is to find a path strategy to help the mission vehicle make the best decision between balancing energy consumption and the amount of data uploaded, and to maximize the data upload efficiency during a single flight from the perspective of optimizing the movement trajectory. The main factors considered in this process are the data upload amount and mission energy consumption. The expression of the fourth mathematical model is as follows:

[0037]

[0038] in, It refers to the transmission rate of the data transmission link established between the task machine in time slot t and the mobile IoT device m. The mission aircraft consumes a constant amount of energy for each flight maneuver; let the current time slot of the system be t, t=0, ,2 …T, Let m be the length of a single time slot, and T be the time when the system terminates; m∈{1,2…M}.

[0039] Step (6) specifically includes the following steps:

[0040] (6a) Determine the set of states S:

[0041] The expression for the state set S is as follows:

[0042]

[0043] in, The system state at time slot t is determined by... , , It consists of three parts; The status of each mobile IoT device in the data upload task is represented to guide the data upload task objectives. This includes the location and remaining data volume information of all mobile IoT devices. Let M be the two-dimensional coordinates of the last mobile IoT device M in the region during the upload task in the horizontal direction on time slot t. , With the current remaining amount of data to be uploaded ,but The expression is as follows:

[0044]

[0045] By simulating the field of view of the mission machine, the data transmission rate of mobile IoT devices at different locations within the effective transmission distance of the mission machine is represented in relation to the location of the mission machine. Based on the system characteristics, with the mission machine as the center, it is assumed that the field of view range is n grids. The location outside the mission area is represented by black grids, and the corresponding matrix value is set to -1 to indicate that it is outside the mission area. The data upload rate at different locations is represented based on the difference in distance from the mission machine. The upload rate at the mission machine is the highest, and the corresponding matrix element is set to 50. The matrix value corresponding to the distance of one grid is set to 20, the matrix value corresponding to the distance of two grids is set to 5, and the matrix value corresponding to the remaining white grids is set to 0.

[0046] Then a new vision matrix is ​​constructed. , Let be the value of the element in the i-th row and j-th column of the field of view matrix, where And assume that the spatial position of the i-th row and j-th column in the field of view matrix Y corresponds to the horizontal and vertical coordinates in the overall grid space as follows: , , Let the distance from the grid of the mission machine to the position of the element in the i-th row and j-th column of the field of view matrix at time slot t be the distance of the field of view matrix element. The numerical expression is as follows:

[0047]

[0048] The obtained field-of-view matrix Y contains A square matrix of elements, flattening the elements in the field of view matrix Y, and the horizontal coordinate of the joint mission machine in the grid space. The ordinate of the mission machine in the grid space Form a one-dimensional column vector The expression is as follows:

[0049]

[0050] Including the current battery level of the mission machine The x-coordinate of the mobile charger in the grid space The vertical coordinate of the mobile charger in the grid space And the Euclidean distance between the mission machine and the mobile charger. , The expression is as follows:

[0051]

[0052] (6b) Determine the action set A:

[0053] The task machine in the system has four possible movement directions in the gridded space, and the motion space expression is as follows:

[0054]

[0055] in, The action performed by the mission aircraft in time slot t, and at a constant speed during flight. There are four flight directions; during flight, energy is obtained from the mobile charger transmitter via wireless charging, and the two processes of scene data uploading and energy acquisition are carried out in parallel.

[0056] (6c) Determine the reward function :

[0057] For the task machine, the impact of different behavioral decisions on various aspects of the system is reflected in the reward mechanism. The immediate reward obtained by the task machine in time slot t is represented as... The expression is as follows:

[0058]

[0059] in, , , To adjust , , Weighting factors between them; With the core objective of maximizing data throughput throughout the entire data upload process, a positive reward is assigned to each action based on the data throughput during a single communication session, as expressed in the following expression:

[0060]

[0061] in, This represents the amount of data uploaded by the mobile IoT device m to the task machine at time t;

[0062] With the goal of maximizing the data upload efficiency of the entire data upload process, negative rewards are given to the task machine's movement behavior at each step to urge the task machine to improve its path selection ability, reduce unnecessary energy loss and promote the convergence of the optimal path.

[0063] With the goal of maximizing the battery life of the task machine during data upload tasks, a positive behavioral reward is assigned to the wireless charging resulting from the task machine's movement decisions. The expression for this reward is as follows:

[0064]

[0065] in, This represents the remaining battery power of the mission machine at time slot t. This represents the energy received by the mission machine in a single time slot. The threshold for determining whether the mission machine has entered a low-energy state. and All are constant coefficients.

[0066] Step (7) specifically includes the following steps:

[0067] (7a) Initialize the estimated network neural parameters and target value network neural parameters ,make Initialize the experience replay pool with a capacity of D; initialize the network learning rate. attenuation coefficient ;

[0068] (7b) Based on the current state-action pair input , The system state at time slot t. For the action performed by the task machine in time slot t, the output is an estimate of the Q value, i.e. ,in The network generates the estimated values ​​for the neural parameters of the target value network; the target value network generates the action to select the next state. ,in The state values ​​are the input values ​​to the target value network. The action value is input to the target value network, and then the next state-action pair of the target value network is determined. Q value The objective value of the Double DQN algorithm is defined as:

[0069]

[0070] (7c) Based on the current state Execute action It then selects actions based on the improved Double DQN algorithm to transition to a new state. According to the reward function Calculate the single-step reward value and store the conversion in the experience replay pool. ;

[0071] Repeat steps (7b) to (7c) until the number of memories stored in the experience replay pool equals D, then proceed to step (7d).

[0072] The Double DQN algorithm has two neural network structures: an estimation network and a target value network. Given a training step size step, training is performed once every step.

[0073] The total reward in each training round is defined as R, and the reward value obtained by executing a single flight path is expressed as follows:

[0074]

[0075] The improved Double DQN algorithm refers to the greedy coefficient for action selection. Improvements are made by ensuring the action that maximizes the reward value at each step based on the optimization metric. During the execution of the ε-greedy policy, the AI ​​perceives the action as numerically... The probability of choosing the current optimal solution is given, and the remaining 1- The probability of exploring other actions. The numerical changes of the coefficients are expressed as follows:

[0076]

[0077] in, Let K be the current flight round number. The maximum number of flight rounds reached at that time;

[0078] The training dataset for offline learning is formed by randomly selecting z experiences from an experience replay pool of size D; the maximum number of task rounds is F.

[0079] (7d) Randomly select z memories from the experience replay pool, z=32; denote the k-th state transition sequence in the mini-batch data as ;

[0080] (7e) Obtained according to step (7d) Calculate the target Q value and loss value ,make The loss function is expressed as follows:

[0081]

[0082] (7f) Minimize the loss function using gradient descent It can be expressed as follows:

[0083]

[0084]

[0085] In the formula, The sign of the partial derivative. Represents minimizing the error function right Find the partial derivative. The input to the estimation network is... , The squared difference between the calculated Q-value and the target value network's Q-value affects the estimated network's neural parameters. Find the partial derivative. To update the estimated network parameters after the update, the target value network parameters are updated every step. Replace with ;

[0086] When steps (7d) to (7f) are completed once, a single training session is completed. Steps (7b) to (7f) are repeated during each round of flight mission execution. After F rounds, the training is completed, and the learning process of the Double DQN algorithm ends.

[0087] (7g) The training algorithm ends, and the optimal path strategy is saved. Record reward values, flight steps, and data uploads: estimate network neural parameters over F training rounds. and target value network neural parameters The update strategy is performed in the direction of maximizing the total reward value R, eventually finding the optimal path strategy. .

[0088] As can be seen from the above technical solution, the beneficial effects of the present invention are as follows: First, compared with the prior art, by utilizing far-field wireless charging technology and combining a highly mobile drone equipped with a high-gain radio frequency antenna as a mobile charger, a more convenient charging service is provided for the task machine that assists the base station in the Internet of Things communication system to perform data uploading. Under the condition of high data real-time requirements, the endurance of the task machine is significantly improved through a more convenient charging method, and high data uploading efficiency of the Internet of Things communication system assisted by drones is achieved. Second, based on the high-dimensionality of the problem state space, the present invention uses the deep reinforcement learning method Double DQN with a dual deep neural network structure in the path optimization method and improves the exploration coefficient. This method enables the task machine to learn autonomously by interacting with the environment, adjust the action selection according to the reward signal, and thus gradually optimize the path. At the same time, it can also adapt to changes in the environment and task, and achieve better adaptability and generalization ability. Attached Figure Description

[0089] Figure 1 This is a flowchart of the method of the present invention;

[0090] Figure 2 This is a diagram illustrating a grid scene in an embodiment of the present invention;

[0091] Figure 3 This is a model block diagram of the improved Double DQN algorithm in this invention;

[0092] Figure 4 This is a graph showing the convergence effect of reward value training in this invention;

[0093] Figure 5 This is a training effect diagram of the number of flight steps in a single flight mission in this invention;

[0094] Figure 6 This is a training effect diagram of the data payload of a single flight mission in this invention. Detailed Implementation

[0095] like Figure 1 As shown, a method for optimizing the data upload path of a wireless rechargeable drone based on DRL is proposed. This method includes the following steps in sequence:

[0096] (1) Construct an IoT communication and wireless charging scenario system, which includes a mission machine, a mobile charger and M mobile IoT devices;

[0097] (2) Establish a first mathematical model for the consumption of the mission aircraft's onboard battery;

[0098] (3) Establish a second mathematical model for the data upload channel;

[0099] (4) Establish a mathematical model for energy replenishment during the wireless charging process, i.e., the third mathematical model;

[0100] (5) Based on the IoT communication and wireless charging scenario system, the first mathematical model, the second mathematical model and the third mathematical model, a fourth mathematical model is established for the path optimization objective;

[0101] (6) Based on the IoT communication and wireless charging scenario system, the first mathematical model, the second mathematical model, the third mathematical model, and the fourth mathematical model, determine the state set S, the action set A, and the reward function. ;

[0102] (7) Based on the state set S, action set A, and reward function The optimal path strategy is obtained by using an improved Double DQN algorithm for offline learning. .

[0103] Step (1) specifically refers to: The IoT communication and wireless charging scenario system is denoted as the system. The system space is divided into an N*N grid, where each grid is a square unit with a side length of c, and N is the number of grids in the horizontal and vertical directions. A task machine is deployed within the scenario to perform the upload task, a mobile charger provides wireless charging service, and M mobile IoT devices each have a value... The amount of data is waiting to be uploaded;

[0104] Let the current time slot of the system be t, where t=0. ,2 …T, Let T be the length of a single time slot, and T be the moment when the system terminates. Then, the position of mobile IoT device m at time slot t is: It is represented as, where m∈{1,2…M}, This represents the x-coordinate of the mobile IoT device m in the grid space. This represents the ordinate of the mobile IoT device m in the grid space. This represents the fixed height of the mobile IoT device m; assuming the remaining amount of data to be uploaded by the mobile IoT device m in time slot t is represented as... ;

[0105] The mobile charger, acting as an energy supplier, departs from its docking point at the start of the mission, following a pre-determined flight path and speed. Move and provide power to the mission machine, real-time location information.

[0106] express, This represents the x-coordinate of the mobile charger in the grid space. This represents the ordinate of the mobile charger in the grid space. A fixed flight altitude for the mobile charger; at the start of the mission, the mission aircraft departs from the docking point and flies at a constant speed. Flight, real-time location It means that, among them, This represents the x-coordinate of the mission machine in the grid space. This represents the ordinate of the mission machine in the grid space. This is the fixed flight altitude for the mission aircraft. Here, meters per second rice, meters per second rice.

[0107] Step (2) specifically refers to: the maximum battery capacity of the mission machine is , This represents the remaining charge in time slot t. The mission aircraft consumes a constant amount of energy for each flight maneuver it performs. Without considering charging, a mathematical model is established for the battery level of the mission machine in time slot t+1, denoted as the first mathematical model, with the following expression:

[0108]

[0109] In the formula, This indicates the remaining battery power in time slot t+1.

[0110] Here, kilojoules 1,000 joules.

[0111] Step (3) specifically refers to: during the process of establishing a communication connection between the mission machine and the mobile IoT device and starting data upload, the mission machine flies at a sufficient altitude. At this time, the line-of-sight wireless transmission communication between the mobile IoT device and the mission machine is guaranteed. The expression for the channel gain between the two is as follows:

[0112]

[0113] in, This represents the channel gain at a reference distance of 1m. The Euclidean distance between the mission machine and the mobile IoT device m in time slot t is expressed as follows:

[0114]

[0115] in, This represents the Euclidean distance between the mission machine and the mobile IoT device m on the horizontal axis. This represents the Euclidean distance between the mission machine and the mobile IoT device m on the vertical axis; The mission aircraft is at a fixed flight altitude; a mathematical model for data transmission at time slot t is established, denoted as the second mathematical model:

[0116]

[0117] in, W is the transmission rate of the data transmission link established between the task machine in time slot t and the mobile IoT device m, where W is the signal bandwidth. It refers to the transmit power of IoT devices. It is noise power.

[0118] Here, , , , .

[0119] Step (4) specifically refers to: the mobile charger is equipped with a high-gain radio frequency antenna for transmitting wireless power, providing energy supply services to the mission machine via a fixed deployment trajectory; during mission execution, the distance between the mission machine and the mobile charger varies in different time slots, and the power of wireless charging will drop sharply when the distance between them increases. Assuming full energy conversion efficiency in the wireless transmission link, the power obtained at the mission machine is calculated using the Friis free space propagation model. Expressed as follows:

[0120]

[0121] in, It is the transmitter power. and It refers to the antenna gain at the transmitting and receiving ends. It is the transmission wavelength. It is the Euclidean distance between the transmitter and receiver, expressed as follows:

[0122]

[0123] in, This represents the Euclidean distance between the mission unit and the mobile charger in the horizontal direction within time slot t. This represents the Euclidean distance between the mission unit and the mobile charger in the longitudinal direction along time slot t. The constant height difference between the mission device and the mobile charger is represented; a mathematical model for energy replenishment during the wireless charging process is established and denoted as the third mathematical model, the expression of which is as follows:

[0124]

[0125] In the formula, This represents the energy received by the mission machine in a single time slot. The length of a single time slot.

[0126] Here, , , , rice.

[0127] Step (5) specifically refers to the following: In the dynamic scenario of optimizing the mission vehicle path planning problem, the variables are the location information of the mobile IoT device, the mobile charger, and the mission vehicle, the device data upload queue, and the mission vehicle battery status information; the optimization goal is to find a path strategy to help the mission vehicle make the best decision between balancing energy consumption and the amount of data uploaded, and to maximize the data upload efficiency during a single flight from the perspective of optimizing the movement trajectory. The main factors considered in this process are the data upload amount and mission energy consumption. The expression of the fourth mathematical model is as follows:

[0128]

[0129] in, It refers to the transmission rate of the data transmission link established between the task machine in time slot t and the mobile IoT device m. The mission aircraft consumes a constant amount of energy for each flight maneuver; let the current time slot of the system be t, t=0, ,2 …T, Let m be the length of a single time slot, and T be the time when the system terminates; m∈{1,2…M}.

[0130] Step (6) specifically includes the following steps:

[0131] (6a) Determine the set of states S:

[0132] The expression for the state set S is as follows:

[0133]

[0134] in, The system state at time slot t is determined by... , , It consists of three parts; The status of each mobile IoT device in the data upload task is represented to guide the data upload task objectives. This includes the location and remaining data volume information of all mobile IoT devices. Let M be the two-dimensional coordinates of the last mobile IoT device M in the region during the upload task in the horizontal direction on time slot t. , With the current remaining amount of data to be uploaded ,but The expression is as follows:

[0135]

[0136] By simulating the field of view of the mission machine, the data transmission rate of mobile IoT devices at different locations within the effective transmission distance of the mission machine is represented in relation to the location of the mission machine. Based on the system characteristics, with the mission machine as the center, it is assumed that the field of view range is n grids. The location outside the mission area is represented by black grids, and the corresponding matrix value is set to -1 to indicate that it is outside the mission area. The data upload rate at different locations is represented based on the difference in distance from the mission machine. The upload rate at the mission machine is the highest, and the corresponding matrix element is set to 50. The matrix value corresponding to the distance of one grid is set to 20, the matrix value corresponding to the distance of two grids is set to 5, and the matrix value corresponding to the remaining white grids is set to 0.

[0137] Then a new vision matrix is ​​constructed. , Let be the value of the element in the i-th row and j-th column of the field of view matrix, where And assume that the spatial position of the i-th row and j-th column in the field of view matrix Y corresponds to the horizontal and vertical coordinates in the overall grid space as follows: , , Let the distance from the grid of the mission machine to the position of the element in the i-th row and j-th column of the field of view matrix at time slot t be the distance of the field of view matrix element. The numerical expression is as follows:

[0138]

[0139] The obtained field-of-view matrix Y contains A square matrix of elements, flattening the elements in the field of view matrix Y, and the horizontal coordinate of the joint mission machine in the grid space. The ordinate of the mission machine in the grid space Form a one-dimensional column vector The expression is as follows:

[0140]

[0141] Including the current battery level of the mission machine The x-coordinate of the mobile charger in the grid space The vertical coordinate of the mobile charger in the grid space And the Euclidean distance between the mission machine and the mobile charger. , The expression is as follows:

[0142]

[0143] (6b) Determine the action set A:

[0144] The task machine in the system has four possible movement directions in the gridded space, and the motion space expression is as follows:

[0145]

[0146] in, The action performed by the mission aircraft in time slot t, and at a constant speed during flight. There are four flight directions; during flight, energy is obtained from the mobile charger transmitter via wireless charging, and the two processes of scene data uploading and energy acquisition are carried out in parallel.

[0147] (6c) Determine the reward function :

[0148] For the task machine, the impact of different behavioral decisions on various aspects of the system is reflected in the reward mechanism. The immediate reward obtained by the task machine in time slot t is represented as... The expression is as follows:

[0149]

[0150] in, , , To adjust , , The weighting factors between them; here, , , ; With the core objective of maximizing data throughput throughout the entire data upload process, a positive reward is assigned to each action based on the data throughput during a single communication session, as expressed in the following expression:

[0151]

[0152] in, This represents the amount of data uploaded by the mobile IoT device m to the task machine at time t;

[0153] With the goal of maximizing the data upload efficiency of the entire data upload process, negative rewards are given to the task machine's movement behavior at each step to urge the task machine to improve its path selection ability, reduce unnecessary energy loss and promote the convergence of the optimal path.

[0154] With the goal of maximizing the battery life of the task machine during data upload tasks, a positive behavioral reward is assigned to the wireless charging resulting from the task machine's movement decisions. The expression for this reward is as follows:

[0155]

[0156] in, This represents the remaining battery power of the mission machine at time slot t. This represents the energy received by the mission machine in a single time slot. The threshold for determining whether the mission machine has entered a low-energy state. and All are constant coefficients.

[0157] Step (7) specifically includes the following steps:

[0158] (7a) Initialize the estimated network neural parameters and target value network neural parameters ,make Initialize the experience replay pool with a capacity of D; initialize the network learning rate. attenuation coefficient Here, , .

[0159] (7b) Based on the current state-action pair input , The system state at time slot t. For the action performed by the task machine in time slot t, the output is an estimate of the Q value, i.e. ,in The network generates the estimated values ​​for the neural parameters of the target value network; the target value network generates the action to select the next state. ,in The state values ​​are the input values ​​to the target value network. The action value is input to the target value network, and then the next state-action pair of the target value network is determined. Q value The objective value of the Double DQN algorithm is defined as:

[0160]

[0161] (7c) Based on the current state Execute action It then selects actions based on the improved Double DQN algorithm to transition to a new state. According to the reward function Calculate the single-step reward value and store the conversion in the experience replay pool. ;

[0162] Repeat steps (7b) to (7c) until the number of memories stored in the experience replay pool equals D, then proceed to step (7d).

[0163] The Double DQN algorithm has two neural network structures: an estimation network and a target value network. Given a training step size step, training is performed once every step.

[0164] The total reward in each training round is defined as R, and the reward value obtained by executing a single flight path is expressed as follows:

[0165]

[0166] Here, K=8000, D=55000, z=32, F=16000, step=25.

[0167] like Figure 3 As shown, the improved Double DQN algorithm refers to the greedy coefficient for action selection. Improvements are made by ensuring the action that maximizes the reward value at each step based on the optimization metric. During the execution of the ε-greedy policy, the AI ​​perceives the action as numerically... The probability of choosing the current optimal solution is given, and the remaining 1- The probability of exploring other actions. The numerical changes of the coefficients are expressed as follows:

[0168]

[0169] in, Let K be the current flight round number. The maximum number of flight rounds reached at that time;

[0170] The training dataset for offline learning is formed by randomly selecting z experiences from an experience replay pool of size D; the maximum number of task rounds is F.

[0171] (7d) Randomly select z memories from the experience replay pool, z=32; denote the k-th state transition sequence in the mini-batch data as ;

[0172] (7e) Obtained according to step (7d) Calculate the target Q value and loss value ,make The loss function is expressed as follows:

[0173]

[0174] (7f) Minimize the loss function using gradient descent It can be expressed as follows:

[0175]

[0176]

[0177] In the formula, The sign of the partial derivative. Represents minimizing the error function right Find the partial derivative. The input to the estimation network is... , The squared difference between the calculated Q-value and the target value network's Q-value affects the estimated network's neural parameters. Find the partial derivative. To update the estimated network parameters after the update, the target value network parameters are updated every step. Replace with ;

[0178] When steps (7d) to (7f) are completed once, a single training session is completed. Steps (7b) to (7f) are repeated during each round of flight mission execution. After F rounds, the training is completed, and the learning process of the Double DQN algorithm ends.

[0179] (7g) The training algorithm ends, and the optimal path strategy is saved. Record reward values, flight steps, and data uploads: estimate network neural parameters over F training rounds. and target value network neural parameters The update strategy is performed in the direction of maximizing the total reward value R, eventually finding the optimal path strategy. .

[0180] This invention aims to improve the mission execution efficiency of a mission drone (hereinafter referred to as the mission drone) in a communication system that assists IoT devices in data uploading. It leverages the concept of far-field wireless charging and the flexibility and ease of deployment of drones by using a charging drone (hereinafter referred to as a mobile charger) equipped with a high-gain radio frequency antenna to provide charging services to the mission drone. For the aforementioned communication and wireless charging system, an improved deep reinforcement learning algorithm, Double DQN, is used to implement a flight path optimization strategy for the mission drone that balances data uploading tasks with power replenishment needs, thereby improving the mission execution efficiency of the mission drone.

[0181] Based on the characteristics of the path optimization problem, this invention addresses the state set S, action set A, and reward function. Furthermore, the design was carried out, and the actual field of view of the mission machine was simulated. Combined with effective system information, the input information of the neural network was designed. , , The algorithm uses a greedy coefficient. The algorithm incorporates a design that increases the number of effective samples acquired in the early stages by increasing the number of flight rounds. It also designs independent estimation and target value networks. During the learning process, both networks gradually update their parameters to reduce estimation errors, effectively mitigating the impact of overestimation and improving algorithm accuracy.

[0182] like Figure 2 As shown, the IoT communication and wireless charging system is denoted as the system. The system space is divided into an N*N grid, where each grid is a square cell with a side length of c, and N is the number of grids in the horizontal and vertical directions. Within the scenario, a drone is deployed to perform an upload task, a mobile charger provides wireless charging service, and M mobile IoT devices each have a value... The data volume is waiting to be uploaded. In this embodiment, N=15, c=10 meters, M=10. kB.

[0183] like Figure 4 As shown, Figure 4 The horizontal axis represents the number of mission rounds, and the vertical axis represents the reward value per flight round. From Figure 4 It can be seen that as the number of task rounds increases, the reward value gradually increases and tends to stabilize, with the training effect reaching its optimal level at around 14,000 rounds. The neural parameters of the estimation network and the target value network are shown below. , Update complete, resulting in a flight path strategy that maximizes mission execution efficiency. .

[0184] like Figure 5 As shown, Figure 5 The horizontal axis represents the number of mission rounds, and the vertical axis represents the number of flight steps of the mission aircraft in a single flight round. From Figure 5 It can be seen that as the number of mission rounds increases, the number of flight steps of the mission aircraft oscillates and gradually increases between rounds 0 and 10,000. Since the initial onboard battery power is insufficient to support the completion of the entire data upload mission, the mission aircraft learns to approach the charger at appropriate times to replenish its battery during this stage. Between rounds 10,000 and 16,000, the number of flight steps gradually decreases and eventually converges to 72 steps per flight. During this stage, the mission aircraft further learns from the generated optimized strategy set towards fewer flight steps and ultimately finds the path flight strategy with minimum energy consumption. .

[0185] like Figure 6 As shown, Figure 6 The horizontal axis represents the number of mission rounds, and the vertical axis represents the total data upload volume per flight round. From Figure 6 It can be seen that as the number of task rounds increases, the total amount of data uploaded gradually increases and stabilizes at a maximum value of 10,000kB when the number of rounds reaches around 12,000, thus realizing a strategy that can efficiently complete the data upload task.

[0186] In summary, compared with existing technologies, this invention utilizes far-field wireless charging technology and combines a highly mobile drone equipped with a high-gain radio frequency antenna as a mobile charger to provide more convenient charging services for mission drones that assist base stations in performing data uploads in IoT communication systems. Under conditions with high data real-time requirements, the more convenient charging method significantly improves the mission drone's endurance, achieving high data upload efficiency in IoT communication systems assisted by drones.

Claims

1. A method for optimizing the data upload path of a wireless rechargeable drone based on DRL, characterized in that: The method includes the following steps in sequence: (1) Construct an IoT communication and wireless charging scenario system, which includes a mission machine, a mobile charger and M mobile IoT devices; (2) Establish a first mathematical model for the consumption of the mission aircraft's onboard battery; (3) Establish a second mathematical model for the data upload channel; (4) Establish a mathematical model for energy replenishment during the wireless charging process, i.e., the third mathematical model; (5) Based on the IoT communication and wireless charging scenario system, the first mathematical model, the second mathematical model and the third mathematical model, a fourth mathematical model is established for the path optimization objective; (6) Based on the IoT communication and wireless charging scenario system, the first mathematical model, the second mathematical model, the third mathematical model, and the fourth mathematical model, determine the state set S, the action set A, and the reward function. ; (7) Based on the state set S, action set A, and reward function The optimal path strategy is obtained by using an improved Double DQN algorithm for offline learning. ; Step (2) specifically refers to: the maximum battery capacity of the mission machine is , This represents the remaining charge in time slot t. The mission aircraft consumes a constant amount of energy for each flight maneuver it performs. Without considering charging, a mathematical model is established for the battery level of the mission machine in time slot t+1, denoted as the first mathematical model, with the following expression: ; In the formula, This indicates the remaining battery power in time slot t+1; Step (3) specifically refers to: during the process of establishing a communication connection between the mission machine and the mobile IoT device and starting data upload, the mission machine flies at a sufficient altitude. At this time, the line-of-sight wireless transmission communication between the mobile IoT device and the mission machine is guaranteed. The expression for the channel gain between the two is as follows: ; in, This represents the channel gain at a reference distance of 1m. The Euclidean distance between the mission machine and the mobile IoT device m in time slot t is expressed as follows: ; in, This represents the Euclidean distance between the mission machine and the mobile IoT device m on the horizontal axis. This represents the Euclidean distance between the mission machine and the mobile IoT device m on the vertical axis; The mission aircraft is at a fixed flight altitude; a mathematical model for data transmission at time slot t is established, denoted as the second mathematical model: ; in, W is the transmission rate of the data transmission link established between the task machine in time slot t and the mobile IoT device m, where W is the signal bandwidth. It refers to the transmit power of IoT devices. It is noise power; Step (4) specifically refers to: the mobile charger is equipped with a high-gain radio frequency antenna for transmitting wireless power, providing energy supply services to the mission machine via a fixed deployment trajectory; during mission execution, the distance between the mission machine and the mobile charger varies in different time slots, and the power of wireless charging will drop sharply when the distance between them increases. Assuming full energy conversion efficiency in the wireless transmission link, the power obtained at the mission machine is calculated using the Friis free space propagation model. Expressed as follows: ; in, It is the transmitter power. and It refers to the antenna gain at the transmitting and receiving ends. It is the transmission wavelength. It is the Euclidean distance between the transmitter and receiver, expressed as follows: ; in, This represents the Euclidean distance between the mission unit and the mobile charger in the horizontal direction within time slot t. This represents the Euclidean distance between the mission unit and the mobile charger in the longitudinal direction along time slot t. The constant height difference between the mission device and the mobile charger is represented; a mathematical model for energy replenishment during the wireless charging process is established and denoted as the third mathematical model, the expression of which is as follows: ; In the formula, This represents the energy received by the mission machine in a single time slot. The length of a single time slot; Step (5) specifically refers to the following: In the dynamic scenario of optimizing the mission vehicle path planning problem, the variables are the location information of the mobile IoT device, the mobile charger, and the mission vehicle, the device data upload queue, and the mission vehicle battery status information; the optimization goal is to find a path strategy to help the mission vehicle make the best decision between balancing energy consumption and the amount of data uploaded, and to maximize the data upload efficiency during a single flight from the perspective of optimizing the movement trajectory. The main factors considered in this process are the data upload amount and mission energy consumption. The expression of the fourth mathematical model is as follows: ; in, It refers to the transmission rate of the data transmission link established between the task machine in time slot t and the mobile IoT device m. The mission aircraft consumes a constant amount of energy for each flight maneuver; let the current time slot of the system be t, t=0, ,2 …T, Where m is the length of a single time slot, T is the time of system termination; m∈{1,2…M}; Step (6) specifically includes the following steps: (6a) Determine the set of states S: The expression for the state set S is as follows: ; in, The system state at time slot t is determined by... , , It consists of three parts; The status of each mobile IoT device in the data upload task is represented to guide the data upload task objectives. This includes the location and remaining data volume information of all mobile IoT devices. Let M be the two-dimensional coordinates of the last mobile IoT device M in the region during the upload task in the horizontal direction on time slot t. , With the current remaining amount of data to be uploaded ,but The expression is as follows: ; By simulating the field of view of the task machine, the data transmission rate of mobile IoT devices at different locations within the effective transmission distance of the task machine is represented in relation to the location of the task machine. Based on the characteristics of the system, with the task machine as the center, it is assumed that the field of view range is n grids. The location outside the task area is represented by black grids, and the corresponding matrix value is set to -1 to indicate that it is outside the task area. The data upload rate at different locations is represented based on the difference in distance from the task machine. The upload rate at the task machine is the highest, and the corresponding matrix element is set to 50. The matrix value corresponding to the distance of one grid is set to 20, the matrix value corresponding to the distance of two grids is set to 5, and the matrix value corresponding to the remaining white grids is set to 0. Then a new vision matrix is ​​constructed. , Let be the value of the element in the i-th row and j-th column of the field of view matrix, where And assume that the spatial position of the i-th row and j-th column in the field of view matrix Y corresponds to the horizontal and vertical coordinates in the overall grid space as follows: , , Let the distance from the grid of the mission machine to the position of the element in the i-th row and j-th column of the field of view matrix at time slot t be the distance of the field of view matrix element. The numerical expression is as follows: ; The obtained field-of-view matrix Y contains A square matrix of elements, flattening the elements in the field of view matrix Y, and the horizontal coordinate of the joint mission machine in the grid space. The ordinate of the mission machine in the grid space Form a one-dimensional column vector The expression is as follows: ; Including the current battery level of the mission machine The x-coordinate of the mobile charger in the grid space The vertical coordinate of the mobile charger in the grid space And the Euclidean distance between the mission machine and the mobile charger. , The expression is as follows: ; (6b) Determine the action set A: The task machine in the system has four possible movement directions in the gridded space, and the motion space expression is as follows: ; in, The action performed by the mission aircraft in time slot t, and at a constant speed during flight. There are four flight directions; during flight, energy is obtained from the mobile charger transmitter via wireless charging, and the two processes of scene data uploading and energy acquisition are carried out in parallel. (6c) Determine the reward function : For the task machine, the impact of different behavioral decisions on various aspects of the system is reflected in the reward mechanism. The immediate reward obtained by the task machine in time slot t is represented as... The expression is as follows: ; in, , , To adjust , , Weighting factors between them; With the core objective of maximizing data throughput throughout the entire data upload process, a positive reward is assigned to each action based on the data throughput during a single communication session, as expressed in the following expression: ; in, This represents the amount of data uploaded by the mobile IoT device m to the task machine at time t; With the goal of maximizing the data upload efficiency of the entire data upload process, negative rewards are given to the task machine's movement behavior at each step to urge the task machine to improve its path selection ability, reduce unnecessary energy loss and promote the convergence of the optimal path. With the goal of maximizing the battery life of the task machine during data upload tasks, a positive behavioral reward is assigned to the wireless charging resulting from the task machine's movement decisions. The expression for this reward is as follows: ; in, This represents the remaining battery power of the mission machine at time slot t. This represents the energy received by the mission machine in a single time slot. The threshold for determining whether the mission machine has entered a low-energy state. and All are constant coefficients.

2. The method for optimizing the data upload path of a wireless rechargeable UAV based on DRL according to claim 1, characterized in that: Step (1) specifically refers to: The IoT communication and wireless charging scenario system is denoted as the system. The system space is divided into an N*N grid, where each grid is a square unit with a side length of c, and N is the number of grids in the horizontal and vertical directions. A task machine is deployed within the scenario to perform the upload task, a mobile charger provides wireless charging service, and M mobile IoT devices each have a value... The amount of data is waiting to be uploaded; Let the current time slot of the system be t, where t=0. ,2 …T, Let T be the length of a single time slot, and T be the moment when the system terminates. Then, the position of mobile IoT device m at time slot t is: It is represented as, where m∈{1,2…M}, This represents the x-coordinate of the mobile IoT device m in the grid space. This represents the ordinate of the mobile IoT device m in the grid space. This represents the fixed height of the mobile IoT device m; assuming the remaining amount of data to be uploaded by the mobile IoT device m in time slot t is represented as... ; The mobile charger, acting as an energy supplier, departs from its docking point at the start of the mission, following a pre-determined flight path and speed. Move and provide power to the mission machine, real-time location information. ; express, This represents the x-coordinate of the mobile charger in the grid space. This represents the ordinate of the mobile charger in the grid space. A fixed flight altitude for the mobile charger; at the start of the mission, the mission aircraft departs from the docking point and flies at a constant speed. Flight, real-time location It means that, among them, This represents the x-coordinate of the mission machine in the grid space. This represents the ordinate of the mission machine in the grid space. This refers to the fixed flight altitude of the mission aircraft.

3. The method for optimizing the data upload path of a wireless rechargeable UAV based on DRL according to claim 1, characterized in that: Step (7) specifically includes the following steps: (7a) Initialize the estimated network neural parameters and target value network neural parameters ,make Initialize the experience replay pool with a capacity of D; initialize the network learning rate. attenuation coefficient ; (7b) Based on the current state-action pair input , The system state at time slot t. For the action performed by the task machine in time slot t, the output is an estimated value of the Q value, i.e. ,in The network generates the estimated values ​​for the neural parameters; the target value network generates the action to select the next state. ,in The state values ​​are the input values ​​to the target value network. The action value is input to the target value network, and then the next state-action pair of the target value network is determined. Q value The objective value of the Double DQN algorithm is defined as: ; (7c) Based on the current state Execute action It then selects actions based on the improved Double DQN algorithm to transition to a new state. According to the reward function Calculate the single-step reward value and store the converted value in the experience replay pool. ; Repeat steps (7b) to (7c) until the number of memories stored in the experience replay pool equals D, then proceed to step (7d). The Double DQN algorithm has two neural network structures: an estimation network and a target value network. Given a training step size step, training is performed once every step. The total reward in each training round is defined as R, and the reward value obtained by executing a single flight path is expressed as follows: ; The improved Double DQN algorithm refers to the greedy coefficient for action selection. Improvements are made by ensuring the action that maximizes the reward value at each step based on the optimization metric. During the execution of the ε-greedy policy, the AI ​​perceives the action as numerically... The probability of choosing the current optimal solution is given, and the remaining 1- The probability of exploring other actions. The numerical changes of the coefficients are expressed as follows: ; in, Let K be the current flight round number. The maximum number of flight rounds reached at that time; The training dataset for offline learning is formed by randomly selecting z experiences from an experience replay pool of size D; the maximum number of task rounds is F. (7d) Randomly select z memories from the experience replay pool, z=32; denote the k-th state transition sequence in the mini-batch data as ; (7e) Obtained according to step (7d) Calculate the target Q value and loss value ,make The loss function is expressed as follows: ; (7f) Minimize the loss function using gradient descent It can be expressed as follows: ; ; In the formula, The sign of the partial derivative. Represents minimizing the error function right Find the partial derivative. The input to the estimation network is... , The squared difference between the calculated Q-value and the target value network's Q-value affects the estimated network's neural parameters. Find the partial derivative. To update the estimated network parameters after the update, the target value network parameters are updated every step. Replace with ; When steps (7d) to (7f) are completed once, a single training session is completed. Steps (7b) to (7f) are repeated during each round of flight mission execution. After F rounds, the training is completed, and the learning process of the Double DQN algorithm ends. (7g) The training algorithm ends, and the optimal path strategy is saved. Record reward values, flight steps, and data uploads: estimate network neural parameters over F training rounds. and target value network neural parameters The update strategy is performed in the direction of maximizing the total reward value R, eventually finding the optimal path strategy. .

Citation Information

Patent Citations

  • Unmanned aerial vehicle collection path planning method based on hierarchical deep reinforcement learning

    CN113190039A

  • Internet of Things information collection method based on unmanned aerial vehicle-wireless charging platform

    CN115696255A