A Multi-UAV Data Collection Method and System Based on Federated Reinforcement Learning
Through the multi-UAV data collection method based on federated reinforcement learning, the drone trajectory and communication resources are optimized, and the problems of high energy consumption and poor data security of sensor nodes are solved, achieving low-cost and efficient data transmission and security enhancement.
Patent Information
- Application Number
- CN202310156117.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-02-21
AI Technical Summary
In the prior art, sensor nodes have high energy consumption, difficulty in long-distance communication transmission, and poor data security, resulting in high system energy consumption and high data leakage risk.
The multi-UAV data collection method based on federal reinforcement learning is adopted to collect data through ground sensor equipment, dispatch drones to collect data, and optimize the drone trajectory and communication resource allocation to minimize system energy consumption and ensure data security.
It realizes low-cost, high-rate, and low-latency data transmission, reduces system energy consumption, and solves data leakage problems through distributed federated learning, improving data security and algorithm convergence performance.
Smart Images

Figure CN116205390B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of wireless communication, and more particularly to a multi-UAV data collection method and system based on federated reinforcement learning. Background Art
[0002] In recent years, with the rapid development of Internet of Things technology, a large number of sensor nodes have been deployed at some key nodes to monitor the environment. On the one hand, maintaining a very large-scale network is very challenging because sensor nodes consume a large amount of energy during operation, and frequent replacement or charging of sensor nodes will result in high costs. Currently, a better solution is to use backscatter communication technology, which enables sensor nodes to transmit data by reflecting radio frequency signals. Although its transmission power consumption is very low, it cannot perform long-distance transmission, which brings great difficulties to the data collection of the system. On the other hand, it is very difficult to ensure the security performance of data during the process of data collection or processing. It may be stolen by criminals, resulting in data leakage and other problems. Therefore, how to ensure long-distance and fast transmission of data and ensure the security performance of data is an urgent problem to be solved. Summary of the Invention
[0003] The present invention proposes a multi-UAV data collection method and system based on federated reinforcement learning. First, ground sensing devices collect the data they need to transmit, and then the ground center dispatches UAVs to collect the data collected by the sensors in the sensor area. When the UAVs collect the data sent by the ground sensors, they need to optimize their own trajectories and reasonably allocate communication resources, so as to save the energy consumption of the entire system.
[0004] To solve the above technical problems, the present invention adopts the following technical solutions:
[0005] A multi-UAV data collection method based on federated reinforcement learning, comprising the following steps:
[0006] Step S101: The ground sensor device collects the nearby data information, and the data information is collected by the ground sensor device sensing the surrounding information;
[0007] Step S102: The ground center dispatches UAVs to collect the data of the ground sensor device according to the number of available UAVs;
[0008] Step S103: While collecting all the data of the ground sensing devices, the UAVs optimize the UAV trajectories and allocate resources with the goal of minimizing the energy consumption of the entire system;
[0009] Step S104: Determine whether the maximum number of training times has been reached;
[0010] Step S105: Output the optimal trajectories of each drone and the resource allocation situation.
[0011] Further, in step S2, each drone senses the nearby sensor nodes and uses the greedy algorithm to select the sensor node with the maximum rate as the collection target. To ensure the integrity of the data, among them, the drone can only collect the data of the next sensor node after collecting the data of this sensor node.
[0012] Further, in step S3, the Federated Learning dueling DDQN algorithm based on the federated reinforcement learning method is used to optimize the energy consumption of the entire system.
[0013] Further, in step S3, the total energy consumed by the drone to complete the total data collection task can be expressed as:
[0014] E = E * + E com
[0015] Where E * is the energy consumption of the entire system, including the total on-board energy consumed by all drones in the total time slots N, and E com is the communication energy of the drone during the data collection process.
[0016] Further, the total on-board energy E * consumed by all drones in the total time slots N can be expressed as:
[0017]
[0018] Where P m,n (V) is the propulsion energy consumption to ensure the flight of drone m in time slot n, M is the number of drones, and the total data collection time slot of the drone is N.
[0019] Further, the propulsion energy consumption to ensure the flight of drone m in time slot n can be expressed as:
[0020]
[0021] Where P0 and P i are two constants, representing the blade profile power and induced power in the hovering state of the drone respectively, V is the flight speed of the drone, U tip is the tip speed of the rotor blade, v0 is the average rotor induced speed in the hovering state, d0 represents the fuselage drag ratio, s represents the rotor solidity, ρ represents the air density, and A represents the rotor disk area.
[0022] Furthermore, the communication energy consumption of the UAV during data collection can be expressed as:
[0023]
[0024] where p l,n is the transmission power of the l-th ground sensor node in time slot n, L is the number of ground sensor nodes, and the total transmission power of the ground sensing equipment is P, that is, 0 < p l,n ≤ P.
[0025] Furthermore, the trajectory of the UAV is optimized and the communication resources of the entire system are reasonably optimized to minimize the energy consumption E of the entire system. The specific process is as follows:
[0026] Step 1. Each UAV senses the nearby sensor nodes and uses the greedy algorithm to select the sensor node with the maximum rate as the collection target. Among them, the UAV can only collect the data of the next sensor node after completing the data collection of the current sensor node;
[0027] Step 2. The established problem is transformed into a Markov decision problem, and then the method of federated reinforcement learning is used to solve it. A complete Markov decision process can be composed of four parts, namely <S, A, γ, r n >, where S is the state space, A is the action space, γ is the state transition probability when the UAV executes the task, and r n is the reward function when the UAV executes the inspection task;
[0028] Step 3. Determine whether the UAV has completed all data collection tasks. If not, execute Step 1. If so, all data collection tasks end;
[0029] Step 4. Determine whether the maximum number of iterations has been reached. If not, repeat Steps 1-3 until the algorithm reaches the maximum number of iterations. If so, output the optimal trajectory and resource allocation results and end the program.
[0030] Furthermore, the process of optimizing the trajectory of the UAV and the communication resources of the system using the method of federated reinforcement learning is as follows:
[0031] a. Assume that the state of UAV m in time slot n is S n ={q m,n , α l,n}, and it takes an action A n ={o m,n , α l,n} from the action space A and transfers to the next state S n+1 ={q m,n+1 , α l,n+1}, and obtain a reward, and then save the result of the state transition (S n , A n , r n , S n+1 ) into the experience pool; where q m,n represents the position coordinates of the UAV m at time slot n, represents the power discrete matrix, o m,n = {n, s, e, w} represents the flight direction of the UAV, where n, s, e, w represent north, south, east, and west respectively, and S n+1 is the state of the UAV m at time slot n + 1;
[0032] b. Randomly select N1-step samples from the experience pool, and use the gradient descent method to reduce the loss function of the neural network to optimize the trajectory and communication resource allocation of the UAV, so as to obtain a greater reward, where the loss function is defined as
[0033]
[0034] where r n+1 represents the reward obtained by the UAV at time slot n + 1, λ represents the discount factor, θ * and θ represent the factors affecting the neural network model parameters, Q(S n , A n |θ) represents the Q value obtained by the UAV in the current state S n taking the action A n in the current network, represents the Q value of the UAV taking the action n+1 in the current state S in the target network;
[0035] c. The UAV sends the trained model parameters to the aggregation end, and then the aggregation end aggregates and processes the model parameters and sends them to each UAV. The aggregation end can be served by a UAV performing tasks, and the energy consumed in the process of model parameter exchange can be ignored. Assume that the model parameter of the UAV m at time slot n is w m,n , then the UAV aggregation end obtains the neural network model parameters w n+1 trained by all UAVs in the next time slot through aggregation and weighted processing, and transmits w n+1 to each UAV through downlink communication at time slot n + 1,
[0036] where w n+1 is specifically expressed as:
[0037]
[0038] where Let \(\upsilon\) be the number of model parameters of all UAVs, and \(\upsilon_m\) be the number of model parameters of UAV \(m\).
[0039] d. Determine whether the UAV has completed the data collection task of the sensing device. If not, the UAV executes step a. If so, the data collection of this sensor node is completed.
[0040] The present invention also provides a multi-UAV data collection method based on federated reinforcement learning, including a sensor data collection module, a UAV data collection module, a UAV trajectory optimization and resource allocation module, and a result output module. Among them,
[0041] The sensor data collection module senses and collects the data around it by ground sensor devices.
[0042] The UAV data collection module dispatches UAVs to collect the data of ground sensor devices.
[0043] The UAV trajectory optimization and resource allocation module optimizes the UAV trajectory and allocates resources with the goal of minimizing the energy consumption of the entire system.
[0044] The result output module outputs the optimal trajectory and resource allocation result of the optimized UAV inspection.
[0045] Compared with the prior art, the present invention has at least the following beneficial effects:
[0046] 1. Utilizing the flexibility of UAVs and the high data transmission rate of air-ground communication, by optimizing the UAV trajectory and the communication resources of the entire system, the energy consumption required for task completion of the entire system is effectively saved, featuring low cost, high rate, low latency, and energy saving.
[0047] 2. Using the method of distributed federated learning enables multiple UAVs to share model parameters during training when performing data collection tasks, not only solving the problem of data leakage during data collection by UAVs, but also accelerating the convergence performance of the algorithm compared with the centralized multi-UAV data collection solution.
[0048] 3. When multiple UAVs cooperate in data collection, the problems of safe flight between UAVs and flying out of the boundary are considered, making the multiple UAVs closer to the real data collection scenario when performing data collection tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 This is a schematic diagram of the system of the present invention;
[0051] Figure 2 This is the overall flowchart of the present invention;
[0052] Figure 3 This is the process of UAV trajectory optimization and resource allocation based on the federated reinforcement learning algorithm of the present invention;
[0053] Figure 4 This is the process of selecting data collection points by the greedy strategy of the present invention;
[0054] Figure 5 This is the training process of UAV trajectory optimization and resource allocation of the present invention. Detailed implementation manners
[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0056] It should be noted that the experimental methods described in the following implementation schemes are all conventional methods unless otherwise specified, and the reagents and materials can be obtained from commercial channels unless otherwise specified; in the description of the present invention, the terms "lateral", "longitudinal", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.
[0057] In addition, the terms "horizontal", "vertical", "hanging", etc. do not mean that the components are required to be absolutely horizontal or hanging, but can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined.
[0058] In the description of the present application, it should also be noted that unless otherwise clearly specified and limited, the terms "set", "installed", "connected", and "linked" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.
[0059] The present invention aims at the problems of difficult charging of sensor nodes, difficult long-distance communication transmission, and easy data leakage during data transmission, and proposes a method and system for collecting drone data based on federated reinforcement learning, which not only solves the problems of the inability of ground sensor devices to transmit over long distances and limited battery power of sensor devices, but also solves the problem of data leakage during data transmission. It has the advantages of energy saving, safety, and high-speed transmission.
[0060] As Figure 2 shown, this embodiment provides a method for collecting multi-drone data based on federated reinforcement learning, including the following steps:
[0061] Step S101: The ground sensor device collects the nearby data information, and the data information is collected by the ground sensor device sensing the surrounding information.
[0062] Step S102: The ground center dispatches drones to collect the data of the ground sensor devices according to the number of available drones.
[0063] Step S103: While collecting all the data of the ground sensing devices, the drones optimize the drone trajectories and allocate resources with the goal of minimizing the energy consumption of the entire system.
[0064] Step S104: Determine whether the maximum number of training times has been reached.
[0065] Step S105: Output the optimal trajectories of each drone and the resource allocation situation.
[0066] Assume that the system includes M drones and L ground sensor nodes, and these L ground sensor nodes are distributed in an area of K×KKm. Assume that the total data collection time slot for the drones is N; the position coordinates of the m-th drone at time slot n are q m,n =[x m,n , y m,n , z m,n , where x m,n is the abscissa of the m-th drone at time slot n, y m,n is the ordinate of the m-th drone at time slot n, and z m,nis the horizontal height of the drone m from the ground at time slot n; the position coordinates of the l-th ground sensor node are w l =[x l , y l , 0], the total transmission power of the ground sensing device is P, that is, 0 < p l,n ≤ P where p l,n is the transmission power of the l-th ground sensor node at time slot n.
[0067] In step S102, each drone senses the nearby sensor nodes and uses the greedy algorithm to select the sensor node with the maximum rate as the collection target. To ensure the integrity of the data, among them, the drone can only collect the data of the next sensor node after collecting the data of this sensor node.
[0068] In step S103, the Federated Learning dueling DDQN algorithm based on the federated reinforcement learning method is used to optimize the energy consumption of the entire system.
[0069] In step S103, the energy consumption of the entire system consists of two parts. One part is the propulsion energy consumption to ensure the flight of the drone, and the other part is the communication energy consumed by the drone during data collection. Among them, at time slot n, the propulsion energy consumption to ensure the flight of the drone m can be expressed as
[0070]
[0071] where P0 and P i are two constants, representing the blade profile power and induced power in the hovering state of the drone respectively, V is the flight speed of the drone, U tip is the tip speed of the rotor blade, and v0 is the average rotor induced speed in the hovering state. In addition, d0 represents the fuselage drag ratio, s represents the rotor solidity, ρ represents the air density, and A represents the rotor disk area. Then, the total onboard energy E * consumed by all drones in the total time slots N can be expressed as:
[0072]
[0073] The communication energy consumption of the drone during data collection can be expressed as
[0074]
[0075] Therefore, the total energy consumed by the drone to complete the total data collection task can be expressed as:
[0076] E = E * + E com
[0077] Optimize the trajectory of the UAV and the communication resources of the entire system reasonably to minimize the energy consumption E of the entire system. The specific process is as follows:
[0078] Step 1. Each UAV senses the nearby sensor nodes and selects the sensor node with the maximum rate as the collection target using the greedy algorithm. Among them, the UAV can only collect the data of the next sensor node after completing the data collection of this sensor node;
[0079] Step 2. Transform the established problem into a Markov decision problem, and then use the method of federated reinforcement learning to solve it. A complete Markov decision process can be composed of four parts, namely <S, A, γ, r n >, where S is the state space, A is the action space, γ is the state transition probability when the UAV executes the task, and r n is the reward function when the UAV executes the inspection task;
[0080] In this design, the state space is S = {q m,n , α l,n}, where q m,n represents the position coordinates of UAV m at time slot n, represents the power discrete matrix. The action space is A = {o m,n , α l,n}, where o m,n = {n, s, e, w} represents the flight direction of the UAV, and n, s, e, w represent north, south, east, and west respectively. The state transition probability γ represents the probability of being in state S n at time slot n, and the UAV transfers to the next state S n according to the action selection strategy and executes the selected action a n+1 . The reward function can be expressed as
[0081]
[0082] where a is a negative number, which represents the penalty for the UAV going out of bounds or collisions between UAVs during the execution of the task, and r m,n is the data collection rate of UAV m at time slot n, β is the weight coefficient, which is a constant. Among them, the data collection rate r m,n of UAV m at time slot n can be expressed as:
[0083]
[0084] where B is the bandwidth of the system, h m,n represents the channel gain of the system, l is the channel gain between UAV m and the ground sensing node l at time slot n during the transmission process, and p l$P_{ul}$ is the transmit power for the user's uplink communication, and $N_0$ is the noise power spectral density.
[0085] After establishing the Markov decision, the method of federated reinforcement learning is used to reasonably optimize the trajectory of the UAV and the communication resources of the system, so as to minimize the energy consumption of the entire system.
[0086] Step 3. Determine whether the UAV has completed all data collection tasks. If not, execute Step 1. If so, all data collection tasks end;
[0087] Step 4. Determine whether the maximum number of iterations has been reached;
[0088] If not, repeat Steps 1-3 until the algorithm reaches the maximum number of iterations. If so, output the optimal trajectory and resource allocation results and end the program.
[0089] The process of optimizing the trajectory of the UAV and the communication resources of the system using the method of federated reinforcement learning is as follows:
[0090] a. Assume that the state of UAV $m$ at time slot $n$ is $S$ n $=\{q$ m,n , $\alpha$ l,n $\}$, it takes an action $A$ n $=\{o$ m,n , $\alpha$ l,n $\}$ from the action space $A$, transfers to the next state $S$ n+1 $=\{q$ m,n+1 , $\alpha$ l,n+1 $\}$, and obtains a reward. Then, save the result of the state transition ($S$ n , $A$ n , $r$ n , $S$ n+1 ) to the experience pool
[0091] b. Randomly select $N_1$-step samples from the experience pool, and use the method of gradient descent to reduce the loss function of the neural network, thereby optimizing the trajectory of the UAV and the communication resource allocation, so as to obtain a greater reward. The loss function is defined as
[0092]
[0093] where $r$ n+1 represents the reward obtained by the UAV at time slot $n + 1$, $\lambda$ represents the discount factor, $\theta$ * and $\theta$ represent the factors affecting the neural network model parameters, and $Q(S$ n , $A$ n $|\theta)$ represents the $Q$-value obtained when the UAV takes action $A$ n in the current state $S$ n in the current network. Indicates the UAV in the current state S in the target network n+1 Take an action The Q value of;
[0094] c. The UAV sends the trained model parameters to the aggregation end, and then the aggregation end aggregates and processes the model parameters and sends them to each UAV. The aggregation end can be served by a UAV performing a task, and the energy consumed during the exchange of model parameters can be ignored. Assume that the model parameter of UAV m in time slot n is w m,n , then the UAV aggregation end obtains the neural network model parameters trained by all UAVs in the next time slot through aggregation and weighting as w n+1 , and transmits w n+1 to each UAV through downlink communication in time slot n+1,
[0095] where w n+1 Specifically expressed as:
[0096]
[0097] where is the size of the model parameter quantity of all UAVs, and υ represents the size of the model parameter quantity of UAV m;
[0098] d. Determine whether the UAV has completed the data collection task of this sensing device. If not, the UAV executes step a. If so, the data collection of this sensor node is completed.
[0099] As Figure 3 shown, the steps for the UAV to start collecting the data of the ground sensor node are:
[0100] Step 201: The process starts.
[0101] Step 202: Each UAV uses the greedy algorithm to determine its data collection sensor node.
[0102] Step 203: The UAV obtains the distances from other UAVs and the boundary at the current position.
[0103] Step 204: The UAV optimizes its own trajectory and resource allocation.
[0104] Step 205: Determine whether the UAV has completed the data collection task of this sensor node.
[0105] Step 206: Determine whether the UAV has completed the data collection tasks of all sensor nodes.
[0106] Step 207: The process ends.
[0107] Among them, in step 202, the UAV uses the greedy algorithm to determine its data collection sensor node asFigure 4 As shown, the specific steps are as follows:
[0108] Step 301: The process starts.
[0109] Step 302: Number the sensor nodes.
[0110] Step 303: Calculate the data rate between the UAV and each sensor node, delete the serial numbers of the completed sensor nodes, and sort the sensor nodes from largest to smallest according to the rate.
[0111] Step 304: Determine whether the data collection of the sensor node with the maximum rate is completed. If so, execute Step 303; otherwise, execute Step 305.
[0112] Step 305: Output the device number.
[0113] Step 306: The process ends.
[0114] As Figure 5 shown, the specific steps for the UAV to optimize its own trajectory and resource allocation in Step 204 are as follows:
[0115] Step 401: The UAV selects the current action according to the policy selection mechanism, and senses the distances between each UAV and the distance from the distance boundary in the next time slot.
[0116] Step 402: The UAV obtains the reward according to the state during the flight.
[0117] Step 403: Store the state transition information in the experience pool.
[0118] Step 404: Train the loss function.
[0119] Step 405: Each UAV sends the model parameters to the aggregation end.
[0120] Step 406: The aggregation end should send the processed model parameters to each UAV after aggregation processing.
[0121] This embodiment also provides a multi-UAV data collection method based on federated reinforcement learning, which is characterized in that it includes a sensor data collection module, a UAV data collection module, a UAV trajectory optimization and resource allocation module, and a result output module. Among them,
[0122] The sensor data collection module senses and collects the data around it by the ground sensor device;
[0123] The UAV data collection module dispatches UAVs to collect the data of the ground sensor device;
[0124] The UAV trajectory optimization and resource allocation module optimizes the UAV trajectory and allocates resources with the goal of minimizing the energy consumption of the entire system;
[0125] The result output module outputs the optimal trajectory of the optimized UAV inspection and the resource allocation result.
[0126] In summary, the present invention has the following technical effects;
[0127] 1. Considering the problems of limited transmission distance of the surface sensing device and limited energy of the total system, the mobility of the UAV and the high-probability LoS channel model are utilized to achieve long-distance data transmission and high-rate data transmission.
[0128] 2. Considering the problem of data leakage and other security issues that are likely to occur when the UAV collaborates with the ground sensing device to complete data collection, the method of federated learning is used to protect data privacy when multiple UAVs collaborate to collect data. Only the model parameters need to be uploaded, and there is no longer a need for a large amount of data transmission, reducing the communication overhead.
[0129] 3. Considering the safe flight problems of out-of-bounds and collision that are likely to occur when multiple UAVs perform data collection, it is more in line with the actual data collection scenario.
[0130] 4. According to the total resources of the system and the problem of limited energy, the trajectory of the UAV and the communication resource allocation are jointly optimized, so as to minimize the total energy consumption of the system on the premise of meeting all data collections.
[0131] The above embodiments are only illustrative examples of the technical solutions of the present invention. The methods and devices involved in the present invention are not limited only to the content described in the above embodiments, but are subject to the scope defined by the claims. Any modification, supplement, or equivalent replacement made by those skilled in the art to which the present invention pertains on the basis of this embodiment is within the scope protected by the claims of the present invention.
Claims
1. A multi - UAV data collection method based on federated reinforcement learning, characterized in that, It includes the following steps: Step S101: The ground sensor device collects nearby data information, where the data information is collected by the sensor device on the ground sensing the information around it; Step S102: The ground center dispatches drones to collect the data of the ground sensor device according to the number of available drones; Step S103: While collecting all the data of the ground sensing devices, the drones optimize the drone trajectories and resource allocation with the goal of minimizing the energy consumption of the entire system; Step S104: Judge whether the maximum number of training times is reached; Step S105: Output the optimal trajectories of each drone and the resource allocation situation; Reasonably optimize the trajectory of the drone and the communication resources of the entire system to minimize the energy consumption of the entire system , and the specific process is as follows: Step 1. Each drone senses the nearby sensor nodes and uses the greedy algorithm to select the sensor node with the maximum rate as the collection target. Among them, the drone can only collect the data of the next sensor node after collecting the data of this sensor node; Step 2. Convert the established problem into a Markov decision problem, and then use the method of federated reinforcement learning to solve it. A complete Markov decision process consists of four parts, namely , where is the state space, is the action space, is the state transition probability when the UAV performs tasks, is the reward function when the UAV performs inspection tasks; Step 3. Judge whether the drone has completed all data collection tasks. If not, execute Step 1. If so, all data collection tasks end; Step 4. Judge whether the maximum number of iterations is reached. If not, repeat Steps 1-3 until the algorithm reaches the maximum number of iterations. If so, output the optimal trajectory and resource allocation result and end the program; The process of optimizing the trajectories of drones and the communication resources of the system using the method of federated reinforcement learning is as follows: a. Hypothetical UAV At time slot The state is It takes an action from the action space and transfers to the next state while obtaining a reward. Then the result of the state transition is saved in the experience pool; where represents the position coordinates of the UAV at time slot , represents the power discrete matrix, represents the flight direction of the UAV represent north, south, east, and west respectively is the state of the UAV at time slot +1; b. Randomly select from the experience pool step samples, and use the gradient descent method to reduce the loss function of the neural network to optimize the trajectory and communication resource allocation of the UAV, so as to obtain greater rewards, where the loss function is defined as wherein represents the reward obtained by the UAV in the time slot, represents the discount factor, and represents a factor affecting the parameters of the neural network model, represents the value obtained when the UAV in the current network takes an action in the current state and the value represents the value obtained when the UAV in the target network takes an action in the current state and the value.
2. The multi-UAV data collection method based on federated reinforcement learning according to claim 1, wherein, In the said Step S102, each drone senses the nearby sensor nodes and uses the greedy algorithm to select the sensor node with the maximum rate as the collection target. To ensure the integrity of the data, among them, the drone can only collect the data of the next sensor node after collecting the data of this sensor node; 3. A multi-UAV data collection method based on federated reinforcement learning according to claim 1, characterized in that, In the said Step S103, the Federated Learning dueling DDQN algorithm based on the federated reinforcement learning method is used to optimize the energy consumption of the entire system; 4. A multi-UAV data collection method based on federated reinforcement learning according to claim 1, characterized in that In the said Step S103, the total energy consumed by the drone to complete the total data collection task can be expressed as: Among them, The energy consumption of the entire system includes the total on-board energy consumed by all UAVs in the total time slots under the condition, and the communication energy of the UAV during the data collection process.
5. A multi-UAV data collection method based on federated reinforcement learning according to claim 4, characterized in that, The total on-board energy consumed by all UAVs in the total time slot can be expressed as: can be expressed as: Among them, To ensure the propulsion energy consumption of the UAV during the time slot , ensure that the UAV The propulsion energy consumption during flight, is the number of UAVs, and the total data collection time slot completed by the UAVs is .
6. A multi-UAV data collection method based on federated reinforcement learning according to claim 5, characterized in that During a time slot , ensure that the propulsion energy consumption of the drone in flight can be expressed as: Among them, and are two constants, representing the blade profile power and induced power in the hovering state of the drone respectively, is the flight speed of the drone, is the tip speed of the rotor blade, is the average rotor induced velocity in the hovering state, represents the fuselage drag ratio, represents the rotor solidity, represents the air density, represents the rotor disk area.
7. A multi-UAV data collection method based on federated reinforcement learning according to claim 4, characterized in that The energy consumption of the drone during data collection can be expressed as: Among them, is the transmission power of the th ground sensor node in time slot . is the number of ground sensor nodes, and the total transmission power of the ground sensing device is , that is .
8. A multi-UAV data collection method based on federated reinforcement learning according to claim 1, characterized in that, The process of optimizing the trajectories of drones and the communication resources of the system using the method of federated reinforcement learning also includes: c. The UAV sends the trained model parameters to the aggregation end, and then the aggregation end aggregates and processes the model parameters and sends them to each UAV. The aggregation end can be served by a UAV performing tasks, and the energy consumed during the process of model parameter exchange can be ignored. Assume the UAV at time slot has model parameters of , then the neural network model parameters for all UAVs to train in the next time slot obtained by the UAV aggregation end through aggregation and weighting processing are , and at time slot , it transmits to each UAV through downlink communication Among them Specifically expressed as: wherein is the number of model parameters of all drones, indicating the drone in terms of the number of model parameters; d. Judge whether the drone has completed the data collection task of this sensing device. If not, the drone executes Step a. If so, the data collection of this sensor node is completed; 9. A multi-UAV data collection method based on federated reinforcement learning, characterized in that, It includes a sensor data collection module, a drone data collection module, a drone trajectory optimization and resource allocation module, and a result output module. Among them, The sensor data collection module senses and collects the data around it by the ground sensor device; The drone data collection module dispatches drones to collect the data of the ground sensor device; The drone trajectory optimization and resource allocation module optimizes the drone trajectories and resource allocation with the goal of minimizing the energy consumption of the entire system; The result output module outputs the optimal trajectory and resource allocation result of the optimized drone inspection; Reasonably optimize the trajectory of the drone and the communication resources of the entire system to minimize the energy consumption of the entire system , and the specific process is as follows: Step 1. Each drone senses the nearby sensor nodes and selects the sensor node with the maximum rate as the collection target using the greedy algorithm. Among them, the drone can only collect the data of the next sensor node after completing the data collection of this sensor node; Step 2. Convert the established problem into a Markov decision problem, and then use the method of federated reinforcement learning to solve it. A complete Markov decision process consists of four parts, namely , where is the state space, is the action space, is the state transition probability when the UAV performs tasks, is the reward function when the UAV performs the inspection task; Step 3. Determine whether the drone has completed all data collection tasks. If not, execute Step 1. If so, all data collection tasks end; Step 4. Determine whether the maximum number of iterations has been reached. If not, repeat Steps 1-3 until the algorithm reaches the maximum number of iterations. If so, output the optimal trajectory and resource allocation results and end the program. The process of optimizing the trajectory of the drone and the communication resources of the system using the method of federated reinforcement learning is as follows: a. Assume the drone at time slot is in the state of , it takes an action from the action space , transfers to the next state , and obtains a reward, then saves the result of the state transition into the experience pool; where represents the position coordinates of the drone at time slot , represents the power discrete matrix represents the flight direction of the drone respectively represent north, south, east, and west is the state of the drone at time slot +1; b. Randomly select from the experience pool step samples, and use the gradient descent method to reduce the loss function of the neural network, thereby optimizing the trajectory of the UAV and the allocation of communication resources, so as to obtain greater rewards, where the loss function is defined as Among them indicates the reward obtained by the UAV in the time slot, represents the discount factor, and represents the factor affecting the parameters of the neural network model, represents the value obtained by the UAV in the current state taking action in the current network, value, represents the value of the UAV taking action in the current state in the target network, value.
Citation Information
Patent Citations
Unmanned aerial vehicle three-dimensional trajectory planning method and system based on deep reinforcement learning
CN115580845A
Method and device for monitoring vehicle's brake system in autonomous driving system
US20210331655A1