A multi-drone collaborative multi-hop real-time relay link design method and device
Through the multi-UAV collaborative multi-hop real-time relay link design method, using the Dec-POMDP model and HyAR-TD3 algorithm optimization strategy, the problem of limited data transmission and energy supply range of UAV clusters is solved, and low-latency and efficient data transmission and energy supply are achieved.
Patent Information
- Application Number
- CN202410953724.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-16
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-07-16
AI Technical Summary
The limited perception range of drones in multi-drone scenarios leads to limited data transmission and energy supply range, and the timeliness of data packets is difficult to measure. Existing technologies cannot effectively solve the data transmission and energy supply problems between drone clusters.
A multi-UAV cooperative multi-hop real-time relay link design method is adopted. By acquiring environmental information, generating basic area units, setting update degree information, and establishing a Dec-POMDP model, the HyAR-TD3 algorithm is used to optimize the strategy model and deploy it to each UAV to achieve data transmission and energy supply.
It solves the limitations of the drone cluster's perception range and data transmission range, reduces data packet delay and drone energy consumption, and is suitable for multi-drone collaborative systems with different perception ranges and numbers of users.
Smart Images

Figure CN118826837B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of multi-UAV collaboration, and in particular to a multi-UAV collaboration multi-hop real-time relay link design method and a multi-UAV collaboration multi-hop real-time relay link design device. Background Art
[0002] With the rapid development of wireless communication technology, IoT devices are becoming increasingly prevalent in modern society. This influx of IoT devices has significantly impacted existing communication networks. Frequent data exchange not only increases network traffic load but also accelerates the energy consumption of IoT devices. Traditional energy replenishment methods, such as battery replacement and the use of renewable energy, have their own drawbacks. Wireless data-energy integrated transmission technology, as a new energy transmission method, can effectively solve the energy supply problem of IoT devices. Furthermore, the large number of IoT devices has significantly promoted the development and application of real-time status update systems in various fields such as intelligent transportation, environmental monitoring, and smart agriculture. In these systems, data generated by sensor nodes often needs to be collected and processed promptly to meet real-time user needs. However, current considerations of data real-time performance only consider the information age at the sensor node level. When sensor nodes continuously sense large amounts of data, information age alone is insufficient to characterize the real-time nature of the collected data. Therefore, a more practical metric is needed to characterize the timeliness of each data packet in the sensor node.
[0003] Drone-assisted data acquisition systems are currently attracting significant attention in the industrial sector. Drones, thanks to their flexibility and controllability, can establish line-of-sight communication links with target communication nodes at a very low cost. However, due to limitations in their sensing range, in real-world production, drones can only acquire sensor information within their range. In multi-drone scenarios, the transmission of control information and data between drones is also limited. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-UAV collaborative multi-hop real-time relay link design method to solve at least one of the above-mentioned technical problems.
[0005] One aspect of the present invention provides a method for designing a multi-drone collaborative multi-hop real-time relay link, the method comprising:
[0006] Step 1: Obtain basic information about the environment to be transferred;
[0007] Step 2: Generate multiple basic area units based on the basic information of the environment to be transmitted;
[0008] Step 3: Set the update degree information for each basic area unit;
[0009] Step 4: Establishing a Dec-POMDP model framework based on the basic information of the environment to be transmitted and the update degree information of each basic area unit;
[0010] Step 5: Use the HyAR-TD3 algorithm to train and optimize the parameters of the Dec-POMDP model framework to obtain the trained policy model;
[0011] Step 6: Deploy the trained policy model to each drone.
[0012] Optionally, the basic environment information to be transmitted includes the number of drones, drone performance information, the number of sensors, sensor performance information and the location of each sensor.
[0013] Optionally, the step 2: generating a plurality of basic area units according to the basic information of the environment to be transmitted includes:
[0014] Step 21: Acquire a sensor node set, where the sensor node set includes a plurality of sensor nodes;
[0015] Step 22: Calculate the sample density of all sensor nodes based on the sensor node set;
[0016] Step 23: Obtain one of the sensor nodes as the first centroid according to the sample density of all sensor nodes;
[0017] Step 24: Eliminate other sensor nodes that meet the first condition from the sensor node set according to the first centroid, thereby obtaining a sensor node set after elimination;
[0018] Step 25: Determine whether there are still sensor nodes in the sensor node set after elimination. If so,
[0019] Repeat steps 22 to 24 until no other sensor nodes exist in the sensor node set.
[0020] Optionally, the following formula is used to obtain the update degree information of each basic area unit:
[0021] in,
[0022] τ c (t) represents the update degree of the cth basic area unit in the tth time slot, H 4 Indicates that in the current time slot, a drone provides data services to the PoI in the cth area. Indicates that no drone provides data services for the c-th PoI in the current time slot.
[0023] Optionally, the hybrid action space model includes:
[0024] State space S:
[0025] The state of the system environment at the tth time slot is expressed as:
[0026] s(t)={{q n (t)},{α m (t)},{e k (t)},{L},{e n (t)}}; where,
[0027] s(t) represents the state of the system environment at the tth time slot, q n (t) is the position of all UAVs in the tth time slot, α m (t) is the data freshness of all data packets in the tth time slot, e k (t) is the capacity of all sensor nodes in the tth time slot, L is the capacity of all relay links in the tth time slot, e n (t) is the energy consumption of all UAVs in the tth time slot;
[0028] Observation space O:
[0029] The observation of drone n in the tth time slot is expressed as:
[0030]
[0031] in,
[0032] o n (t) is the observation of UAV n in the tth time slot, q n (t) is the position of all UAVs at the tth time slot, represents the number of data packets collected at the data center in the tth time slot, represents the sum of the data freshness of the data packets collected at the data center in the t-th time slot, Indicates that the drone n in the tth time slot can access the data center directly or through a relay link, represents the forward drone n at the tth time slot f location, represents the number of data packets within the sensing range of drone n in the tth time slot, represents the backward drone n in the tth time slot b If there are multiple backward UAVs, their average position is taken. represents the total number of data packets within the sensing range of all UAVs in the backward link of the tth time slot, {τ c (t)} represents the update degree of all basic area units in the t-th time slot;
[0033] Action space A:
[0034] The direction of drone n in the tth time slot can be expressed as:
[0035] in,
[0036] Indicates the direction of drone n, N f To move the drone in the direction The direction of discretization, represents the direction of drone n in the t-th time slot, π represents pi;
[0037] The action of drone n in the tth time slot can be expressed as:
[0038] in,
[0039] a n (t) represents the action of drone n in the tth time slot, represents the direction of drone n in the tth time slot, θ n (t) indicates the t Time slot drones n The time division factor of
[0040] Reward function R:
[0041] The return of drone n in the tth time slot is expressed as:
[0042] in,
[0043] Indicates the reward shaping part, Represents the objective function part.
[0044] Optionally, the step 5 of respectively training and optimizing the parameters of each Dec-POMDP model using the HyAR-TD3 algorithm to obtain a trained Dec-POMDP model includes:
[0045] The following training optimization is performed for the parameters of each Dec-POMDP model:
[0046] A1, randomly initialize the parameters ω, ε1, ε2 of the Actor network and the two Critic networks;
[0047] A2. Randomly initialize the embedding table, encoder, and decoder parameters ζ, χ e , χ d ;
[0048] A3. Prepare an experience replay pool D;
[0049] A4. Repeat the warm-up phase;
[0050] A5, repeated training stage;
[0051] A6. End.
[0052] Optionally, the A4, repeated preheating stage includes:
[0053] A41, using random strategies to collect experience and store it in experience replay pool D;
[0054] A42, train CVAE every certain number of steps;
[0055] A43: Until the maximum number of preheating steps is reached, otherwise return to A4.
[0056] Optionally, the A5, repeated training phase includes:
[0057] A51, repeat each time slot:
[0058] A511, get the environment state s;
[0059] A512. Take the state s as the input of Actor and get the implicit action z k ,
[0060] A513. According to the similarity principle, z k Input into the embedding table to obtain discrete action k, z k and s are input to the decoder to obtain the continuous action x k ;
[0061] A512, mix the action k, x k Input into the environment and get the next state s - , immediate return r and end mark Done;
[0062] A513, will Deposit into experience replay pool D;
[0063] A512, after a certain number of steps, train the TD3 network;
[0064] A513: Until the maximum number of steps in a training round is reached, otherwise return to A51;
[0065] A52. After a certain number of rounds, train CVAE.
[0066] A53. Until the maximum number of training rounds is reached, otherwise return to A5.
[0067] Optionally, the optimization goal of the strategy model is:
[0068]
[0069] The present application also provides a multi-UAV cooperative multi-hop real-time relay link design device, which includes:
[0070] A basic information acquisition module for the environment to be transmitted, wherein the basic information acquisition module is used to acquire basic information of the environment to be transmitted;
[0071] A basic area unit division module, wherein the basic area unit division module is used to generate a plurality of basic area units according to basic information of the environment to be transmitted;
[0072] An update degree setting module, configured to set update degree information for each basic area unit;
[0073] A hybrid action space model establishment module, the hybrid action space model establishment module is used to establish a Dec-POMDP model according to the basic information of the environment to be transmitted and the update degree information of each basic area unit;
[0074] A training optimization module, which is used to train and optimize the parameters of the Dec-POMDP model framework through the HyAR-TD3 algorithm to obtain a trained policy model;
[0075] A deployment module is used to deploy the trained policy model to each drone.
[0076] Beneficial effects
[0077] The multi-drone collaborative multi-hop real-time relay link design method of this application can solve the problems of limited data service range caused by the limited perception range of drones, limited range of control information transmission between drone clusters, and low latency requirements for data packets. The method of the present invention has the following advantages:
[0078] 1. The perception range of drone clusters is limited, and the range of data transmission and control information transmission between clusters is limited, which is more in line with actual production and life.
[0079] 2. Using drone clusters to form a multi-hop real-time relay link in the air can transmit data packets directly from the sensor back to the data center, greatly reducing the delay waste caused by drones carrying data packets.
[0080] 3. By jointly designing the trajectory planning of the drone cluster and the time segmentation scheme for each time slot, the average data freshness of the data packets is further reduced and the average energy consumption of the drones is reduced.
[0081] 4. The proposed algorithm is applicable to multi-UAV cooperative multi-hop real-time relay link systems with different sensing ranges, different numbers of users, and different flight altitudes. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 This is a flow chart of a method for designing a multi-UAV collaborative multi-hop real-time relay link according to an embodiment of the present application.
[0083] Figure 2 It is used to implement Figure 1 The electronic device diagram of the multi-UAV cooperative multi-hop real-time relay link design method is shown.
[0084] Figure 3 This is a system model diagram of an embodiment of the present application;
[0085] Figure 4 This is a schematic diagram of a chain relay link according to an embodiment of the present application;
[0086] Figure 5 This is a schematic diagram of a tree-like relay link according to an embodiment of the present application;
[0087] Figure 6 This is a flowchart of an algorithm according to an embodiment of the present application;
[0088] Figure 7 This is a schematic diagram of the CVAE structure of an embodiment of the present application;
[0089] Figure 8 This is a flowchart of an embodiment of the present application;
[0090] Figure 9 This is a schematic diagram of the TD3 structure of an embodiment of the present application;
[0091] Figure 10 A flight trajectory diagram of a drone cluster according to an embodiment of the present application;
[0092] Figure 11 This is a graph showing how the time division factor changes over time according to an embodiment of the present application;
[0093] Figure 12 This is a graph showing how average data freshness changes with the number of users in one embodiment of the present application;
[0094] Figure 13 This is a graph showing how average data freshness changes with increasing perception range in one embodiment of the present application;
[0095] Figure 14 This is a graph showing how average data freshness changes with increasing flight altitude according to an embodiment of the present application;
[0096] Figure 15 This is a trend chart showing how the average energy consumption of a drone changes with different parameters according to an embodiment of the present application. DETAILED DESCRIPTION
[0097] In order to make the purpose, technical solutions and advantages of the implementation of this application clearer, the technical solutions in the embodiments of this application will be described in more detail below in conjunction with the drawings in the embodiments of this application. In the drawings, the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The described embodiments are part of the embodiments of this application, not all of the embodiments. The embodiments described below with reference to the drawings are exemplary and are intended to be used to explain this application, and should not be understood as limitations on this application. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. The embodiments of this application are described in detail below in conjunction with the drawings.
[0098] Figure 1 This is a flow chart of a method for designing a multi-UAV collaborative multi-hop real-time relay link according to an embodiment of the present application.
[0099] like Figure 1 The multi-UAV cooperative multi-hop real-time relay link design method shown includes:
[0100] Step 1: Obtain basic information about the environment to be transferred;
[0101] Step 2: Generate multiple basic area units based on the basic information of the environment to be transmitted;
[0102] Step 3: Set the update degree information for each basic area unit;
[0103] Step 4: Establishing a Dec-POMDP model framework based on the basic information of the environment to be transmitted and the update degree information of each basic area unit;
[0104] Step 5: Use the HyAR-TD3 algorithm to train and optimize the parameters of the Dec-POMDP model framework to obtain the trained policy model;
[0105] Step 6: Deploy the trained policy model to each drone.
[0106] In this embodiment, the basic environment information to be transmitted includes the number of drones, drone performance information, the number of sensors, sensor performance information, and the location of each sensor.
[0107] In this embodiment, the step 2: generating a plurality of basic area units according to the basic information of the environment to be transmitted includes:
[0108] Step 21: Acquire a sensor node set, where the sensor node set includes a plurality of sensor nodes;
[0109] Step 22: Calculate the sample density of all sensor nodes based on the sensor node set;
[0110] Step 23: Obtain one of the sensor nodes as the first centroid according to the sample density of all sensor nodes;
[0111] Step 24: Eliminate other sensor nodes that meet the first condition from the sensor node set according to the first centroid, thereby obtaining a sensor node set after elimination;
[0112] Step 25: Determine whether there are still sensor nodes in the sensor node set after elimination. If so,
[0113] Repeat steps 22 to 24 until no other sensor nodes exist in the sensor node set.
[0114] In this embodiment, the following formula is used to obtain the update degree information of each basic area unit:
[0115] in,
[0116] τ c (t) represents the update degree of the cth basic area unit in the tth time slot, H 4 Indicates that in the current time slot, a drone provides data services to the PoI in the cth area. Indicates that no drone provides data services for the c-th PoI in the current time slot.
[0117] In this embodiment, the hybrid action space model includes:
[0118] State space S:
[0119] The state of the system environment at the tth time slot is expressed as:
[0120] s(t)={{q n (t)},{α m (t)},{e k (t)},{L},{e n (t)}}; where,
[0121] s(t) represents the state of the system environment at the tth time slot, q n (t) is the position of all UAVs in the tth time slot, α m (t) is the data freshness of all data packets in the tth time slot, e k (t) is the capacity of all sensor nodes in the tth time slot, L is the capacity of all relay links in the tth time slot, e n (t) is the energy consumption of all UAVs in the tth time slot;
[0122] Observation space O:
[0123] The observation of drone n in the tth time slot is expressed as:
[0124]
[0125] in,
[0126] o n (t) is the observation of UAV n in the tth time slot, q n (t) is the position of all UAVs at the tth time slot, represents the number of data packets collected at the data center in the tth time slot, represents the sum of the data freshness of the data packets collected at the data center in the t-th time slot, Indicates that the drone n in the tth time slot can access the data center directly or through a relay link, represents the forward drone n at the tth time slot f location, represents the number of data packets within the sensing range of drone n in the tth time slot, represents the backward drone n in the tth time slot b If there are multiple backward UAVs, their average position is taken. represents the total number of data packets within the sensing range of all UAVs in the backward link of the tth time slot, {τ c (t)} represents the update degree of all basic area units in the t-th time slot;
[0127] Action space A:
[0128] The direction of drone n in the tth time slot can be expressed as:
[0129]
[0130] The action of drone n in the tth time slot can be expressed as
[0131]
[0132] Reward function R:
[0133] The return of drone n in the tth time slot is expressed as
[0134] in,
[0135] Indicates the reward shaping part, Represents the objective function part.
[0136] In this embodiment, step 5: respectively training and optimizing the parameters of each Dec-POMDP model using the HyAR-TD3 algorithm to obtain the trained Dec-POMDP model includes:
[0137] The parameters of each Dec-POMDP model are trained and optimized as follows:
[0138] A1, randomly initialize the parameters ω, ε1, ε2 of the Actor network and the two Critic networks;
[0139] A2. Randomly initialize the embedding table, encoder, and decoder parameters ζ, χ e , χ d ;
[0140] A3. Prepare an experience replay pool D;
[0141] A4. Repeat the warm-up phase;
[0142] A5, repeated training stage;
[0143] A6. End.
[0144] In this embodiment, the A4, repeated preheating stage includes:
[0145] A41, using random strategies to collect experience and store it in experience replay pool D;
[0146] A42, train CVAE every certain number of steps;
[0147] A43: Until the maximum number of preheating steps is reached, otherwise return to A4.
[0148] In this embodiment, the A5, repeated training phase includes:
[0149] A51, repeat each time slot:
[0150] A511, get the environment state s;
[0151] A512. Take the state s as the input of Actor and get the implicit action z k ,
[0152] A513. According to the similarity principle, z k Input into the embedding table to obtain discrete action k, z k and s are input to the decoder to obtain the continuous action x k ;
[0153] A512, mix the action k, x k Input into the environment and get the next state s- , immediate return r and end mark Done;
[0154] A513, will Deposit into experience replay pool D;
[0155] A512, after a certain number of steps, train the TD3 network;
[0156] A513: Until the maximum number of steps in a training round is reached, otherwise return to A51;
[0157] A52. After a certain number of rounds, train CVAE;
[0158] A53. Until the maximum number of training rounds is reached, otherwise return to A5.
[0159] In this embodiment, the application exploits the limited sensing range of drones and the limited transmission range of control information between clusters to develop appropriate strategies and construct a multi-hop relay link with a suitable structure to better provide data and energy services to sensor nodes. By jointly designing the trajectory planning of the drone cluster and the time division factor of each time slot, the application aims to minimize the average data freshness and the average energy consumption of the drones. Data simulation is used to verify the effectiveness of the proposed algorithm.
[0160] In this embodiment, the tasks that each drone of this application needs to perform are:
[0161] The drone cluster departs from the data center and flies to the target area, collecting data from all sensor nodes in the target area while transmitting energy downlink, and returns to the data center after completing the mission.
[0162] In this embodiment, the tasks to be performed by each sensor node are as follows:
[0163] Ground sensor nodes k∈K=(1,2,…,K) continuously sense and monitor their surroundings and generate data packets with a certain timeliness. The timeliness of each data packet is represented by data freshness α. Specifically, data freshness α increases over time until the packet is retrieved to the data center. Obviously, the higher the data freshness α, the lower its timeliness.
[0164] In addition, the process of sensor node k sensing data is defined as a composite Poisson process in N represents the number of data packets sensed by sensor node k s The parameter is λ s Poisson distribution, X k,i ∈N(μ s ,(σ s ) 2) represents the amount of data X contained in the i-th data packet perceived by sensor node k k,i Obey Gaussian distribution. Each time a sensor node generates a data packet, it consumes a certain amount of energy e s ,When the energy reserve of the sensor node k is exhausted, the ,device will stop working and no longer sense the monitoring ,environment.
[0165] In this embodiment, during a mission, the data transmission relationship between the drone swarm and the sensor is as follows:
[0166] The task of the drone cluster is divided into three stages: 1) (drone cluster) UAVs depart from the data center and head to their respective target areas; 2) UAVs interact with SNs according to certain strategies to complete the data transmission task; 3) UAVs return to the data center. For convenience, during a mission, the data packets generated in the third stage are not used as the target of this round of missions. UAVs only need to collect the data packets generated in the first two stages of this round of missions and the third stage of the previous round of missions. During the entire mission, time is divided into T equal time slots t0. Within the same time slot, the channel state is considered unchanged. For any data packet, the data freshness update process in the t∈T=(1,2,…,T)th time slot can be expressed as
[0167]
[0168] Among them H 1 ,H 2 ,H 3 They represent the beginning of the data packet, the data packet is in the data queue of the sensor node, and the data packet is in the data center.
[0169] In this embodiment, the positions of the drone and the sensor are represented by a three-dimensional Cartesian coordinate system. Specifically, the position of the drone n∈N=(1,2,…,N) in the tth time slot can be expressed as q n =(x n (t),y n (t),H), the position of sensor k can be expressed as q k =(x k (t),y k (t),0), where the altitude of the drone is fixed as H. In addition, the location of the data center is denoted by q o =(x o ,y o ,0) indicates.
[0170] In this embodiment, the sensing range R of any UAV n is sare all limited. Sensor k can only transmit data and energy when it is within the sensing range of drone n. Drone n transmits energy to sensor node k, and sensor node k transmits data to drone n. It should be noted that in order to make the data uploaded by sensor node k more timely, drone n does not carry the collected data itself, but directly transmits it back to the data center through multi-hop forwarding to reduce the timeliness waste caused by the drone carrying the flight. Each relay node is a drone. The condition for any two drones n and n′ to form a relay relationship is ||q n (t)-q n′ (t)||<2R s , any drone n can directly transmit data back to the data center under the condition that ||q n (t)-q o ||<R s If the range limit is exceeded, not only will the data relay behavior fail to be completed, but the status of each drone will also not be shared.
[0171] In this embodiment, each relay link is represented by L={o,n1,n2,…,n i} indicates that the fewer hops a drone n has to take to reach data center o in the relay link, the closer it is to the front of the communication relay link. It is known that in order to transmit data in real time, the communication link structures that a drone cluster may have include chain and tree relay links. Obviously, when the number of drones in a chain or tree relay link is constant, the former can provide data services to sensor nodes in more distant areas, while the latter can simultaneously serve sensor nodes in a wider area. Let ζ n (t) represents the communication enable signal of drone n in the tth time slot. n When (t) = 1, it means that in the tth time slot, UAV n can establish a communication link with the data center directly or through a multi-hop relay link. Otherwise, n (t)=0.
[0172] In this embodiment, there is a line-of-sight link between the drone cluster and between the drone and the sensor node. The channel gain between drone n and sensor node k in the tth time slot can be expressed as:
[0173]
[0174] Where g0 represents the channel gain when the reference distance is 1m. Let θ n (t) represents the time division factor of the t-th time slot UAV n. In any time slot t, the first θ n (t)t0 hovers and is used for digital energy transmission, after (1-θ n(t))t0 is used for flight, and the flight speed is constant at v max Let κ n (t)∈K represents the sensor set that still has data packets within the sensing range of drone n in the tth time slot, then any sensor node k∈κ n (t) The amount of data transmitted uplink can be expressed as:
[0175]
[0176] where p k represents the transmit power of the sensor node, σ 2 represents the receiving noise of the UAV, M FM The total bandwidth is the total bandwidth. The data transmission between the UAV and the sensor nodes within its sensing range adopts the FDMA mechanism. The UAV transmits energy downlink in the form of broadcast. The energy consumed by the sensor node k to upload data can be expressed as The energy received from the downlink transmission of UAV n is Then the energy of sensor node k in the tth time slot can be expressed as:
[0177]
[0178] The energy consumed by the downlink transmission of UAV n can be expressed as The energy consumption required for hovering and flying are: and where p u (v) represents the power consumption of the drone when hovering or flying:
[0179]
[0180] Where P0 and P1 represent the blade shape power and derived power of the drone when it is hovering, v represents the speed of the drone, s tip represents the tip speed of the blade, s0 is the average rotor induced speed during hovering, d0 represents the airframe drag ratio, ρ represents the air density, μ0 represents the rotor solidity, and Z represents the area of the rotor disk. In summary, the energy consumption of UAV n in the tth time slot can be expressed as:
[0181]
[0182] In this embodiment, during the entire mission, the drone cluster not only needs to control the time division factor θ of each time slot, but also needs to determine the flight direction of each time slot. Then the position of drone n at any time slot t can be expressed as:
[0183]
[0184] In this embodiment, let the number of UAVs at the front and the back in a certain communication relay link be n and n respectively. b , n e , if in the tth time slot Then a drone will appear b The data transmission process has ended, but the drone n in the subsequent relay link e There are still In order to deal with the conflict, this application decides to adopt a free time transmission mechanism. Specifically, the free transmission mechanism means that any UAV in a relay link L will manage the behavior of the forward relay and the data from the backward relay through a pre-set link layer protocol: for the behavior of the forward relay n′, UAV n will divide the data into two groups according to the time division factor θ decided by itself. n (t) is used to allocate the transmission and flight time of the data; for the data of the backward relay n″, the drone n will receive and store the data packets in the data buffer queue according to the arrival time of the data packets, and will continue to forward the data to the forward relay or data center during the remaining flight time of this time slot. Here, the present invention assumes that the Doppler effect caused by the movement of drone n can be well compensated.
[0185] In this embodiment, if a certain time slot sensor node k is simultaneously within the sensing range of multiple drones, let κ k (t)∈N represents the set of these drones. For the public data transmission phase of all drones in the set, the sensor node k selects the drone that uploads And it will receive all drones n∈κ k (t) The energy of downlink transmission, i.e.
[0186] In this embodiment, Figure 3 This is the system model diagram of the present invention. Figure 2 As shown, the algorithm of the present invention is implemented based on the following system: in a multi-UAV cooperative multi-hop relay real-time communication system including N unmanned aerial vehicles (UAVs) and K ground sensor nodes (SNs), the UAV cluster starts from the data center, flies to the target area, collects data from all sensor nodes in the target area, and transmits energy downlink at the same time, and returns to the data center after completing the mission.
[0187] Figure 4 This is a schematic diagram of the chain relay link of the present invention. Figure 4 As shown in the figure, this is one of the relay structures for the drone cluster to directly transmit data back to the data center. In the case of the same relay node, this relay structure can provide data services for sensor nodes in more distant areas.
[0188] Figure 5 This is a schematic diagram of the tree-like relay link of the present invention. Figure 5 As shown in the figure, this is one of the relay structures for the drone cluster to directly transmit data back to the data center. In the case of the same relay node, this relay structure can simultaneously provide data services to sensor nodes in a wider area.
[0189] Figure 6 It is an algorithm flow chart of the multi-UAV cooperative multi-hop real-time relay link system of the present invention.
[0190] In this embodiment, considering that drones have limited perception range and can only make decisions based on information within their range, and can only receive control information when connected directly to a data center or via a multi-hop relay, the present invention uses the Dec-POMDP method to model this scenario. Let all information in the environment interacting with the drone cluster be the global state S. Each drone is controlled by its own dedicated agent, which can be considered an agent. In each time slot, the agent first perceives the global state S, obtains the observation O through the observation function Z(S), and then uses the locally deployed policy π(O) to determine the action A for the current time slot. The actions A of all agents act together on the environment, causing changes in the environment. This process is recorded by the state transition function P, and the quality of the transition process is evaluated by the reward function R.
[0191] In this embodiment, multiple basic area units are generated according to the basic information of the environment to be transmitted as follows:
[0192] For drone swarms, sensor nodes are not always independent of each other. Their location distribution characteristics always make drones tend to fly to and provide services to sensor nodes within a region, rather than to a single node. In other words, the inline features of sensor nodes can, to a certain extent, better determine the behavior of drone swarms. Therefore, this application will study how to extract the inline features of sensor nodes.
[0193] Clustering algorithms in unsupervised learning are often used to mine intrinsic features between data. To this end, this paper adopts a classic and effective unsupervised clustering algorithm—the improved K-means algorithm based on Canopy. Compared with the traditional K-means algorithm, this algorithm comprehensively considers global node information and adaptively generates the number of clusters and the location of the initial centroid, greatly improving the stability and accuracy of clustering. Specifically, the average distance of sample elements in a node set D∈K is defined as:
[0194]
[0195] Define the sample density of a sensor node k in the node set D:
[0196]
[0197] When ||q k -q i When ||<MeanDis(D), f=1, otherwise f=0. When a sample element k in the node set D is used as the centroid, the average distance of the samples within the cluster can be expressed as:
[0198]
[0199] It means the average distance of samples of a node set when k is the centroid of the node set. Let s(k) represent the inter-cluster distance of sample element k, which means the distance from sample element k to the high-density sample element. If there is such a sample element k′ such that ρ(k) < ρ(k′), then the distance to the sample closest to sample k is taken. If there is no such sample, then the distance to the sample farthest from sample k is taken: Formula section (next section)
[0200]
[0201] Let ω(k) represent the maximum weight product. The larger its value is, the more suitable the sample element k is as the centroid:
[0202]
[0203] According to the above definition, the process of the Capony-based improved K-Means algorithm is as follows:
[0204]
[0205]
[0206] At this point, we have obtained a cluster unit generated by the intrinsic characteristics of the sensor nodes. This unit will serve as a basic area unit and, combined with the update degree, provide more meaningful reference data for the behavior of the drone cluster.
[0207] For any basic area unit PoI, the higher the frequency of access to the basic area unit PoI, the lower the demand for digital energy services, and vice versa. To this end, the present invention defines an indicator to measure the frequency of access to basic area unit PoIs - the update rate τ, which is used to represent the time interval from the last time a basic area unit PoI was provided with digital energy services to the present. The update process is as follows:
[0208]
[0209] where τ c (t) represents the update degree of the PoI in the cth region in the tth time slot, H 4 and They respectively indicate whether there is a drone providing data service for the c-th PoI in the current time slot.
[0210] In this embodiment, building a Dec-POMDP model for each UAV includes:
[0211] Dec-POMDP reformulation:
[0212] 1) State space S
[0213] The state of the system environment at the tth time slot can be expressed as:
[0214] s(t)={{q n (t)},{α m (t)},{e k (t)},{L},{e n (t)}}
[0215] The elements are the positions of all drones, the data freshness of all data packets, the capabilities of all sensor nodes, the status of all relay links, and the energy consumption of all drones. Obviously, for a single drone n, not all states of the system environment can be known. The observation o of each drone n n (t) will be obtained by the observation function Z(S).
[0216] 2) Observation space O
[0217] The observation of drone n in the tth time slot can be expressed as:
[0218]
[0219] in Indicates the number of data packets collected at the data center. Indicates the total data freshness of the data packets collected at the data center. Indicates that drone n can connect to the data center directly or through a relay link. represents the forward drone n f location, Indicates the number of data packets within the sensing range of drone n, Represents the backward drone n b If there are multiple backward UAVs, their average position is taken. represents the total number of data packets within the sensing range of all UAVs in the backward link, {τ c (t)} represents the update degree of PoI in all regions.
[0220] 3) Action Space A
[0221] Considering that for drone n, adjacent flight directions often have similar effects, we will set the direction of drone n Discretize into N f directions, then the direction of drone n in the tth time slot can be expressed as The action of drone n in the tth time slot can be expressed as:
[0222]
[0223] 4) Reward function R
[0224] The reward of drone n in the tth time slot can be expressed as:
[0225]
[0226] in, and Denote the reward shaping part and the objective function part respectively. Let v1 v2 denote the scalar product operation of v1 and v2, g(ν) is a function that represents the relationship value, when v is true g(v) = 1, otherwise g(v) = 0, then the reward function of the reward shaping part can be expressed as
[0227]
[0228] in Indicates that there is a data packet within the sensing range of drone n or its successor drones, and it can connect directly or through a multi-hop relay link to the data center. In this case, communication is encouraged; otherwise, flight is encouraged; Indicates that there is a data packet within the sensing range of the successor drone of drone n, but it cannot connect directly or through a multi-hop relay link to the data center. In this case, drone n is encouraged to fly to the midpoint of the connection between its predecessor and successor drones; Indicates that there is no data packet within the perception range of either UAV n or the UAVs in its subsequent links. In this case, UAV n is encouraged to derive the direction u according to the update degree of the regional PoI. n (t) to fly. Specifically, let q c (t) represents the centroid position of the PoI in the cth region in the tth time slot, then
[0229]
[0230] in Represents a very small positive number to prevent the denominator from being 0. The sig(x) function is an improved sigmoid function, and its specific form is as follows:
[0231]
[0232] In addition, the reward function of the objective function can be expressed as
[0233]
[0234] in is a proportional factor used to control energy consumption to be consistent with the average data freshness. It represents a time slot before the end of this round of mission, during which the drone cluster is encouraged to collect as many data packets as possible; Indicates that the number of data packets collected at the data center is close to the total number of data packets, and UAV n returns to the vicinity of the data center, which means that the mission of UAV n in this episode is successfully completed. It is worth noting that in the reward function of the objective function, we consider that the task is completed when the number of packets at the data center is close to the total number, rather than when all packets are collected. This is because in actual tests, we found that when setting a target margin After that, it is easier for the agent to complete the task, and in subsequent iterations the agent will gradually collect all data packets.
[0235] In this embodiment, the remodeled Dec-POMDP problem is a mixed action space problem, and there is a coupling relationship between discrete actions and mixed actions - drone n moves along the current direction The flight distance is affected by the time division factor θ n (t), i.e., the discrete action The effect of the continuous action θ n (t). Generally speaking, it is difficult for traditional DRL methods to simultaneously solve the scalability problem of discrete actions, the modal inconsistency problem of mixed action space, and the coupling problem of discrete actions and continuous actions. In order to solve the above problems, this paper will adopt a new hybrid action space DRL algorithm-HyAR-TD3, which solves the scalability problem of discrete problems by introducing the embedding table Embedding, and unifies the modalities of discrete actions and continuous actions. Then, CVAE is introduced to solve the coupling problem of discrete and continuous actions, and the fused mixed actions are mapped to the latent space. Then, the TD3 algorithm is used to train the strategy of mapping from the state space to the latent space, and finally the similarity principle is used to realize the mapping from the latent space to the original mixed action space. The algorithm structure diagram is shown in the figure below. Figure 8 、 9 、10 as shown:
[0236] Figure 8 The schematic diagram of the structure of the CVAE of the present invention is shown in FIG. The structure mainly consists of three parts: encoder, decoder and embedding table. The trainable parameters are χ e χ dThe embedding table takes the unique hot encoding k of the discrete action as input and obtains the discrete latent action z k . It is concatenated with the state s as the conditional vector of CVAE and combined with the continuous action x k Multiply as the input of the encoder, and the output is the continuous latent action mean μ and the continuous latent action standard deviation σ of the same dimension. Re-parameterize the latent action mean and standard deviation to get the latent action z k , and multiplied by the conditional vector of CVAE as the input of the decoder, and finally the reconstructed continuous action x is obtained k- and state transfer residual s - The loss function of CVAE can be expressed as
[0237]
[0238] The first term of the CVAE loss function is the reconstruction error, and the second term is when the input is s,k,x k The KL divergence between the encoder output distribution and the standard normal distribution, and the third term represents the state transfer residual. The data required to update the CVAE network (s, k, x k ,s - ) comes from the experience replay pool D.
[0239] Figure 9 The actual workflow diagram of the present invention is as follows: after obtaining the state s of the environment, it is fed into the Actor network as input to obtain the discrete latent action z k and continuous latent actions On the one hand, the discrete latent action is fed into the embedding table as input, and the discrete action k with the highest matching degree is obtained by using the similarity principle. On the other hand, the discrete latent action z k The input state s is concatenated into the conditional vector of CVAE, and the continuous hidden action Multiply as the input of the decoder to get the continuous action x k Discrete action k and continuous action x k Act on the environment together to get the next state s - , immediate return r and the end mark Done. Finally, s, z k ,z xk ,k,x k ,r,s_,Done form Transition and are sent to the experience replay pool D for learning and training of CVAE and TD3 networks.
[0240] Figure 10 The TD3 structure diagram of the present invention mainly consists of an Actor network and its target network Actor_, two Critic networks and their respective target networks Critic_, and their respective trainable parameters are ω, ω- , ε1, ε2, and The input of the Actor network is the current state s, and the output is the discrete latent action z of the current time slot. k and continuous latent action z xk The output is concatenated with the current state s and used as the input of two Critic networks to obtain two state-action value functions Q1 and Q2, respectively. i The opposite of the value will be used as the loss function to update the Actor network (Q i =Q1). The training of the two critic networks will be implemented by all target networks together. Specifically, the next state s - As the input of Actor_, we get the implicit action z of the next time slot. k- and Similarly, the hidden action and the next state s - Splicing as the input of the two Critic_ networks to obtain the state action value function Q of the next state 1- and Q 2- , let Q - =min(Q 1- ,Q 2- ), then the target value y target =r+γQ - .
[0241] All trainable parameters of the target network will be replaced by the parameters of their corresponding training network after a certain number of training steps.
[0242]
[0243]
[0244] The present application is further described in detail below by way of examples. It should be understood that the examples do not constitute any limitation to the present application.
[0245] In the simulation process of this application, a block with a side length of L is set. m =400m area, with a data center located in the center of the map There are N drones and K sensor nodes, where the default values of N and K are 3 and 30 respectively. The total mission time is divided into T = 140 time slots, and the length of each time slot is t0 = 1s. The sensing range of the drone is R s The default value is 60m, the default value of height H is 50m, and the safe flight distance d safe =8m, maximum flight speed v max =10m / s, transmission power p n =10W. Transmitting power p of sensor nodek =-30dBm, initial energy reserve e k (0) = Q0 = 0.5J, the Poisson distribution parameter λ in the process of sensing data packets s =0.05, normal distribution parameter μ s =0.7, The energy consumption of sensing data packets is e s =10dBm. The channel gain g0 at a reference distance of 1m is g0 = 0.01, and the receiving noise of the drone is σ 2 =-110dBm, transmission bandwidth M FM =1MHz, energy transmission efficiency η = 0.8. Among the parameters related to drone flight energy consumption, d0 = 0.48, P0 = 99.66W, P1 = 120.16W, s0 = 0.002, ρ = 1.225kg / m 3 , s tip =120m / s, Z=0.5s 2 , μ0=0.0001. The following table shows the parameter settings of HyAR-TD3:
[0246]
[0247]
[0248] In order to demonstrate the performance of the algorithm proposed in this application, four other benchmark schemes are compared here: 1) Circular trajectory method. The flight trajectories of the three drones are all circles, with the center as the data center and the radii as R s ,2R s ,3R s , when there is a data packet in the sensor within the sensing range of drone n in the tth time slot θ n (t)=1, otherwise θ n (t) = 0. 2) CS method. First, the K-means algorithm is used to cluster the sensor nodes. Then, the cluster center of each cluster is used as the vertex. The Hamiltonian circuit obtained by the CW saving method will be used as the trajectory of one of the drones n. In order to ensure real-time transmission of data, the other two drones will be located at the three points of the line connecting drone n and the data center. The setting strategy of the time division factor is the same as the circular trajectory method. 3) Nearest neighbor search method. In each time slot, drone n flies with the sensor node closest to it as the target. The other two drones are located as relay nodes. The other two drones will be located at the three points of the line connecting drone n and the data center. The setting strategy of the time division factor is the same as the circular trajectory method. 4) Taboo search. First, the K-means algorithm is used to cluster the sensor nodes. Then, the distance from all cluster centers to the data center is greater than R. sClusters are constructed, and one or more vertices are set additionally to ensure that when there is a drone at the cluster center collecting data, a multi-hop relay link can be built through these vertices. Then, all cluster centers and the additional vertices are used as vertices of the VRP problem, and the tabu search algorithm is used to find the optimal solution trajectory. The setting strategy of the time division factor is the same as that of the circular trajectory method.
[0249] Figure 10 The flight trajectory diagram of the drone cluster is shown. Static trajectory diagrams cannot fully demonstrate the collaborative relationship between the clusters. Let the blue, orange, and green curves represent the flight paths of drones 1, 2, and 3, respectively. It can be seen that at the beginning of the mission, drones 2 and 3 headed for clusters 4 and 6, respectively. Drone 2 used drone 1 as a forward relay, and drone 3 used drone 2 as a forward relay. Then, drone 2 headed for cluster 5, and drone 3 returned to cluster 4, while drone 1 continued to provide relay services for drones 3 and 2 near the data center. After collecting data packets from clusters 4, 5, and 6, drone 1, which was closer to cluster 1, headed first to the area where cluster 1 resided, while drone 2 headed to the area where cluster 2 resided. Drone 3, being farther away, arrived in the second quadrant of the map after drones 1 and 2 were already in their respective areas. Drone 3 was midway between the two drones and the data center, providing relay services for drones 1 and 2. At this point, drones 1, 2, and 3 formed a tree-like relay link. At the end of the mission, drones 2 and 3 returned directly to the data center, ending their flights. Drone 1, after collecting all the data from cluster 3, also returned to the data center, thus completing the mission. In summary, we can see that the algorithm proposed in this paper can comprehensively consider the distribution location of each cluster. By designing appropriate trajectory planning and scheduling, when any drone in the drone cluster recycles data, there will always be other drones to help it build a multi-hop relay link connected to the data center. This makes it unnecessary to temporarily store data packets in the drone's cache, thereby achieving the expected objective function - reducing the average data freshness.
[0250] Figure 11 reflects Figure 10Figure 2 shows the changes in the time-slicing factor for each time slot during the drone's flight. It can be seen that when the time-slicing factor approaches 1, the drone is essentially in a data recovery state or acting as a data relay. Since the drone cluster needs to build multi-hop relay links when providing data services to clusters other than cluster 3, it can be found that when the time-slicing factor of any drone approaches 1, the time-slicing factors of the remaining drones are also close to 1. Of course, we can also see that when the time-slicing factor of drones 1 and 3 approaches 1 during time slot t∈[50,70], while the time-slicing factor of drones 2 approaches 0, this is because drones 1 and 2 form a chain relay link. Although drone 3 passes near this chain relay link, drone 1 does not need drone 3's relay service and has a better option—drone 2. Therefore, drone 3 still flies at full capacity, and its time-slicing factor approaches 0.
[0251] Figure 12 、 13 , 14 respectively show the average data freshness as the different parameters change (the default value is R s =60m, K=30, H=50m). Specifically, Figure 12 The parameter that changes is the number of users K. It can be seen that, except for the circular trajectory method, the average data freshness of the other algorithms increases when the number of users increases. Among them, the performance from best to worst is the algorithm proposed in this article, the taboo search algorithm, the CS method and the nearest neighbor search method. This is because when the number of users increases, the complexity of their distribution positions also increases significantly. The drone cluster needs to spend more energy to ensure that there is a real-time multi-hop relay link to connect to the data center when recovering data. In addition, it can be found that the algorithm proposed by the present invention deteriorates faster after the number of users increases significantly. This is because the additional user location information is not recorded by the trained model, so the performance of the environment uncertainty brought by the increase in users is not satisfactory. Of course, the model trained by the present invention itself is constantly learning online. When the environment changes, it is naturally necessary to further interact with the environment to learn the information brought by the environmental changes, so as to adjust the effect of the model. Figure 11 The parameter that changes is the perception range R s It can be seen that compared with other benchmark algorithms, the average data freshness obtained by the algorithm proposed in this paper decreases more significantly as the perception range increases. This is because the change in the perception range does not essentially introduce new information, but only affects the distribution of states and thus determines the effect of the model output. The state vector designed in this paper has higher robustness to changes. Figure 12 The variable parameter is the flight altitude H. It can also be seen that the average data freshness gradually decreases with the increase of flight altitude, and the algorithm proposed in this paper performs better at the same altitude. However, compared with Figure 11Generally speaking, the rate of change of average data freshness with altitude (increment of average data freshness / increment of altitude) is smaller than the rate of change of average data freshness with perception range (increment of average data freshness / increment of perception range). This is because the change of altitude mainly affects only the search radius of the drone cluster on the ground, while the change of perception range not only affects the search radius of the drone cluster on the ground, but also affects the difficulty of forming multi-hop relay links between drone clusters. Therefore, compared with the change of altitude, the average data freshness performance of the drone cluster is more sensitive to the change of perception range. In general, compared with other benchmark algorithms, the algorithm proposed in this invention has better performance in terms of the number of users K, perception range R, and the number of users K. s The average data freshness performance in the three dimensions of the UAV, flight altitude H, and flight altitude H is more outstanding, and can better cope with the challenges brought about by the information asymmetry between UAVs and the limited control information delivery caused by their limited perception range.
[0252] Figure 15 The trend graph of the average energy consumption of the UAV cluster obtained by the algorithm proposed in this invention as a function of different parameters is shown. The parameters that change are R s ∈{54,60,66,72,78}, H∈{32,38,44,50,56} and K∈{24,30,36,42,48}, the common default parameter value is R s = 60m, H = 50m, K = 30. Considering that the average energy consumption of a drone swarm is primarily determined by flight energy and hovering energy, while the flight trajectory of other benchmark algorithms is not affected by parameter changes, we only demonstrate the performance of the proposed algorithm in terms of average energy consumption for drone swarms. It can be seen that the average energy consumption of a drone swarm increases with the number of users, decreases with the increase in perception range, and increases with the increase in flight altitude. Furthermore, it can be found that the average energy consumption of a drone swarm is most sensitive to changes in the number of users, followed by perception range, and finally flight altitude. This is primarily because the trained model does not incorporate new sensor node information when the number of users increases. Therefore, the average energy consumption increases significantly when the number of users increases. Compared to the flight altitude of drones, the perception range of drones not only affects the search radius of the ground but also the difficulty of drone swarms collaborating to form multi-hop relay links. Therefore, the average energy consumption is more sensitive to changes in perception range. In summary, given a fixed number of users, the average energy consumption of a drone swarm can be reduced by lowering the flight altitude or increasing the perception range.
[0253] It should be noted that the above explanations of the method embodiment are also applicable to the device of this embodiment and will not be repeated here.
[0254] The present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, the above-mentioned multi-UAV collaborative multi-hop real-time relay link design method is implemented.
[0255] The present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the above-mentioned multi-UAV collaborative multi-hop real-time relay link design method.
[0256] Figure 2 This is an exemplary structural diagram of an electronic device that can implement the multi-UAV cooperative multi-hop real-time relay link design method provided according to an embodiment of the present application.
[0257] like Figure 2 As shown, the electronic device includes an input device 501, an input interface 502, a central processing unit 503, a memory 504, an output interface 505, and an output device 506. The input interface 502, the central processing unit 503, the memory 504, and the output interface 505 are interconnected via a bus 507. The input device 501 and the output device 506 are connected to the bus 507 via the input interface 502 and the output interface 505, respectively, and are then connected to other components of the electronic device. Specifically, the input device 501 receives input information from the outside and transmits the input information to the central processing unit 503 via the input interface 502; the central processing unit 503 processes the input information based on the computer-executable instructions stored in the memory 504 to generate output information, temporarily or permanently stores the output information in the memory 504, and then transmits the output information to the output device 506 via the output interface 505; the output device 506 outputs the output information to the outside of the electronic device for use by the user.
[0258] That is to say, Figure 2 The electronic device shown may also be implemented as comprising: a memory storing computer executable instructions; and one or more processors, which can implement the combination of the computer executable instructions when executing the computer executable instructions. Figure 1 The described multi-hop real-time relay link design method for multi-UAV cooperation.
[0259] In one embodiment, Figure 2 The electronic device shown can be implemented to include: a memory 504 configured to store executable program code; and one or more processors configured to run the executable program code stored in the memory 504 to execute the multi-UAV cooperative multi-hop real-time relay link design method in the above-mentioned embodiment.
[0260] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0261] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0262] Computer-readable media include permanent and non-permanent, removable and non-removable media, and media can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), data versatile disk (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0263] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and a part of the module, program segment or code includes one or more executable instructions for realizing the prescribed logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes identified in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or overall flow chart can be implemented using a dedicated hardware-based system that performs the prescribed function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0264] The processor referred to in this embodiment may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0265] The memory can be used to store computer programs and / or modules. The processor implements various functions of the device / terminal equipment by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area. The program storage area can store an operating system and at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created based on the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory can include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0266] In this embodiment, if the module / unit integrated in the device / terminal equipment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. Although the present application is disclosed as above with reference to preferred embodiments, it is not intended to limit the present application. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
[0267] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0268] In addition, it is obvious that the word "comprising" does not exclude other units or steps. Multiple units, modules or devices recited in the device claims can also be implemented by one unit or the entire device through software or hardware.
[0269] Although the present invention has been described in detail above using general descriptions and specific embodiments, it will be apparent to those skilled in the art that modifications and improvements may be made based on the present invention. Therefore, such modifications and improvements, which do not depart from the spirit of the present invention, are intended to be within the scope of protection claimed herein.
Claims
1. A multi-drone collaborative multi-hop real-time relay link design method, characterized in that: The multi-UAV cooperative multi-hop real-time relay link design method includes: Step 1: Obtain basic information about the environment to be transmitted, including the number of drones, drone performance information, the number of sensors, sensor performance information, and the location of each sensor; Step 2: Generate multiple basic area units based on the basic information of the environment to be transmitted; Step 3: Set update information for each basic area unit, where the update information is used to represent the time interval from the last time the basic area unit was provided with data services to the present time; Step 4: Establishing a Dec-POMDP model framework based on the basic information of the environment to be transmitted and the update degree information of each basic area unit; Step 5: Use the HyAR-TD3 algorithm to train and optimize the parameters of the Dec-POMDP model framework to obtain the trained policy model; Step 6: Deploy the trained policy model to each drone.
2. The multi-UAV cooperative multi-hop real-time relay link design method according to claim 1, characterized in that: The step 2: generating a plurality of basic area units according to the basic information of the environment to be transmitted includes: Step 21: Acquire a sensor node set, where the sensor node set includes a plurality of sensor nodes; Step 22: Calculate the sample density of all sensor nodes based on the sensor node set; Step 23: Obtain one of the sensor nodes as the first centroid according to the sample density of all sensor nodes; Step 24: Eliminate other sensor nodes that meet the first condition from the sensor node set according to the first centroid, thereby obtaining a sensor node set after elimination; Step 25: Determine whether there are still sensor nodes in the sensor node set after elimination. If so, Steps 22 to 24 are repeated until no other sensor nodes exist in the sensor node set.
3. The multi-UAV cooperative multi-hop real-time relay link design method according to claim 2, characterized in that: The following formula is used to obtain the update degree information of each basic regional unit: ;in, Indicates the t Time slot c The update degree of the basic regional unit, Indicates that there is a drone in the current time slot c Regional PoI provides digital energy services, Indicates that there is no drone in the current time slot. c Each regional PoI provides digital energy services.
4. The multi-UAV cooperative multi-hop real-time relay link design method according to claim 3, characterized in that: Building a Dec-POMDP model for each drone includes: State Space : No. t The state of the time slot system environment is expressed as: ;in, Indicates the t The state of the system environment in each time slot, For all drones t The position of the time slot, For the t Data freshness of all data packets in a time slot, For the t The capabilities of all sensor nodes in a time slot, L For the t The situation of all relay links in time slots, For the t Energy consumption of all drones in a time slot; Observation Space : No. t Time slot drones n The observation is expressed as: ;in, For the t Time slot drones n Observation, For all drones t The position of the time slot, Indicates the t The number of data packets collected at the data center in the time slot, Indicates the t The sum of the data freshness of the data packets collected at the data center in the time slot, Indicates the t time slot drones n Can be connected directly or via a relay link to the data center, Indicates the t forward drones in time slots location, Indicates the t time slot drones n The number of packets within the sensing range, Indicates the t Backward UAVs with time slots If there are multiple backward UAVs, their average position is taken. Indicates the t The total number of data packets within the sensing range of all drones in the backward link of time slots, Indicates the t The update degree of all basic area units in a time slot; Action Space : No. t Time slot drones n The direction can be expressed as: ,in, Indicates drone n direction, To drone n Direction The direction of discretization, Indicates the t Time slot drones n direction, represents pi; No. t Time slot drones n The action can be expressed as: ,in, Indicates the t Time slot drones n Actions, Indicates the t Time slot drones n direction, Indicates the t Time slot drones n The time division factor, Reward Function : No. t Time slot drones n The return is expressed as: ;in, Indicates the reward shaping part, Represents the objective function part.
5. The multi-UAV cooperative multi-hop real-time relay link design method according to claim 4, characterized in that: The step 5: respectively training and optimizing the parameters of each Dec-POMDP model by the HyAR-TD3 algorithm to obtain the trained Dec-POMDP model includes: The following training optimization is performed for the parameters of each Dec-POMDP model: A1. Randomly initialize the parameters of the Actor network and the two Critic networks , , ; A2. Randomly initialize the parameters of the embedding table, encoder, and decoder , , ; A3. Prepare an experience replay pool D; A4. Repeat the warm-up phase; A5, repeated training stage; A6. End.
6. The multi-UAV cooperative multi-hop real-time relay link design method according to claim 5, characterized in that: The A4, repeated preheating stage includes: A41, using random strategies to collect experience and store it in experience replay pool D; A42, train CVAE every certain number of steps; A43: Until the maximum number of preheating steps is reached, otherwise return to A4.
7. The multi-UAV cooperative multi-hop real-time relay link design method according to claim 6, characterized in that: The A5, repeated training phase includes: A51, repeat each time slot: A511, get the environment status ; A512, the status As input to the Actor, get the implicit action ; A513. According to the similarity principle, Input into the embedding table to obtain discrete actions k ,Will 、 and Input to the decoder to get continuous action ; A512, Mixing Actions k , Input into the environment and get the next state , immediate returns r and end marker ; A513, will Deposit into experience replay pool D; A512, after a certain number of steps, train the TD3 network; A513: Until the maximum number of steps in a training round is reached, otherwise return to A51; A52. After a certain number of rounds, train CVAE; A53. Until the maximum number of training rounds is reached, otherwise return to A5.
8. The multi-UAV cooperative multi-hop real-time relay link design method according to claim 7, characterized in that: The optimization goal of the strategy model is: , , , , , , , in, For the safe flight distance of the drone, The maximum flight speed of the drone.
9. A multi-drone collaborative multi-hop real-time relay link design device, characterized in that: The multi-UAV cooperative multi-hop real-time relay link design device includes: A basic information acquisition module for the environment to be transmitted, wherein the basic information acquisition module is used to acquire basic information of the environment to be transmitted, wherein the basic information of the environment to be transmitted includes the number of drones, drone performance information, the number of sensors, sensor performance information, and the location of each sensor; A basic area unit division module, wherein the basic area unit division module is used to generate a plurality of basic area units according to basic information of the environment to be transmitted; An update rate setting module, the update rate setting module is used to set update rate information for each basic area unit, the update rate information is used to represent the time interval from the last time the basic area unit was provided with data service to the present; A hybrid action space model establishment module, the hybrid action space model establishment module is used to establish a Dec-POMDP model according to the basic information of the environment to be transmitted and the update degree information of each basic area unit; A training optimization module, which is used to train and optimize the parameters of the Dec-POMDP model framework through the HyAR-TD3 algorithm to obtain a trained policy model; A deployment module is used to deploy the trained policy model to each drone.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle cooperative autonomous navigation method and system
CN118424282A