Topology reconstruction method of drone swarm for emergency rescue missions

By decomposing the drone agent network through a multi-agent reinforcement learning algorithm, drone movement and data packet transmission are optimized, which solves the problem of low communication efficiency of the drone cluster communication network after topology reconstruction in emergency rescue missions and realizes efficient collaborative communication.

CN119277473BActive Publication Date: 2025-10-03BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310816905.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-05
Publication Date
2025-10-03
Estimated Expiration
2043-07-05

AI Technical Summary

Technical Problem

The existing drone swarm communication network has low communication efficiency after topology reconstruction during emergency rescue missions. Traditional methods are unable to effectively solve the problems of network segmentation and communication quality degradation caused by drone node failure.

Method used

A multi-agent reinforcement learning algorithm is used to construct a topology reconstruction network. The drone agent network is decomposed into mobile sub-agents and transmission sub-agents, which are responsible for the drone movement direction and the selection of the next-hop transmission node for the data packet, respectively. The reward function is used to optimize the data packet transmission time and constraints to achieve collaborative communication between drones.

Benefits of technology

The communication efficiency of the UAV cluster communication network in emergency rescue is improved, the impact of topology damage on collaborative communication is reduced, and network performance is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119277473B_ABST
    Figure CN119277473B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for reconfiguring the topology of a drone cluster for emergency rescue missions, belonging to the field of drone communication technology. The method solves the problem of low communication efficiency after existing topology reconstruction. The method includes constructing and training a topology reconstruction network, which includes a global hybrid network and local hybrid networks and proxy networks corresponding to each drone. The proxy network receives the local state information of the corresponding drone at each moment and outputs the action function value of the action at each moment. The local hybrid network receives the action function value of the corresponding drone and outputs a local joint action value. The global hybrid network receives the local joint action values ​​of all drones and outputs a cluster joint action value. When a drone fails, the local state information and the previous action of each undamaged drone are obtained and transmitted to the corresponding trained proxy network. The corresponding action is selected according to the ε-greedy strategy to perform topology reconstruction. This achieves topology reconstruction with high communication efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of unmanned aerial vehicle (UAV) communication technology, and in particular to a method for reconfiguring the topology of an UAV cluster under emergency rescue missions. Background Art

[0002] With the rapid development of drone technology, drones are gaining widespread application in fields such as agricultural plant protection, express delivery, and disaster relief due to their adaptability, low cost, and flexible configuration. When disasters such as landslides, fires, and earthquakes occur, the primary task is to obtain disaster information and establish temporary communication networks using drones to provide information support for subsequent rescue efforts. However, due to the complex and harsh post-disaster environment, drone nodes often fail. This not only affects communication quality in the emergency rescue area but also fragments the network into multiple disconnected partitions, hindering data exchange between nodes and even causing network failure, severely impacting the coordination capabilities of the entire drone swarm.

[0003] When a drone relay network is paralyzed due to node failure, topology reconstruction is necessary. By adjusting drone positions and optimizing the network topology from drones to ground stations, a unique information transmission path is planned for each drone. This improves network performance while restoring connectivity and minimizing the impact of drone node failure on drone collaboration. Given the large solution space and numerous complex solutions posed by the dynamic topology of drones, traditional heuristic algorithms struggle to find solutions. Therefore, finding a suitable solution warrants close attention.

[0004] Existing topology reconstruction issues primarily focus on the wireless sensor field. Due to energy depletion and the effects of harsh environments, sensor nodes are prone to failure, resulting in network disconnection. Depending on the scale of network damage, these algorithms are categorized into large-scale fault repair algorithms and small-scale fault repair algorithms. Small-scale fault repair algorithms can be further divided into two types: those that distinguish node importance and those that do not. Both algorithms repair the network by changing the positions of existing drones. The former focuses on restoring network coverage as quickly as possible but increases information transmission overhead, while the latter focuses on repairing network connectivity but reduces efficiency. Large-scale fault repair algorithms primarily rely on adding relay nodes to complete network repairs. Numerous studies have focused on reducing the number of relay nodes and improving network robustness after reconnection.

[0005] Currently, there is limited research specifically focused on UAV topology reconstruction technology. UAV swarm communication networks differ from sensor networks in terms of application requirements and scenarios. The dynamic topology and inter-machine collaboration of UAV networks make it difficult to directly apply wireless sensor network topology repair methods to UAV swarm communication networks. Therefore, it is crucial to design a topology reconstruction method that is efficient and adapts to the characteristics of UAV swarms. Summary of the Invention

[0006] In view of the above analysis, an embodiment of the present invention aims to provide a method for reconfiguring the topology of a drone cluster under emergency rescue missions, so as to solve the problem that the existing method does not consider the cooperative communication of drones, resulting in low communication efficiency after topology reconstruction.

[0007] An embodiment of the present invention provides a method for reconfiguring the topology of a drone cluster in an emergency rescue mission, comprising the following steps:

[0008] Based on the multi-agent reinforcement learning algorithm, a topology reconstruction network is constructed and trained. The topology reconstruction network includes: a global hybrid network and a local hybrid network and agent network corresponding to each drone. The agent network is used to receive the local state information of the corresponding drone at each moment and the corresponding action at the previous moment and output the action function value of each moment. The local hybrid network is used to receive the action function value of the corresponding drone and output the local joint action value. The global hybrid network is used to receive the local joint action value of all drones and output the cluster joint action value.

[0009] When a drone fails, the local state information and the previous actions of the undamaged drones are obtained and passed into the corresponding trained agent network. The action function value of each action of the drone at the current moment is output, and the corresponding action is selected according to the ε-greedy strategy for topology reconstruction.

[0010] Based on a further improvement of the above method, the proxy network is divided into a mobile sub-proxy network and a transmission sub-proxy network according to the drone's movements, which are used to select the drone's movement direction and the next-hop transmission node for the data packet, respectively; the movement directions include: east, south, west, north and hovering, and the next-hop transmission node for the data packet is the neighboring node within the drone's communication range.

[0011] Based on the further improvement of the above method, the mobile sub-agent network and the transmission sub-agent network output two action function values, which are simultaneously passed into the local hybrid network of the corresponding drone to output the local joint action value of the drone; the network parameters of the local hybrid network are generated by a super network that takes the local state information of the drone as input.

[0012] Based on the further improvement of the above method, the network parameters of the global hybrid network are generated by a super network that takes the global state information of the drone cluster as input.

[0013] Based on a further improvement of the above method, the mobile sub-agent network and transmission sub-agent network of each drone have the same reward function, which is used to calculate the reward value obtained by each drone after performing an action; the reward value includes: packet transmission reward value, constraint condition reward value, reconstruction cost reward value and transmission end reward value.

[0014] Based on the further improvement of the above method, the packet transmission reward value is the inverse of the packet transmission time of each hop, the constraint reward value is the preset negative reward value corresponding to the violated constraint condition, the reconstruction cost reward value is the preset negative reward value corresponding to the selected moving direction, and the transmission end reward value is the preset positive reward value for the packet transmission to the ground base station.

[0015] Based on the further improvement of the above method, the packet transmission time is optimized with the goal of minimizing the packet transmission time, which is expressed as follows:

[0016]

[0017] in, Indicates that the data packet is from node k m To node k m+1 The transmission time of this hop, M represents the total number of hops of the data packet transmission path ξ, and W represents the set of flight positions of each UAV.

[0018] Based on further improvements to the above method, data packet transmission also meets the following constraints: the position spacing of the drone between adjacent moments is the product of the flight speed and the time interval; the next-hop transmission node of the data packet is within the communication range of the drone; the remaining queue capacity of the next-hop transmission node of the data packet is greater than or equal to the size of the data packet; the sum of the energy consumed by the drone in performing actions at each moment within the time period is less than or equal to the total energy of the drone; the position of the drone at each moment is within the preset flight area.

[0019] Based on the further improvement of the above method, the transmission time of each hop data packet is calculated by the following steps:

[0020] Calculate the path loss of the direct link based on the next-hop transmission node type and transmission distance of the data packet;

[0021] Calculate the signal-to-noise ratio of the direct link based on the path loss value, the UAV’s transmit power, and the noise power;

[0022] Calculate the transmission rate of the direct link based on the signal-to-noise ratio and the UAV bandwidth;

[0023] The packet transmission time is calculated based on the size and transmission rate of the packet at each hop.

[0024] Based on the further improvement of the above method, the path loss of the direct transmission link is calculated according to the type of the next-hop transmission node and the transmission distance of the data packet, including:

[0025] If the next-hop transmission node type of the data packet is a ground base station, the probability of line-of-sight and non-line-of-sight links is calculated based on the environment, drone altitude, and transmission distance. The path loss of the line-of-sight and non-line-of-sight links is calculated based on the transmission distance, carrier frequency, speed of light, and additional loss of the non-line-of-sight link. The path loss of the direct link is calculated based on the probability and path loss of the line-of-sight and non-line-of-sight links.

[0026] If the next-hop transmission node type of the data packet is a drone, the path loss of the line-of-sight link is calculated based on the transmission distance, carrier frequency, and speed of light, and used as the path loss of the direct transmission link.

[0027] Compared with the existing technology, the present invention can achieve at least one of the following beneficial effects: for the scenario of emergency rescue of drone cluster communication network, comprehensively considering the movement characteristics of drones and communication task requirements, based on the multi-agent reinforcement learning algorithm, the drone's agent network only needs local information, reducing the problem search space; using the divide-and-conquer idea to decompose the agent network of each drone into two sub-agents to reduce the action space, one is responsible for the drone movement decision, and the other is responsible for the data packet next-hop transmission selection decision, so as to better realize the collaboration between drones; with minimizing the data packet transmission time as the optimization goal, the drone's action reward is calculated from four aspects: data packet transmission time, constraints, reconstruction cost and transmission end, and the drone mission requirements are considered while repairing the topology, further optimizing the communication network performance, reducing the impact of topology damage on drone collaborative communication, and improving the communication efficiency of topology reconstruction.

[0028] In the present invention, the above-mentioned technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of the present invention will be described in the following description, and some advantages will become apparent from the description or be learned through practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings are only for the purpose of illustrating particular embodiments and are not to be considered limiting of the present invention. Like reference symbols denote like parts throughout the drawings.

[0030] Figure 1 This is a flow chart of a method for reconfiguring the topology of a drone cluster under an emergency rescue mission in an embodiment of the present invention;

[0031] Figure 2 A schematic diagram of the structure of a topology reconstruction network in an embodiment of the present invention;

[0032] Figure 3 Schematic diagram of a drone cluster network in an embodiment of the present invention. DETAILED DESCRIPTION

[0033] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, and are not used to limit the scope of the present invention.

[0034] A specific embodiment of the present invention discloses a method for reconfiguring the topology of a drone cluster under an emergency rescue mission based on an emergency rescue environment. Figure 1 As shown, the following steps are included:

[0035] S11. Based on the multi-agent reinforcement learning algorithm, a topology reconstruction network is constructed and trained. The topology reconstruction network includes: a global hybrid network and a local hybrid network and an agent network corresponding to each drone; the agent network is used to receive the local state information of the corresponding drone at each moment and the corresponding action at the previous moment and output the action function value of the action at each moment, the local hybrid network is used to receive the action function value of the corresponding drone and output the local joint action value, and the global hybrid network is used to receive the local joint action value of all drones and output the cluster joint action value.

[0036] It should be noted that, considering the need for drone clusters to complete tasks through collaborative communication, it is necessary not only to complete the repair of the drone cluster communication network in a short period of time, but also to ensure the collaborative ability of the drone communication network in the event of node failure. That is, drones must make decisions in a short period of time that can both ensure topology repair and select the appropriate next-hop node to improve communication efficiency, which leads to an increase in the action space. This embodiment uses the idea of ​​divide and conquer to decompose the drone's actions into movement and data packet transmission. Correspondingly, the proxy network of each drone is divided into a mobile sub-proxy network and a transmission sub-proxy network, which are used to select the drone's movement direction and the next-hop transmission node for data packets, respectively. This solves the problems of large drone action space and inter-machine collaboration, so that the drone cluster network still has high communication efficiency after topology reconstruction.

[0037] Specifically, the movement directions of the drone include: east, south, west, north and hovering, and the next-hop transmission node of the data packet is the neighboring node within the communication range of the drone.

[0038] The structural diagram of the topology reconstruction network is as follows Figure 2As shown, the two sub-agent networks for each drone are independent, and parameters of the same sub-agent network can be shared between different drones. The sub-agent network structure consists of three layers: an input layer consisting of a multi-layer perceptron neural network (MLP), an intermediate layer consisting of a gated recurrent neural network (GRU), and an output layer consisting of a multi-layer perceptron neural network (MLP). These layers select actions under an ε-greedy strategy and share the same reward function, which is used to calculate the reward for each drone after executing an action. A local hybrid network receives the two action function values ​​output by the mobile sub-agent network and the transmission sub-agent network corresponding to each drone, monotonically mixes them, and outputs a local joint action value. The local joint action value of all drones is then output by the global hybrid network as the joint action value of the cluster. The parameters of the local and global hybrid networks are generated by the supernetwork, but the input to the supernetwork of the local hybrid network is the local state information of the corresponding drone, while the input to the supernetwork of the global hybrid network is the global state information of the drone cluster. The hypernetwork that generates weights is composed of a single linear layer and an absolute value activation function in turn, ensuring that the weights of the local and global hybrid networks are non-negative. The biases are generated in the same way, but the hypernetwork that generates biases does not have an absolute value activation function and is generated by a 2-layer hypernetwork with ReLU nonlinearity.

[0039] It should be noted that the local status information includes: the location of the current drone and the location of the neighboring nodes within its communication range, the remaining queue capacity of the current drone and the remaining queue capacity of the neighboring nodes within its communication range, the remaining energy of the current drone and its neighboring nodes, and the size of the data packets to be transmitted by the current drone and its neighboring nodes; the global status information includes: the location of all drones, the remaining queue capacity, the remaining energy and the size of the data packets to be transmitted.

[0040] After the topology reconstruction network is constructed, it is trained through the following steps to obtain the trained topology reconstruction network:

[0041] ① Initialize the UAV cluster network and obtain training samples based on the agent network corresponding to each UAV.

[0042] It should be noted that if Figure 3 As shown, the drone cluster network includes K drones, denoted as The drone is on a side with length L s km square area, all drones fly at a height of H and a uniform speed of V. The drones hover near the ground base station and establish communication with the ground base station, which is the end point of data packet transmission. In the preset time period, the number of time slots is N, and the time interval is the time slot length δ t , then the flight trajectory of the UAV, that is, the set of flight positions of each UAV is expressed as ω k[n] represents the position of drone k in time slot n.

[0043] The data packets to be transmitted are randomly set in the UAV cluster network; the experience pool is initialized to empty. It should be noted that the experience pool is a global cache area.

[0044] The local state information of each drone at each moment is collected in the order of time intervals within a preset time period, and is input into the mobile sub-agent network and transmission sub-agent network corresponding to the drone, together with the moving direction and next-hop transmission node of the data packet selected at the previous moment. The moving direction and next-hop transmission node of the data packet are selected at each moment through the ε-greedy strategy. Each drone performs actions according to the moving direction and next-hop transmission node of the data packet at each moment, calculates the reward value according to the reward function, and collects the corresponding local state information at the next moment to obtain the action at the next moment.

[0045] The local state information, moving direction, next-hop transmission node of the data packet, reward value and the corresponding local state information at the next moment of each drone are recorded as a record and stored in the experience pool; the records of all drones at the same moment constitute a data set, and multiple data sets are randomly sampled to obtain training samples.

[0046] Specifically, in the UAV cluster network, the record of UAV k at time t includes: local state information Moving direction Next-hop node for data packets Reward value r k and the local state information at the next moment Each data set contains K records. If the drone has no data packets to transmit, the next hop transmission node of the data packet is the node where the drone itself is located.

[0047] The reward value of each record includes the packet transmission reward value, the constraint reward value, the reconstruction cost reward value and the transmission end reward value. Among them, the packet transmission reward value is the inverse of the packet transmission time of each hop. The shorter the packet transmission time, the greater the reward; the constraint reward value is the preset negative reward value corresponding to the violated constraint condition, the reconstruction cost reward value is the preset negative reward value corresponding to the selected moving direction, and the transmission end reward value is the preset positive reward value for the packet transmission to the ground base station.

[0048] Furthermore, the packet transmission time is optimized to minimize the packet transmission time, which is expressed as follows:

[0049]

[0050] in, Indicates that the data packet is from node km To node k m+1 The transmission time of this hop, M represents the total number of hops of the data packet transmission path ξ.

[0051] The established optimization objectives also satisfy the following constraints:

[0052] 1) The distance between the positions of a drone at adjacent moments is the product of the flight speed and the time interval:

[0053] ||ω k [n+1]-ω k [n]||=Vδ t Formula (2)

[0054] 2) The next-hop node for the data packet is within the communication range of the drone:

[0055]

[0056] Where d represents the preset UAV communication range, and Respectively represent the nodes at k m and node k m+1 The position of the drone in time slot n;

[0057] 3) The remaining queue capacity of the next-hop transmission node of the data packet is greater than or equal to the size of the data packet to avoid network congestion:

[0058]

[0059] in, Indicates that at the next hop node k m+1 The remaining queue capacity of the UAV at position n in time slot n, μ represents the packet size;

[0060] 4) The sum of the energy consumed by the drone during each action within the time period is less than or equal to the total energy of the drone:

[0061]

[0062] in, Indicates that at node k m The energy consumed by the drone at position n to perform the action in time slot n, E max represents the total energy of the drone;

[0063] 5) The drone's position at each moment is within the preset flight area:

[0064] φ l ≤ω k [n]≤φ u Formula (6)

[0065] Among them, φ l and φ u Respectively represent the upper and lower bound coordinates of the preset flight area, that is, the side length is L s The boundary coordinates of the square area of ​​km in the actual scene.

[0066] Each of these constraints is associated with a preset negative reward value. Violating any constraint results in a penalty based on the preset negative reward value. If the mobile subagent network chooses a non-hovering direction, it is penalized based on the preset negative reward value. If a data packet is successfully transmitted to the ground base station, it is rewarded based on the preset positive reward value.

[0067] For example, in one-hop transmission, the packet transmission time is 10, and the packet transmission reward value is -10. However, the movement of the UAV exceeds the flight area, violating the fifth constraint condition, the preset negative reward value of which is -5. The selected moving direction is south, and the preset negative reward value is -2. Because the data packet is not transmitted to the ground base station, no positive reward value is obtained. Therefore, the reward value obtained by the UAV after performing the action is -17.

[0068] Furthermore, the data packet in formula (1) is sent from node k m To node k m+1 The transmission time of this hop Calculate by the following steps:

[0069] Calculate the path loss of the direct link based on the next-hop transmission node type and transmission distance of the data packet;

[0070] Calculate the signal-to-noise ratio of the direct link based on the path loss value, the UAV’s transmit power, and the noise power;

[0071] Calculate the transmission rate of the direct link based on the signal-to-noise ratio and the UAV bandwidth;

[0072] The packet transmission time is calculated based on the size and transmission rate of the packet at each hop.

[0073] Specifically, the path loss of the direct link is calculated based on the next-hop transmission node type and transmission distance of the data packet, including:

[0074] The transmission distance of each hop data packet is calculated by formula (7); if the next hop transmission node type of the data packet is a ground base station, the line-of-sight link probability is calculated by formula (8) according to the environment, drone height and transmission distance. The probability of non-line-of-sight link is calculated by formula (9): According to the transmission distance, carrier frequency, speed of light and the additional loss of non-line-of-sight link, the path loss of the line-of-sight link is calculated by formula (10): The path loss of the non-line-of-sight link is calculated using formula (11): According to the probability and path loss of line-of-sight link and non-line-of-sight link, the path loss of direct link is calculated by formula (12):

[0075]

[0076]

[0077]

[0078] Among them, H represents the flight altitude of the drone, Represents node k m UAV and node k m+1 The distance between the base station / UAV at time slot n, i.e. the data packet from node k m To node k m+1 The transmission distance of this hop, η1 and η2, represent probability parameters, which depend on the flight environment. For example, in a suburban environment, η1 is set to 4.88 and η2 is set to 0.43;

[0079]

[0080]

[0081]

[0082] Among them, f c represents the carrier frequency of the UAV communication system, c represents the speed of light, ψ NLOS Indicates the additional loss of the non-line-of-sight link. Its value depends on the environment and is exemplarily set to 20dB.

[0083] If the next-hop transmission node type of the data packet is a drone, the path loss of the line-of-sight link is calculated by formula (10) based on the transmission distance, carrier frequency and speed of light, and used as the path loss of the direct transmission link

[0084] Furthermore, according to the path loss value, the UAV transmission power and the noise power, the signal-to-noise ratio of the direct link is calculated by formula (13):

[0085]

[0086] in, Indicates that at node k m The transmission power of the UAV at the location, σ 2 Represents the noise power.

[0087] Furthermore, according to the signal-to-noise ratio and the bandwidth of the UAV, the transmission rate of the direct link is calculated by formula (14):

[0088]

[0089] in, Indicates that at node k m The bandwidth of the drone at the location.

[0090] Finally, according to the size and transmission rate of each hop data packet, the data packet transmission time is calculated by formula (15):

[0091]

[0092] Where μ represents the size of the data packet.

[0093] Specifically, in formula (5) It is calculated by the following formula:

[0094]

[0095] c1=2v0 2 Formula (17)

[0096]

[0097] Among them, P0 represents the blade profile power when the drone is in a hovering state, P1 represents the induced power when the drone is in a hovering state, and U tip represents the tip speed of the UAV rotor blade, v0 represents the average rotation induced speed of the UAV, d0 represents the fuselage drag ratio, ρ is the air density, s0 represents the rotation stability, and A represents the rotor rotation area.

[0098] Compared with the existing technology, this step takes packet transmission time as the optimization target and calculates the drone's action reward from four aspects: packet transmission time, constraints, reconstruction cost, and transmission completion. This achieves the goal of considering the drone's mission requirements while repairing the topology, further optimizing the communication network performance.

[0099] ② Build a network with the same structure as the topology reconstruction network as the corresponding target network.

[0100] It should be noted that constructing a target network with the same structure as the topology reconstruction network is used to maintain target value stability and prevent overfitting, thereby improving the stability and convergence speed of the training process. During training, the parameters of the topology reconstruction network are updated instantly via gradients, while the parameters of the target network are not updated via gradients. Instead, the parameters of the topology reconstruction network are copied to the target network at regular intervals.

[0101] ③ During the iterative training process, based on the topology reconstruction network, the cluster joint action value of each training sample is obtained, and based on the target network, the maximum joint action value of each training sample is obtained; according to the maximum joint action value of each training sample, the reward value and attenuation factor of the corresponding training sample, the target joint action value of each training sample is obtained; the sum of the squares of the difference between the cluster joint action value and the target joint action value of the training sample is used as the error function, and the parameters of the topology reconstruction network are updated based on the gradient descent method in back propagation, and the parameters of the topology reconstruction network are periodically copied to the target network; when the maximum number of iterations is reached or the preset accuracy is reached, the trained topology reconstruction network is obtained.

[0102] Specifically, based on the K records of each randomly selected training sample and the records in the experience pool, the local state information o corresponding to each record in the training sample at the current moment is obtained. t and the previous action (a t-1,h and a t-1,d ), the corresponding mobile sub-agent network and transmission sub-agent network in the incoming topology reconstruction network (o t and a t-1,h Incoming transport sub-agent network, o t and a t-1,d The two actions of each record are selected through the ε-greedy strategy to obtain the two action function values ​​of each record, and the two action function values ​​of each record are passed into the corresponding local hybrid network in the topology reconstruction network to obtain the local joint action value of each record, which is then passed into the global hybrid network in the topology reconstruction network to obtain the cluster joint action value Q of each training sample. tot (o, a, s; θ); where o represents the set of local state information of the drone in the current state, a represents the set of drone actions in the current state, s represents the global state information in the current state, and θ represents the parameters of the topology reconstruction network.

[0103] Get the local state information o of the next moment corresponding to each record in the training sample t+1 and the current action (a t,h and a t,d ), the corresponding mobile sub-agent network and transmission sub-agent network in the target network (o t+1 and a t,h Incoming transport sub-agent network, o t+1 and a t,d The two maximum action function values ​​of each record are obtained, and the two maximum action function values ​​of each record are passed into the corresponding local hybrid network in the target network to obtain the maximum local joint action value of each record, and then passed into the global hybrid network in the target network to obtain the maximum joint action value of each training sample. Among them, o' represents the set of local state information of the UAV in the next state, a' represents the set of UAV actions in the next state, s' represents the global state information in the next state, θ - Represents the parameters of the target network.

[0104] According to the maximum joint action value of each training sample, the reward value of the corresponding training sample and the attenuation factor, the target joint action value of each training sample is obtained by the following formula:

[0105]

[0106] in, represents the target joint action value of the i-th training sample, r i represents the average value of each record reward value in the i-th training sample, γ represents the decay factor, represents the maximum joint action value of the i-th training sample.

[0107] According to the cluster joint action value of the training sample and the target joint action value, the error function is calculated by the following formula:

[0108]

[0109] Where b represents the total number of training samples, represents the cluster joint action value of the i-th training sample.

[0110] After the loss is obtained, the parameters of the global hybrid network, the local hybrid networks of all drones, and the proxy network are uniformly updated to maximize the local action value of each drone. When the maximum number of iterations is reached or the preset accuracy is achieved, the trained topology reconstruction network is obtained.

[0111] S12. When a UAV fails, the local state information and the actions of each undamaged UAV at the previous moment are obtained and passed into the corresponding trained agent network. The action function value of each action of the corresponding UAV at the current moment is output, and the corresponding action is selected according to the ε-greedy strategy for topology reconstruction.

[0112] It should be noted that in actual situations, once a damaged drone appears, the network is reconstructed using the topology trained in step S11, and the local state information and the previous action of the undamaged drone are input, so that the undamaged drone learns the action sequence according to the current local state information. According to the action function value of each action, the corresponding action is selected according to the ε-greedy strategy to arrive at the corresponding position and transmit the data packet according to the next-hop transmission node, thereby realizing the reconstruction of the network topology and reducing the impact of node failure on the drone communication network.

[0113] Compared with the existing technology, this embodiment provides a method for reconstructing the topology of a drone cluster under an emergency rescue mission. It targets the scenario of emergency rescue of a drone cluster communication network, comprehensively considers the movement characteristics of the drone and the communication mission requirements, and is based on a multi-agent reinforcement learning algorithm so that the drone's agent network only requires local information, reducing the problem search space; the divide-and-conquer idea is used to decompose the agent network of each drone into two sub-agents to reduce the action space, one responsible for the drone's movement decision, and the other responsible for the next-hop transmission selection decision of the data packet, so as to better achieve collaboration between drones; with minimizing the transmission time of the data packet as the optimization goal, the drone's action reward is calculated from four aspects: data packet transmission time, constraints, reconstruction cost, and transmission end. While repairing the topology, the drone mission requirements are considered, the communication network performance is further optimized, the impact of topology damage on drone collaborative communication is reduced, and the communication efficiency of topology reconstruction is improved.

[0114] Those skilled in the art will appreciate that all or part of the process steps of the above-described embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, such as a magnetic disk, an optical disk, a read-only memory, or a random access memory.

[0115] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention should be covered by the scope of protection of the present invention.

Claims

1. A method for reconfiguring the topology of a drone cluster under emergency rescue missions, characterized in that: The following steps are involved: Based on the multi-agent reinforcement learning algorithm, a topology reconstruction network is constructed and trained, and the topology reconstruction network includes: a global hybrid network and a local hybrid network and an agent network corresponding to each drone; the agent network is used to receive the local state information of the corresponding drone at each moment and the corresponding action at the previous moment and output the action function value of the action at each moment, the local hybrid network is used to receive the action function value of the corresponding drone and output the local joint action value, and the global hybrid network is used to receive the local joint action value of all drones and output the cluster joint action value; wherein, the agent network is divided into a mobile sub-agent network and a transmission sub-agent network according to the action of the drone, which are respectively used to select the moving direction of the drone and the next-hop transmission node of the data packet; the moving directions include: east, south, west, north and hovering, and the next-hop transmission node of the data packet is the drone communication node. The mobile sub-agent network and the transmission sub-agent network output two action function values, which are simultaneously passed into the local hybrid network of the corresponding UAV, and output the local joint action value of the UAV; the network parameters of the local hybrid network are generated by a super network with the local state information of the UAV as input; the network parameters of the global hybrid network are generated by a super network with the global state information of the UAV cluster as input; the local state information includes: the position of the current UAV and the position of the neighboring nodes within its communication range, the remaining queue capacity of the current UAV and the remaining queue capacity of the neighboring nodes within its communication range, the remaining energy of the current UAV and its neighboring nodes, and the size of the data packets to be transmitted by the current UAV and its neighboring nodes; the global state information includes: the position of all UAVs, the remaining queue capacity, the remaining energy and the size of the data packets to be transmitted; When a drone fails, the local state information and the previous actions of the undamaged drones are obtained and passed into the corresponding trained agent network. The action function value of each action of the drone at the current moment is output, and the corresponding action is selected according to the ε-greedy strategy for topology reconstruction.

2. The method for reconfiguring the topology of a drone cluster under an emergency rescue mission according to claim 1 is characterized in that: The mobile sub-agent network and transmission sub-agent network of each drone have the same reward function, which is used to calculate the reward value obtained by each drone after performing an action; the reward value includes: data packet transmission reward value, constraint condition reward value, reconstruction cost reward value and transmission end reward value.

3. The method for reconfiguring the topology of a drone cluster under an emergency rescue mission according to claim 2, characterized in that: The data packet transmission reward value is the inverse of the data packet transmission time for each hop, the constraint condition reward value is the preset negative reward value corresponding to the violated constraint condition, the reconstruction cost reward value is the preset negative reward value corresponding to the selected moving direction, and the transmission end reward value is the preset positive reward value for the data packet transmission to the ground base station.

4. The method for reconfiguring the topology of a drone cluster under an emergency rescue mission according to claim 3 is characterized in that: The packet transmission time is optimized to minimize the packet transmission time, which is expressed as follows: in, Indicates that the data packet is from node k m To node k m+1 The transmission time of this hop, M represents the total number of hops of the data packet transmission path ξ, and W represents the set of flight positions of each UAV.

5. The method for reconfiguring the topology of a drone cluster under an emergency rescue mission according to claim 4, characterized in that: The data packet transmission also meets the following constraints: the distance between the positions of the UAVs at adjacent moments is the product of the flight speed and the time interval; the next-hop transmission node of the data packet is within the communication range of the UAV; the remaining queue capacity of the next-hop transmission node of the data packet is greater than or equal to the size of the data packet; the sum of the energy consumed by the UAV in executing actions at each moment in the time period is less than or equal to the total energy of the UAV; The position of the drone at each moment is within the preset flight area.

6. The method for reconfiguring the topology of a drone cluster under emergency rescue missions according to claim 5, characterized in that: The transmission time of each data packet is calculated by the following steps: Calculate the path loss of the direct link based on the next-hop transmission node type and transmission distance of the data packet; Calculate the signal-to-noise ratio of the direct link based on the path loss value, the UAV’s transmit power, and the noise power; Calculate the transmission rate of the direct link based on the signal-to-noise ratio and the UAV bandwidth; The packet transmission time is calculated based on the size and transmission rate of the packet at each hop.

7. The method for reconfiguring the topology of a drone cluster under emergency rescue missions according to claim 6, characterized in that: Calculating the path loss of the direct link based on the next-hop transmission node type and transmission distance of the data packet includes: If the next-hop transmission node type of the data packet is a ground base station, the probability of line-of-sight and non-line-of-sight links is calculated based on the environment, drone altitude, and transmission distance. The path loss of the line-of-sight and non-line-of-sight links is calculated based on the transmission distance, carrier frequency, speed of light, and additional loss of the non-line-of-sight link. The path loss of the direct link is calculated based on the probability and path loss of the line-of-sight and non-line-of-sight links. If the next-hop transmission node type of the data packet is a drone, the path loss of the line-of-sight link is calculated based on the transmission distance, carrier frequency, and speed of light, and used as the path loss of the direct transmission link.

Citation Information

Patent Citations

  • Unmanned aerial vehicle cluster network self-organization system and method based on task cognition

    CN113316118A

  • Unmanned aerial vehicle cluster network intelligent multi-hop routing method based on multi-agent cooperation

    CN114499648A