Heterogeneous UAV swarm clustering method and system based on deep reinforcement learning

By applying a multi-agent clustering method based on deep reinforcement learning in heterogeneous drone groups, the problem of efficient data transmission in heterogeneous drone groups is solved, and the autonomous clustering and intelligent spectrum sharing of the drone groups are realized, and communication efficiency is improved.

CN119743815BActive Publication Date: 2025-05-23NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510248101.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-23
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

In heterogeneous drone groups, how to achieve efficient data transmission while complying with the constraints of latency, throughput and other communication indicators has become a key issue that needs to be solved urgently.

Method used

The heterogeneous drone clustering method based on deep reinforcement learning is adopted. By creating a multi-agent deep reinforcement learning network, the reward functions of the cluster head agent and cluster member agent are initialized, and the strategy model is trained using the multi-agent near-end strategy optimization algorithm and independent near-end strategy optimization method to achieve autonomous clustering of the drone cluster.

Benefits of technology

The automation and intelligent clustering of drone groups have been realized, the spectrum utilization efficiency and data transmission quality have been improved, and the system's computing complexity and dependence on central control nodes have been reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119743815B_ABST
    Figure CN119743815B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for clustering heterogeneous drone swarms based on deep reinforcement learning, belonging to the field of wireless communication technology, the method comprising: based on a heterogeneous drone swarm data transmission network, creating a multi-agent deep reinforcement learning network that interacts with the environment, and initializing the reward function of the cluster head agent and the cluster member agent; obtaining the observation information and possible actions of the cluster head agent, as well as the observation information and possible actions of the cluster member agent, to be used for the strategy model training of the cluster head agent and the cluster member agent respectively; the cluster head agent and the cluster member agent respectively use a multi-agent proximal strategy optimization algorithm and an independent proximal strategy optimization method to train the strategy model; calling the trained strategy model to complete autonomous clustering of the drone swarm. The present invention optimizes the communication efficiency of the drone data transmission network and improves the intelligence and automation level of spectrum sharing of heterogeneous drone swarms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technology, and in particular to a heterogeneous drone swarm clustering method and system based on deep reinforcement learning. Background Art

[0002] With the continuous advancement of drone technology, especially the breakthroughs in swarm intelligence and collaborative operations, drone swarms are playing an increasingly important role in multiple fields such as auxiliary communications, emergency response, battlefield target positioning, environmental reconnaissance and target strikes with their high cost-effectiveness, mobility and strong anti-destruction capabilities. However, with the increase in the number of drones, the contradiction between the scarcity of spectrum resources and the need to achieve reliable communication has become increasingly prominent. In order to solve this problem, the spectrum sharing strategy of drone swarms has emerged, which achieves effective management of spectrum resources by clustering drone swarms into different groups, especially when performing multiple tasks. In particular, in heterogeneous drone swarms, since the cluster contains multiple types of drones, how to achieve efficient data transmission while complying with the constraints of latency, throughput and other communication indicators has become a key issue that needs to be solved urgently.

[0003] In order to realize spectrum sharing of drone swarms, many scholars have conducted research on clustering strategies for drone swarms, with the focus on improving communication network performance and enhancing network scalability.

[0004] Raja Karmakar et al. used the K-means clustering algorithm to dynamically create clusters based on the location information of drones. This method can automatically adjust the cluster structure according to the movement of drones and maintain the stability of the cluster. However, this method has a large computational overhead and is sensitive to initialization parameters, making it unsuitable for deployment in large-scale drone swarms.

[0005] Guanyu Sun et al. proposed a UAV clustering routing protocol based on the improved particle swarm algorithm (IPSO). First, the IPSO algorithm is used to locate the UAV with high precision, and then the UAV is evenly divided into multiple clusters according to its location information, and the cluster head is selected by comprehensively considering factors such as the distance within the cluster, the remaining energy of the node, and the distance from the base station. However, the computational complexity of this protocol is high, and when the number of nodes is large, it takes a long time to calculate.

[0006] Wenjun Xu et al. proposed a clustering algorithm based on predicted location. They first used Gaussian machine learning to predict the location of drones, clustered the drones according to the predicted location, and dynamically adjusted the clustering radius according to the number of RF links. However, due to the dynamic adjustment of the clustering radius, some drones may be assigned to larger clusters, thereby reducing network efficiency.

[0007] The clustering methods mentioned in the above schemes all rely on a central control node, which limits the robustness and adaptability of the system. To overcome this difficulty, distributed algorithms that can achieve autonomous clustering of drone swarms have attracted more and more attention from scholars.

[0008] Xing Na et al. proposed a clustering method for drone swarms based on coalition game, which regards drones as game participants and clusters as coalitions, and uses coalition game theory to form and maintain clusters. However, the game-based method takes a long time to solve when dealing with dynamic networks, and it is difficult to achieve real-time control of network reorganization. Summary of the invention

[0009] The present invention provides a heterogeneous UAV swarm clustering method and system based on deep reinforcement learning, which realizes spectrum sharing through autonomous clustering of UAV swarms, optimizes the communication efficiency of the UAV data transmission network, and improves the intelligence and automation level of spectrum sharing of heterogeneous UAV swarms.

[0010] In a first aspect, the present invention provides a clustering method for heterogeneous drone swarms based on deep reinforcement learning, comprising:

[0011] Step S1: Based on the heterogeneous UAV swarm data transmission network, a multi-agent deep reinforcement learning network that interacts with the environment is created, and the reward functions of the cluster head agent and the cluster member agent are initialized; wherein the heterogeneous UAV swarm data transmission network includes two types of UAVs: cluster head UAVs and cluster member UAVs. After clustering, the cluster member UAVs transmit data back to the corresponding cluster head UAVs, and each cluster member UAV can only join one cluster; each cluster head UAV corresponds to a cluster head agent deployed with a decision model, and each cluster member UAV corresponds to a cluster member agent deployed with a decision model; the decision model is constructed based on a deep reinforcement learning algorithm;

[0012] Step S2: Obtaining observation information and possible actions of the cluster head agent, as well as observation information and possible actions of the cluster member agents, to be used for strategy model training of the cluster head agent and the cluster member agents respectively;

[0013] Step S3: The cluster head agent and the cluster member agents respectively train the strategy model using the multi-agent proximal strategy optimization algorithm and the independent proximal strategy optimization method;

[0014] Step S4: Call the trained strategy model to complete autonomous clustering of the drone swarm.

[0015] In a second aspect, the present invention further provides a heterogeneous drone swarm clustering system based on deep reinforcement learning, comprising:

[0016] An initialization module is configured to create a multi-agent deep reinforcement learning network that interacts with the environment based on a heterogeneous drone swarm data transmission network, and initialize the reward functions of the cluster head agent and the cluster member agent; wherein the heterogeneous drone swarm data transmission network includes two types of drones: cluster head drones and cluster member drones, and the cluster member drones transmit data back to the corresponding cluster head drones after clustering is completed, and each cluster member drone can only join one cluster; each cluster head drone corresponds to a cluster head agent deployed with a decision model, and each cluster member drone corresponds to a cluster member agent deployed with a decision model; the decision model is constructed based on a deep reinforcement learning algorithm;

[0017] An acquisition module configured to acquire observation information and possible actions of the cluster head agent, and observation information and possible actions of the cluster member agents, for use in strategy model training of the cluster head agent and the cluster member agents, respectively;

[0018] The training module is configured such that the cluster head agent and the cluster member agents respectively train the policy model using a multi-agent proximal policy optimization algorithm and an independent proximal policy optimization method;

[0019] The clustering module is configured to call the trained strategy model to complete autonomous clustering of the drone swarm.

[0020] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the heterogeneous drone swarm clustering method based on deep reinforcement learning as described in any one of the above are implemented.

[0021] In a fourth aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described methods for clustering heterogeneous drone swarms based on deep reinforcement learning.

[0022] The heterogeneous UAV swarm clustering method and system based on deep reinforcement learning provided by the present invention have the following beneficial effects compared with the prior art:

[0023] (1) The present invention introduces a heterogeneous multi-agent deep reinforcement learning algorithm into the formation of UAV swarms, realizing the automation and intelligence of heterogeneous UAV swarm clustering;

[0024] (2) It adopts a distributed learning architecture, which eliminates the need for a central control node and improves the system's anti-destruction capabilities;

[0025] (3) By adjusting the formation of the drone swarm, spectrum sharing is achieved among the drone swarm, which improves the spectrum utilization efficiency and data transmission quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0027] Figure 1 It is a flow chart of a heterogeneous UAV swarm clustering method based on deep reinforcement learning provided by the present invention;

[0028] Figure 2 It is a schematic diagram of the structure of the heterogeneous UAV swarm data transmission network provided by the present invention;

[0029] Figure 3 It is a schematic diagram of the cluster member UAV link maintenance probability prediction model provided by the present invention;

[0030] Figure 4 It is a schematic diagram of the framework of the multi-agent deep reinforcement learning network provided by the present invention;

[0031] Figure 5 is an example of a formation generation simulation result provided by the present invention;

[0032] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0034] It should be noted that, in the description of the embodiments of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "include one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.

[0035] The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0036] Combine the following Figure 1-Figure 6 The present invention describes a method and device for clustering a swarm of heterogeneous drones based on deep reinforcement learning provided in an embodiment of the present invention.

[0037] Figure 1 is a flow chart of a heterogeneous UAV swarm clustering method based on deep reinforcement learning provided by the present invention, such as Figure 1 As shown, including but not limited to the following steps:

[0038] Step S1: Based on the heterogeneous UAV swarm data transmission network, a multi-agent deep reinforcement learning network that interacts with the environment is created, and the reward functions of the cluster head agent and cluster member agents are initialized.

[0039] If there are a total of A regular drone. After clustering, ordinary drones have only one cluster head to establish a communication link with them. At this time, ordinary drones are called cluster members of the corresponding cluster. To achieve load balancing, each cluster head has at least one cluster member to establish a link with it. After clustering, ordinary drones are divided into Cluster members collaborate with each other for collaborative perception. However, due to the limitations of onboard computing power and endurance, cluster members usually need to upload the collected environmental perception data to the cluster head with stronger comprehensive capabilities. Considering the coordinated movement of the drone swarm and the constant changes in the environment, cluster members need to autonomously complete clustering and continuously adjust, while the cluster head adaptively finds the optimal data transmission location.

[0040] like Figure 2As shown in the figure, the heterogeneous drone swarm data transmission network contains two types of drones: cluster head drones and cluster member drones. After clustering, cluster member drones transmit data back to the corresponding cluster head, and each cluster member drone can only join one cluster. In particular, cluster member drones can periodically transmit and receive Hello packets to the surrounding area. Each drone is equipped with a set of reinforcement learning algorithms, which realize environmental interaction through the hardware equipment of the drone, and network with each other as independent agents to form a multi-agent deep reinforcement learning network. At the same time, a distributed network architecture is adopted, without the need for a central control node.

[0041] According to the type of UAV, the agents are divided into two categories: cluster head agents (CH agents) and cluster member agents (CM agents). Specifically, each cluster head UAV corresponds to a cluster head agent deployed with a decision model, and each cluster member UAV corresponds to a cluster member agent deployed with a decision model; the decision model is built based on a deep reinforcement learning algorithm, such as an Actor-Critic network.

[0042] Then, the reinforcement learning environment parameters and reward function are initialized according to the real-world task parameters. The reward functions of the cluster head agent and cluster member agents are described as follows:

[0043] 1) Cluster head UAVs have strong communication and data processing capabilities, and can rely on the air access platform to adopt a model training strategy of centralized training and decentralized execution. Therefore, in order to achieve better training results, the cluster head agent adopts a hybrid reward mechanism, including global rewards and local rewards, to encourage the agent to maximize the total data throughput and maximize the minimum signal-to-noise ratio with the cluster member agents. Cluster head agent The global reward and local reward are specifically described as: and .

[0044] in, is a global reward that encourages the agent to maximize the total data throughput, is a local reward that encourages the agent to maximize the minimum signal-to-noise ratio, Cluster head agent With cluster member agents e The signal-to-noise ratio of the channel between represents any cluster, It can be expressed as:

[0045] .

[0046] In summary, the reward function of the CH agent can be expressed as: ;

[0047] in, Used to balance global rewards and local rewards.

[0048] 2) Due to the limitations of communication and data processing capabilities, the cluster member drones adopt a completely decentralized training method and are trained only based on local observation information. e The reward function can be expressed as:

[0049] ;

[0050] in, Encourage agents to establish stable connections with their neighbors. Penalize isolated points and encourage agents to establish connections with neighbors and participate in cluster collaboration. Penalize the agent's delay to reduce the agent's collaboration delay. represents the number of neighbor agents, and is the discount factor, Representing an Agent e Whether it is an isolated point in the routing network, specifically:

[0051] .

[0052] In particular, Representing an Agent e With its neighbors The average link maintenance probability of is. The link maintenance probability calculation model based on mobility prediction is as follows: Figure 3 As shown, assuming that the drone j At the moment Relatively still, while the drone At relative speed Fly at a constant speed in the direction of the dotted arrow. During the flight, both can send and successfully receive Hello message packets from each other. The dotted line represents the drone. The maximum communication distance is 100m, and the drone outside the dotted line cannot communicate with Normal communication. The solid line with a single arrow represents a drone. UAV Mobility prediction part. Based on the above assumptions, the link maintenance probability can be expressed as:

[0053] .

[0054] in, T Indicates the time threshold of Hello packets. Indicates drone The distance from the current position to the communication edge along the current moving direction, Indicates drone The flight speed, especially, and Both can be estimated through the received Hello data packets.

[0055] Step S2: Obtain the observation information and possible actions of the cluster head agent, as well as the observation information and possible actions of the cluster member agents, to be used for the strategy model training of the cluster head agent and the cluster member agents respectively.

[0056] At any time, due to the differences in communication and data processing capabilities between the cluster head UAV and the cluster member UAV, as well as different task requirements, the corresponding environmental observation information and available actions of the two types of agents are also different, which can be specifically described as follows:

[0057] 1) Any cluster head agent exist The observation information in a time step can be expressed as At any time , Agent The observation information can be expressed as: ,in, For intelligent agents Your own position, For intelligent agents The cluster ID of the service, The cluster ID that serves other cluster head agents, For intelligent agents The locations of all cluster member agents in the service cluster, is the location of the unserved cluster member agent. Indicates that except for the agent Other cluster head agents besides Representing an Agent All cluster member agents in the cluster served. All possible actions of the cluster head agent can be expressed as At any time , Agent The action of is a two-dimensional discrete action, that is, ,in, Representing an Agent The moving direction of 1, 2, 3, and 4 represent the agent moving to the east, south, west, and north, respectively. Represents that the agent does not move.

[0058] 2) Cluster member agents can predict the movement of neighboring drones based on the Hello message packet, thereby obtaining the speed and relative position relationship of neighboring drones. The observation information within a time step can be expressed as: At any time , Agent e The observation information can be expressed as: ,in, For intelligent agents e location, For intelligent agents e speed, For intelligent agents e The cluster ID of For intelligent agents e The communication delay, For intelligent agents e The location of neighboring agents, For intelligent agents e Neighbor Agents speed, For intelligent agents e Neighbor Agents All possible actions of the cluster head agent can be expressed as At the moment , Agent e The discrete actions taken are , represents the agent e Deciding to join a cluster .

[0059] Step S3: The cluster head agent and the cluster member agents respectively train the strategy model using the multi-agent proximal strategy optimization algorithm and the independent proximal strategy optimization method;

[0060] In order to realize the autonomous formation of drone swarms, the present invention designs the following Figure 4 Figure 1 shows a multi-agent deep reinforcement learning framework. In reinforcement learning, the optimal action of the agent is generated by an actor-critic network. The Actor network, also known as the policy network, is responsible for selecting actions based on the current state. The Critic network predicts future benefits based on the action-state pairs to evaluate the quality of the selected action.

[0061] Specifically, in a drone group, the cluster head drone and the cluster member drones have different tasks and roles. Therefore, in the present invention, heterogeneous agents are used, that is, a type of agent is designed for each of the cluster head drone and the cluster member drone. Figure 4 As shown in the figure, there are two types of agents: cluster head agents and cluster member agents. These two types of agents can have different observation information and actions in a shared environment. At the same time, different reward functions are designed for the two types of agents to motivate cluster head agents and cluster member agents to act in the direction of their respective cumulative rewards. In this way, the drone swarm can self-organize into a corresponding formation without a central control node.

[0062] The deep reinforcement learning algorithm training process in the present invention is as follows:

[0063] 1) First, initialize the Actor network parameters of the cluster member agents and cluster head agents ( , ) and Critic network parameters ( , ), the corresponding policy network is ( , ), and create a replay cache for each class of agents ( , ).

[0064] 2) In each training iteration, the cluster member agents and cluster head agents interact with the environment, generate their own training trajectories, and record the current rewards , and the state at the next moment , Among them, the current action is based on the old strategy network , After each episode of training, the trajectories of the two types of agents are collected. , , and calculate the estimated value of the value function , and advantage function value , .

[0065] Then the data and Stored in the cache of cluster member agents and cluster head agents respectively and middle.

[0066] 3) Before the policy gradient is updated, the data in the cache is shuffled and divided into batches, of which The batch data can be specifically expressed as:

[0067] .

[0068] 4) Finally, the strategy of each Actor network is updated, with the training direction of maximizing the proximal strategy optimization algorithm pruning objective function, so as to achieve a balance between strategy stability and update. In this process, the update amount of network parameters and It can be described as:

[0069]

[0070]

[0071] in, and Respectively express and gradient.

[0072] 5) After the strategy update is completed, the old Actor network and Critic network are updated:

[0073] , , , .

[0074] Step S4: Call the trained strategy model to complete autonomous clustering of the drone swarm.

[0075] After the strategy model training is completed, the cluster head drone and cluster member drones run the corresponding strategy model and output the actions to be performed according to the current state of the drone: , During operation, each drone performs actions independently without the need for a central control node, and completes autonomous clustering of the drone swarm based on necessary communications with other drones, improving communication efficiency and thus achieving spectrum sharing of the drone swarm.

[0076] The effect of the present invention can be further illustrated by simulation:

[0077] 1. Simulation conditions: Assume that a large heterogeneous drone group is randomly and evenly distributed in the mission area. The drones are divided into two types: cluster heads and cluster members. Cluster heads are distributed at an altitude of 300 meters, while cluster members are randomly distributed between 200-250 meters. The relative flight speed of cluster members is 0-20m / s. The speed of cluster heads is 100m / s without special instructions. The transmission power of cluster members is 20dBm, and the minimum SNR threshold for cluster heads to successfully complete data transmission is 0dB. The center frequency of the carrier is 2.4GHz, and the channel bandwidth is 10MHz. The background noise power is -100dBm. The preparation delay and forwarding delay of the Hello data packet are 0 and 50μs respectively, the Hello data packet time threshold is 1s, the maximum cooperative communication distance of cluster members is 400 meters, and the path attenuation factor is 2.

[0078] 2. Simulation content: 9 cluster heads and 90 cluster members are deployed in an area of ​​3km×3km for clustering. The clustering results are as follows Figure 5 shown.

[0079] The present invention also provides a heterogeneous UAV swarm clustering system based on deep reinforcement learning, comprising:

[0080] An initialization module is configured to create a multi-agent deep reinforcement learning network that interacts with the environment based on a heterogeneous drone swarm data transmission network, and initialize the reward functions of the cluster head agent and the cluster member agent; wherein the heterogeneous drone swarm data transmission network includes two types of drones: cluster head drones and cluster member drones, and the cluster member drones transmit data back to the corresponding cluster head drones after clustering is completed, and each cluster member drone can only join one cluster; each cluster head drone corresponds to a cluster head agent deployed with a decision model, and each cluster member drone corresponds to a cluster member agent deployed with a decision model; the decision model is constructed based on a deep reinforcement learning algorithm;

[0081] An acquisition module configured to acquire observation information and possible actions of the cluster head agent, and observation information and possible actions of the cluster member agents, for use in strategy model training of the cluster head agent and the cluster member agents, respectively;

[0082] The training module is configured such that the cluster head agent and the cluster member agents respectively train the policy model using a multi-agent proximal policy optimization algorithm and an independent proximal policy optimization method;

[0083] The clustering module is configured to call the trained strategy model to complete autonomous clustering of the drone swarm.

[0084] It should be noted that the heterogeneous drone swarm clustering system based on deep reinforcement learning provided in an embodiment of the present invention can execute the heterogeneous drone swarm clustering method based on deep reinforcement learning described in any of the above embodiments during specific operation, which will not be elaborated in this embodiment.

[0085] Figure 6 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630 and a communication bus 640, wherein the processor 610, the communication interface 620 and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute the heterogeneous drone swarm clustering method based on deep reinforcement learning.

[0086] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the heterogeneous drone swarm clustering method based on deep reinforcement learning provided in the above-mentioned embodiments.

[0087] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A clustering method for heterogeneous drone swarms based on deep reinforcement learning, characterized in that: include: Step S1: Based on the heterogeneous UAV swarm data transmission network, a multi-agent deep reinforcement learning network that interacts with the environment is created, and the reward functions of the cluster head agent and the cluster member agent are initialized; wherein the heterogeneous UAV swarm data transmission network includes two types of UAVs: cluster head UAVs and cluster member UAVs. After clustering, the cluster member UAVs transmit data back to the corresponding cluster head UAVs, and each cluster member UAV can only join one cluster; each cluster head UAV corresponds to a cluster head agent deployed with a decision model, and each cluster member UAV corresponds to a cluster member agent deployed with a decision model; the decision model is constructed based on a deep reinforcement learning algorithm; Step S2: Obtaining observation information and possible actions of the cluster head agent, as well as observation information and possible actions of the cluster member agents, to be used for strategy model training of the cluster head agent and the cluster member agents respectively; Step S3: The cluster head agent and the cluster member agents respectively train the strategy model using the multi-agent proximal strategy optimization algorithm and the independent proximal strategy optimization method; the step S3 specifically includes: Initialize the Actor network parameters of cluster member agents and cluster head agents ( , )、Critic network parameters( , ), the corresponding policy network is ( , ), and create a replay cache ( , ); In each training iteration, the agent interacts with the environment, generates training trajectories and records the current reward , and the state at the next moment , ; Among them, the current action is based on the old strategy network , Generate; After each episode of training, collect the trajectories of the two types of agents , , and calculate the estimated value of the value function , and advantage function value , ; Then the data and Stored in the cache of cluster member agents and cluster head agents respectively and middle; Before the policy network gradient is updated, the data in the cache is shuffled and divided into B batches; In the policy network gradient update, the policy of each Actor network is updated with the maximization of the proximal policy optimization algorithm pruning objective function as the training direction; After the policy network is updated, the old Actor network and Critic network are updated; Step S4: Call the trained strategy model to complete autonomous clustering of the drone swarm.

2. The method for clustering heterogeneous drone swarms based on deep reinforcement learning according to claim 1 is characterized in that: Initialize the reward functions of the cluster head agent and cluster member agents, including: Determine the cluster head agent The reward function is: ; in, is a global reward used to encourage the agent to maximize the total data throughput; is a local reward used to encourage the agent to maximize the minimum signal-to-noise ratio between the cluster head and cluster members; is the discount factor; Determine cluster member agents e The reward function is: ; Among them, the parameters It is used to encourage the agent to establish a stable connection relationship with its neighbors. The stable connection relationship is characterized by the average link maintenance probability. Used to penalize isolated points and encourage agents to establish connections with neighbors; parameters Used to impose penalties on the agent's latency.

3. The method for clustering heterogeneous drone swarms based on deep reinforcement learning according to claim 1, characterized in that: Obtain observation information and possible actions of the cluster head agent, including: Any cluster head agent exist The observation information in a time step is expressed as , at any time , Agent The observation information is expressed as: ,in, For intelligent agents Your own position, For intelligent agents The cluster ID of the service, The cluster ID that serves other cluster head agents, For intelligent agents The locations of all cluster member agents in the service cluster, is the location of the unserved cluster member agent, Indicates that except for the agent Other cluster head agents besides Representing an Agent All cluster member agents in the served cluster; All possible actions of the cluster head agent can be expressed as , at any time , Agent The action is a two-dimensional discrete action ,in, Cluster head agent The cluster to be served, Representing an Agent The direction of movement; The cluster head agent makes decisions among possible actions based on the observed information.

4. The method for clustering heterogeneous drone swarms based on deep reinforcement learning according to claim 1, characterized in that: Obtain observation information and possible actions of cluster member agents, including: Any cluster member agent e exist The observation information in a time step can be expressed as: , at any time , Agent e The observation information is expressed as: ,in, For intelligent agents e location, For intelligent agents e speed, For intelligent agents e The cluster ID of For intelligent agents e The communication delay, For intelligent agents e Neighbor Agents location, For intelligent agents e Neighbor Agents speed, For intelligent agents e Neighbor Agents The cluster ID of All possible actions of the cluster head agent can be expressed as , at the time , Agent e The discrete actions taken are , represents the agent e Deciding to join a cluster ; Cluster member agents make decisions among possible actions based on observed information.

5. The method for clustering heterogeneous drone swarms based on deep reinforcement learning according to claim 1, characterized in that: The step S4 comprises: The cluster head drone and cluster member drones run the trained strategy model and output the actions to be performed according to the current state of the drone; Among them, each drone operates independently without the need for a central control node, and autonomously completes the autonomous clustering of the drone swarm.

6. The method for clustering heterogeneous drone swarms based on deep reinforcement learning according to claim 1, characterized in that: Cluster member drones periodically send Hello message packets to the surroundings. Cluster member agents predict the movement of neighboring drones based on the Hello message packets, thereby obtaining the speed and relative position relationship of neighboring drones.

7. A heterogeneous UAV swarm clustering system based on deep reinforcement learning, characterized in that: include: An initialization module is configured to create a multi-agent deep reinforcement learning network that interacts with the environment based on a heterogeneous drone swarm data transmission network, and initialize the reward functions of the cluster head agent and the cluster member agent; wherein the heterogeneous drone swarm data transmission network includes two types of drones: cluster head drones and cluster member drones, and the cluster member drones transmit data back to the corresponding cluster head drones after clustering is completed, and each cluster member drone can only join one cluster; each cluster head drone corresponds to a cluster head agent deployed with a decision model, and each cluster member drone corresponds to a cluster member agent deployed with a decision model; the decision model is constructed based on a deep reinforcement learning algorithm; An acquisition module configured to acquire observation information and possible actions of the cluster head agent, and observation information and possible actions of the cluster member agents, for use in strategy model training of the cluster head agent and the cluster member agents, respectively; The training module is configured as a cluster head agent and a cluster member agent respectively using a multi-agent proximal strategy optimization algorithm and an independent proximal strategy optimization method to train a strategy model; specifically, it includes: Initialize the Actor network parameters of cluster member agents and cluster head agents ( , )、Critic network parameters( , ), the corresponding policy network is ( , ), and create a replay cache ( , ); In each training iteration, the agent interacts with the environment, generates training trajectories and records the current reward , and the state at the next moment , ; Among them, the current action is based on the old strategy network , Generate; After each episode of training, collect the trajectories of the two types of agents , , and calculate the estimated value of the value function , and advantage function value , ; Then the data and Stored in the cache of cluster member agents and cluster head agents respectively and middle; Before the policy network gradient is updated, the data in the cache is shuffled and divided into B batches; In the policy network gradient update, the policy of each Actor network is updated with the maximization of the proximal policy optimization algorithm pruning objective function as the training direction; After the policy network is updated, the old Actor network and Critic network are updated; The clustering module is configured to call the trained strategy model to complete autonomous clustering of the drone swarm.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the heterogeneous drone swarm clustering method based on deep reinforcement learning as described in any one of claims 1 to 6 are implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the heterogeneous drone swarm clustering method based on deep reinforcement learning as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Unmanned aerial vehicle cluster energy-saving anti-interference communication method based on network reinforcement learning

    CN118175551A

  • Multi-task-oriented unmanned aerial vehicle cluster spectrum resource allocation method and device, and medium

    CN118764874A