Unmanned aerial vehicle cluster spectrum resource optimization method and related device

By building a drone cluster spectrum resource optimization model and training a multi-agent deep reinforcement learning model, the problem of low spectrum resource utilization efficiency in the interference environment is solved, and efficient optimization and adaptability of spectrum resources are achieved.

CN120456276APending Publication Date: 2025-08-08NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510580944.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In an interference environment with high dynamic changes in communication frequency requirements, traditional single agent deep reinforcement learning strategies are difficult to effectively optimize the spectrum resources of the drone cluster, resulting in low efficiency of spectrum resource utilization and insufficient flexibility and robustness.

Method used

A drone cluster spectrum resource optimization model is built for dynamic interference changes and dynamic changes in communication frequency requirements, and a multi-agent deep reinforcement learning model is trained based on this model. The drone cluster spectrum resource optimization scheme for each time slot is determined through the multi-agent deep reinforcement learning model, including the selection of channel and transmission power.

Benefits of technology

In the interference environment where communication frequency needs are highly dynamically changed, the rational optimization of spectrum resources is achieved, the spectrum resource utilization efficiency of the drone cluster is improved, flexibility and robustness are enhanced, and adaptability is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120456276A_ABST
    Figure CN120456276A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle cluster spectrum resource optimization method and a related device, and relates to the technical field of unmanned aerial vehicle communication, and the method comprises the steps: firstly constructing an unmanned aerial vehicle cluster spectrum resource optimization model in an interference dynamic change and communication frequency demand dynamic change environment, and training the multi-agent deep reinforcement learning model based on the unmanned aerial vehicle cluster spectrum resource optimization model to obtain a trained multi-agent deep reinforcement learning model, and finally determining an unmanned aerial vehicle cluster spectrum resource optimization scheme of each time slot by using the trained multi-agent deep reinforcement learning model. According to the invention, reasonable optimization of spectrum resources can be realized in an interference environment with high dynamic change of communication frequency demands, and the flexibility, robustness and adaptability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of UAV communication technology, and in particular to a method for optimizing spectrum resources of a UAV cluster and related devices. Background Art

[0002] With the rapid development of drone technology, drones are increasingly being used in military, industrial, and civilian applications, demonstrating significant advantages in complex mission scenarios such as disaster relief, remote monitoring, and communications. However, in an interference-heavy environment where communication frequency demands are highly dynamic, spectrum resources are becoming increasingly scarce, and optimizing spectrum resources for drone swarms has become a daunting challenge. Therefore, in the face of the coexistence of external malicious interference and spectrum resource scarcity, optimizing spectrum resources to ensure the effective operation of drone swarms has become a pressing issue.

[0003] In spectrum resource optimization scenarios, deep reinforcement learning can adaptively select the optimal spectrum resources based on the external electromagnetic environment of the drone, thereby improving communication efficiency. However, in the context of increasingly complex mission scenarios and increasingly demanding mission requirements, traditional single-agent deep reinforcement learning strategies have exposed obvious limitations in spectrum resource optimization for drone swarms. This is because in the spectrum resource optimization task of drone swarms, each drone not only needs to make autonomous decisions under limited global information, but also needs to adapt to the dynamically changing electromagnetic environment and effectively respond to electromagnetic interference threats. However, in a distributed execution environment, traditional single-agent deep reinforcement learning strategies have difficulty rationally optimizing the spectrum resources of multiple drones in complex interference environments with highly dynamic communication frequency requirements, thereby affecting the overall spectrum resource allocation performance. These limitations make it difficult for drones to achieve efficient spectrum resource utilization in actual interference environments. Summary of the Invention

[0004] The purpose of this application is to provide a spectrum resource optimization method and related devices for drone clusters, which can realize spectrum resource optimization in an interference environment with highly dynamic changes in communication frequency demand, achieve efficient spectrum resource utilization, and improve flexibility, robustness and adaptability.

[0005] To achieve the above objectives, this application provides the following solutions:

[0006] In a first aspect, the present application provides a method for optimizing spectrum resources of a drone cluster, the method comprising:

[0007] Construct a spectrum resource optimization model for drone swarms in an environment with dynamic changes in interference and communication frequency demand; the spectrum resource optimization model for drone swarms includes an objective function and constraints, the objective function is to maximize the communication rate, and the constraints include channel constraints, transmit power constraints, and transmission delay constraints;

[0008] A multi-agent deep reinforcement learning model is trained based on the drone swarm spectrum resource optimization model to obtain a trained multi-agent deep reinforcement learning model; the multi-agent deep reinforcement learning model includes multiple policy networks and an evaluation network, the number of the policy networks is the same as the number of drones in the drone swarm, and the output of each policy network is connected to the input of the evaluation network;

[0009] The trained multi-agent deep reinforcement learning model is used to determine the drone cluster spectrum resource optimization plan for each time slot; the drone cluster spectrum resource optimization plan includes the channel and transmission power selected by each drone in the drone cluster when transmitting a data packet in the current time slot.

[0010] In a second aspect, the present application provides a UAV swarm spectrum resource optimization device, the UAV swarm spectrum resource optimization device comprising:

[0011] A model building module is used to construct a spectrum resource optimization model for drone swarms in an environment with dynamic changes in interference and communication frequency demand; the spectrum resource optimization model for drone swarms includes an objective function and constraints, the objective function is to maximize the communication rate, and the constraints include channel constraints, transmit power constraints, and transmission delay constraints;

[0012] A training module is configured to train a multi-agent deep reinforcement learning model based on the drone swarm spectrum resource optimization model to obtain a trained multi-agent deep reinforcement learning model; the multi-agent deep reinforcement learning model includes a plurality of policy networks and an evaluation network, the number of the policy networks being the same as the number of drones in the drone swarm, and the output of each policy network being connected to the input of the evaluation network;

[0013] An optimization module is used to use the trained multi-agent deep reinforcement learning model to determine the drone cluster spectrum resource optimization plan for each time slot; the drone cluster spectrum resource optimization plan includes the channel and transmission power selected by each drone in the drone cluster when transmitting a data packet in the current time slot.

[0014] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned method for optimizing spectrum resources of a drone cluster.

[0015] In a fourth aspect, the present application provides a computer-readable storage medium on which a computer program is stored, which, when executed by a processor, implements the above-mentioned method for optimizing spectrum resources of drone clusters.

[0016] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned method for optimizing spectrum resources of drone clusters.

[0017] According to the specific embodiments provided in this application, this application has the following technical effects:

[0018] The present application provides a method and related devices for optimizing spectrum resources of a drone cluster. First, a drone cluster spectrum resource optimization model is constructed for an environment with dynamic changes in interference and dynamic changes in communication frequency demand. The drone cluster spectrum resource optimization model includes an objective function and constraints. The objective function is to maximize the communication rate, and the constraints include channel constraints, transmission power constraints, and transmission delay constraints. Then, based on the drone cluster spectrum resource optimization model, a multi-agent deep reinforcement learning model is trained to obtain a trained multi-agent deep reinforcement learning model. The multi-agent deep reinforcement learning model includes multiple policy networks and an evaluation network. The number of policy networks is the same as the number of drones in the drone cluster. Finally, the trained multi-agent deep reinforcement learning model is used to determine the drone cluster spectrum resource optimization plan for each time slot. The drone cluster spectrum resource optimization plan includes the channel and transmission power selected by each drone in the drone cluster when transmitting a data packet in the current time slot. This application constructs a spectrum resource optimization model for drone clusters in an environment with dynamic changes in interference and dynamic changes in communication frequency demand, and further trains a multi-agent deep reinforcement learning model based on the drone cluster spectrum resource optimization model. Subsequently, the trained multi-agent deep reinforcement learning model can be used to simultaneously optimize and determine the channel and transmission power selected by each drone when transmitting a data packet in each time slot, thereby optimizing the spectrum resources of the drone cluster. This can achieve reasonable optimization of spectrum resources in an interference environment with highly dynamic changes in communication frequency demand, achieve efficient spectrum resource utilization, and be able to adapt to an interference environment with highly dynamic changes in communication frequency demand, thereby improving flexibility, robustness and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 This is an application environment diagram of a method for optimizing spectrum resources of a drone cluster provided in Example 1 of the present application.

[0021] Figure 2 A flowchart of a method for optimizing spectrum resources of a drone cluster provided in Example 1 of the present application.

[0022] Figure 3 Schematic diagram of the system scenario model provided in Example 1 of this application.

[0023] Figure 4 This is a time slot structure diagram provided in Example 1 of the present application.

[0024] Figure 5 Flowchart of the orthogonal gradient multi-agent deep reinforcement learning algorithm provided in Example 1 of this application.

[0025] Figure 6 This is a schematic diagram comparing cluster throughput in the swept frequency interference mode provided in Example 1 of the present application.

[0026] Figure 7 This is a schematic diagram comparing cluster throughput in the intelligent interference mode provided in Example 1 of the present application.

[0027] Figure 8 This is a schematic diagram comparing spectrum conflict rates in the swept frequency interference mode provided in Example 1 of the present application.

[0028] Figure 9 This is a schematic diagram comparing spectrum conflict rates in the intelligent interference mode provided in Example 1 of the present application.

[0029] Figure 10 A schematic diagram showing the comparison of transmission delays in the swept frequency interference mode provided in Example 1 of the present application.

[0030] Figure 11 A schematic diagram showing the comparison of transmission delays in the intelligent interference mode provided in Example 1 of the present application.

[0031] Figure 12 A schematic diagram of the functional modules of a drone cluster spectrum resource optimization device provided in Example 2 of the present application.

[0032] Figure 13 A schematic diagram of the structure of a computer device provided in Example 3 of the present application. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0034] Example 1

[0035] The method for optimizing spectrum resources of drone clusters provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the terminal communicates with the server via a network. The data storage system can store data that the server needs to process. The data storage system can be set up separately, integrated with the server, or located in the cloud or on other servers. The terminal can send a pending optimization request to the server. After receiving the pending optimization request, the server constructs a drone swarm spectrum resource optimization model for the environment with dynamically changing interference and communication frequency demand. The drone swarm spectrum resource optimization model includes an objective function and constraints. The objective function is to maximize the communication rate, and the constraints include channel constraints, transmit power constraints, and transmission delay constraints. A multi-agent deep reinforcement learning model is trained based on the drone swarm spectrum resource optimization model to obtain a trained multi-agent deep reinforcement learning model. The trained multi-agent deep reinforcement learning model is used to determine the drone swarm spectrum resource optimization plan for each time slot. The drone swarm spectrum resource optimization plan includes the channel and transmit power selected by each drone in the drone swarm when transmitting data packets in the current time slot. The server can provide feedback to the terminal on the optimization result of the drone swarm spectrum resource optimization plan for the optimization request.

[0036] In addition, in some embodiments, the drone cluster spectrum resource optimization method can also be implemented independently by a server or a terminal. For example, the terminal can directly process the pending optimization request, or the server can obtain the pending optimization request from the data storage system and process the pending optimization request.

[0037] The terminals may include, but are not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart car devices. Portable wearable devices may include smart watches, smart bracelets, and head-mounted devices. The server may be implemented as a standalone server or a server cluster consisting of multiple servers, or as a cloud server.

[0038] In an exemplary embodiment, Figure 2 As shown, a method for optimizing spectrum resources of drone clusters is provided. The method is executed by a computer device, specifically a computer device such as a terminal or a server, or a terminal and a server. In the embodiment of the present application, the method is applied to Figure 1 The following steps are used as an example to illustrate the server.

[0039] Step S1, constructing a spectrum resource optimization model for drone clusters in an environment with dynamic changes in interference and dynamic changes in communication frequency demand; the drone cluster spectrum resource optimization model includes an objective function and constraints, the objective function is the maximum communication rate, and the constraints include channel constraints, transmission power constraints and transmission delay constraints.

[0040] Step S2: training a multi-agent deep reinforcement learning model based on the drone cluster spectrum resource optimization model to obtain a trained multi-agent deep reinforcement learning model; the multi-agent deep reinforcement learning model includes multiple policy networks and an evaluation network, the number of the policy networks is the same as the number of drones in the drone cluster, and the output end of each of the policy networks is connected to the input end of the evaluation network.

[0041] Step S3, using the trained multi-agent deep reinforcement learning model to determine the drone cluster spectrum resource optimization plan for each time slot; the drone cluster spectrum resource optimization plan includes the channel and transmission power selected by each drone in the drone cluster when transmitting a data packet in the current time slot.

[0042] Multi-agent deep reinforcement learning provides a decentralized learning framework that enables each agent to independently learn and optimize spectrum resources. Although existing technologies have achieved good results in dynamic environments, these methods have not been fully verified to be effective under rapidly changing interference conditions. Therefore, in interference environments, these methods may find it difficult to ensure long-term robustness and stability. On the other hand, these methods have only demonstrated good spectrum resource allocation results in fixed-task scenarios. Their flexibility and responsiveness need to be further verified when faced with highly dynamic changes in communication frequency demand. Although multi-agent deep reinforcement learning has shown great potential in spectrum resource optimization and alleviated the limitations of single-agent deep reinforcement learning methods in environments with scarce spectrum resources, it still faces significant performance bottlenecks in interference environments when faced with highly dynamic changes in communication frequency demand and limited computing resources.

[0043] In response to the above problems, this embodiment proposes a spectrum resource optimization method for drone clusters based on orthogonal gradient multi-agent deep reinforcement learning (abbreviated as OG-MADRL), which aims to cope with the problem of high dynamic changes in communication frequency demand in interference environments. This method significantly improves the cluster throughput of drone clusters in response to high dynamic changes in communication frequency demand by rapidly optimizing spectrum resource strategies. This method can dynamically optimize and schedule spectrum resources according to real-time task loads when interference conditions change rapidly, significantly enhancing the robustness and adaptability of the system and effectively making up for the shortcomings of existing methods in dynamic interference and high dynamic changes in communication frequency demand. In addition, considering the instability of gradients, this method introduces gradient orthogonal initialization and combines gradient clipping technology to prevent gradient explosion, thereby significantly improving training stability.

[0044] like Figure 3 As shown, the system scenario model of this embodiment includes a drone cluster, a base station and a jammer J. The drone cluster includes N drones with the same function, and each drone is denoted as The base station serves as the data fusion center, and the jammer J can impose various types of interference on the drone.

[0045] To avoid mutual interference between uplink and downlink, the uplink uses a highly secure remote control link channel, that is, this embodiment does not consider the uplink frequency, but only the downlink frequency. In this environment, considering that there are multiple downlink transmission tasks, each UAV needs to perform tasks within T time slots, and the time slot set is recorded as Specifically, drones adjust their channels and transmit power in real time based on dynamic environmental changes, transmitting reconnaissance data in the form of data packets to the base station. The base station centrally processes the reconnaissance data from each drone to support the optimization of mission execution strategies. Each drone adaptively selects a channel, prioritizing interference from jammers, and coordinates internal spectrum resource allocation to avoid mutual interference between drones, maximizing cluster throughput and data transmission rate (i.e., communication rate). To address highly dynamic changes in communication frequency requirements, each drone flexibly adjusts its transmit power allocation to ensure that packet transmission meets latency requirements.

[0046] In each time slot t, the jammer affects the signal transmission between the drone and the base station by adjusting the channel and power of the interference signal. In this environment, this embodiment introduces the sweeping interference mode and the intelligent interference mode to simulate the impact of different external interference sources on the spectrum allocation of the drone. In the sweeping interference mode, the behavior of the jammer is predefined, that is, it interferes with the channel according to a fixed sequence in each time slot t. Let the interference sequence in the sweeping interference mode be Indicates a set of channels interfered by the jammer at time slot t. Unlike the frequency sweep jamming mode, the jamming sequence of the intelligent jamming mode is dynamically adjusted according to the channel selected by the drone in the previous time slot t-1. Specifically, the jammer will first try to interfere with the channel with a higher frequency used by the drone. If it is found that the selected channel is already in the current jamming sequence, it will reselect a channel that has not been interfered with, thereby ensuring a more uniform distribution of interference and avoiding repeated interference with the same channel. Let the jamming sequence in the intelligent jamming mode be It represents a set of channels interfered by the jammer at time slot t.

[0047] This embodiment designs the time slot structure of the UAV cluster communication system as follows: Figure 4 A complete transmission time slot structure Ts is divided into four sub-time slots: (1) Interference sensing sub-time slot T per In this time slot, the UAV uses broadband spectrum sensing technology to obtain the current spectrum environment and interference conditions. (2) Decision update sub-time slot T dec In this time slot, the UAV optimizes the channel and transmission power selection according to the interference situation and the decision made in the previous step. (3) Data transmission sub-time slot T trans In this time slot, the UAV selects the appropriate channel and transmission power for data transmission. (4) Learning sub-time slot T ler ,Based on the feedback information from the environment, the UAV adopts ,an intelligent algorithm to learn the channel change pattern and ,provide training information for the spectrum resource optimization decision ,of the next time slot.

[0048] In this system scenario model, spectrum resources are divided into a limited number of discrete channels, namely It represents the set of channels that each UAV can choose during the communication process. At the same time, the transmission power is also discretized into multiple gears, namely represents the set of transmit powers that each UAV can choose during communication. In each time slot, each UAV needs to strike a balance between selecting the best channel and the optimal transmit power to ensure successful transmission while avoiding selecting channels that are interfered by jammers or occupied by other UAVs.

[0049] In this embodiment, the communication between each drone in the drone cluster and the base station adopts FDMA (frequency division multiple access) technology. This communication method allows each drone to communicate independently on different frequency resources, thereby improving the throughput of the drone cluster.

[0050] In this UAV cluster, a channel model combining small-scale fading and large-scale fading is used to describe the characteristics of the wireless channel. Considering that when the UAV is flying at low altitude, despite the presence of some obstacles, the main channel loss is still dominated by the free space path. Therefore, the large-scale fading part mainly considers the path loss of line of sight (LOS) propagation, and the small-scale fading part adopts the Rice fading model to accurately characterize the rapid fluctuations of the signal caused by LOS propagation and multipath scattering and follows a zero-mean complex Gaussian distribution. By considering large-scale fading and small-scale fading at the same time, the dynamic channel characteristics in the UAV communication environment can be accurately described. At time slot t, the Rice fading factor of UAV n on channel c is Expressed as:

[0051]

[0052] Where K is the Rice factor, which represents the intensity ratio of direct line of sight and non-direct line of sight; θ is the phase of the direct line of sight path; θ NLoS To comply with A random variable with zero mean and complex Gaussian distribution.

[0053] Path loss describes the energy attenuation caused by the increase in distance during the transmission of the signal from drone n to the base station through channel c. This process is represented by the free space path loss model, and its formula is:

[0054]

[0055] in, is the free space path loss of the signal transmitted from UAV n to the base station through channel c at time slot t; is the distance that the signal is transmitted from UAV n to the base station through channel c at time slot t; f s is the carrier frequency of the signal; l is the speed of light constant.

[0056] Combining fast fading and slow fading, the channel gain of channel c selected by drone n at time slot t can be obtained: for:

[0057]

[0058] In this communication environment, the signal-to-noise ratio of the channel c selected by drone n at time slot t is for:

[0059]

[0060] in, is the transmission power of UAV n on channel c at time slot t; is the mutual interference power caused by other UAVs on UAV n on channel c at time slot t; N is the number of UAVs in the UAV cluster; is the transmission power of UAV i on channel c at time slot t; is the channel gain of channel c selected by UAV i at time slot t; is the interference power caused to UAV n when the jammer interferes with channel c at time slot t; is the interference power generated by jammer j on channel c at time slot t; is the channel gain of channel c selected by jammer j at time slot t; σ 2 is the noise power density.

[0061] Therefore, the communication rate of drone n in the selected channel c at time slot t is for:

[0062]

[0063] Where B is the channel bandwidth.

[0064] Assume that at each time slot t, the data packet size required for the UAV n task is The conditions for successful transmission are:

[0065]

[0066] If the actual transmission delay T rel (t) exceeds the maximum transmission delay T max (t), the transmission fails, that is:

[0067] T rel (t)≤T max (t);

[0068] Among them, T rel (t) is set to the ratio of the packet size to the communication rate, i.e.

[0069] The optimization goal of this embodiment is to maximize the communication rate of drone n in each time slot under the condition of limited spectrum resources in a scenario with highly dynamic changes in communication frequency demand. By optimizing the transmit power and channel of each drone, the successful transmission of data packets is achieved while avoiding the waste of spectrum resources. Therefore, the optimization problem of this drone cluster, that is, the spectrum resource optimization model of the drone cluster under the environment of dynamic changes in interference and dynamic changes in communication frequency demand, can be expressed as:

[0070] The objective function is:

[0071]

[0072] in, is a marker with the maximum expected objective function; N is the number of drones in the drone cluster; Select the communication rate of channel c for UAV n to transmit data packets in time slot t.

[0073] The constraints are:

[0074]

[0075] Among them, C1 and C2 are channel constraints; is a binary variable, representing whether UAV n selects at most one channel to transmit data packets in time slot t. If yes, it takes 1, otherwise it takes 0; C n (t) is the channel selected by UAV n in time slot t; C is the number of channels; C3 is the transmit power constraint; is the transmission power selected by UAV n when it selects channel c in time slot t; P1 is the first transmission power; P2 is the second transmission power; P3 is the third transmission power; P L is the Lth transmission power; C4 is the transmission delay constraint; T rel (t) is the actual transmission delay of time slot t; T max (t) is the maximum transmission delay of time slot t.

[0076] Maximize the communication rate of each drone in the drone cluster on its selected channel in each time slot, and avoid mutual interference between drones and enhance external interference capabilities as much as possible. Constraint C1 means that each drone can only select one channel in each time slot, and constraint C2 means that the channel selected by each drone in each time slot must be in the channel set. In the constraint C3, the transmission power selected by each UAV in each time slot must be within the power set In

[15] , constraint C4 indicates that the transmission time of each data packet should follow the transmission delay constraint, that is, the actual transmission delay of the data packet when it is transmitted from the UAV cannot exceed the set threshold (i.e., the maximum transmission delay). When the actual transmission delay of the data packet is greater than the maximum transmission delay, it is considered a transmission failure.

[0077] At this time, this embodiment constructs a drone cluster spectrum resource optimization model for an environment with dynamic changes in interference and dynamic changes in communication frequency demand. The drone cluster spectrum resource optimization model includes an objective function and constraints. The objective function is to maximize the communication rate, and the constraints include channel constraints, transmission power constraints, and transmission delay constraints.

[0078] The following describes the spectrum resource optimization method for drone clusters based on orthogonal gradient multi-agent deep reinforcement learning used in this embodiment:

[0079] In an interference environment, each UAV can only obtain channel interference information. Channel interference information includes the channel selection information of nearby UAVs and channel interference information (i.e., whether the channel is interfered by the jammer). Therefore, a UAV spectrum resource optimization method based on orthogonal gradient multi-agent deep reinforcement learning is adopted for optimization. Each UAV interacts with the environment, makes decisions (i.e., selects actions) based on local observation vectors, and continuously learns to optimize channel selection and transmit power selection strategies. Specifically, it is modeled as a multi-agent partially observable Markov decision process (POMDP), represented as a five-tuple: in, is the state space, is the observation space, is the action space, is the state transition probability, is the reward function.

[0080] (1) Action Space

[0081] At time slot t, the action space of the nth UAV is represented by the channel C n (t) and the transmission power P n (t), then the action of the nth UAV at time slot t is expressed as: a n (t)=(C n (t),P n (t)).

[0082] (2) Observation space

[0083] At time slot t, the local observation vector of the nth UAV is expressed as Among them, I n (t) represents the interference information of the channel observed by the n-th UAV (i.e., the observed interference channel), a n (t-1) represents the last action of the nth drone, a j (t-1) represents the previous action of the j-th UAV.

[0084] (3) State space

[0085] State Space Containing all channel states of the UAV cluster and the action choices of each UAV, it can be expressed as the set of observation spaces of all UAVs at time slot t:

[0086] (4) State transition probability

[0087] The state transition probability P(S(t+1)|S(t),A(t)) is determined by the following process: If drone n chooses action a at time slot t n (t), the environment updates the channel state according to the current interference pattern and the channel selection of the UAV. The state of the next time slot S(t+1) is updated according to the current action and interference pattern, which is described as: S(t+1)~P(S(t+1)|S(t),A(t)), where A(t) is the set of actions selected by each UAV in time slot t.

[0088] (5) Reward r

[0089] If UAV n selects a channel C that is not interfered with by external interference or selected by other UAVs at time slot t, n (t), when the size of the data packet transmitted on the channel can meet the current transmission requirements and the delay requirements, the communication is considered successful. At this time, the drone will receive a positive reward, and the reward value is +1, that is:

[0090]

[0091] in, The reward for drone n to choose channel c at time slot t; Select the communication rate of channel c for drone n in time slot t; The packet size selected by drone n for transmission on channel c in time slot t; T rel (t) is the actual transmission delay of time slot t; T max (t) is the maximum transmission delay of time slot t.

[0092] If the size of the data packet transmitted in this time slot meets the mission requirements but does not meet the latency requirements, then the communication is considered a failure and the drone will receive a negative reward of -1, that is:

[0093] when

[0094] If the size of the data packet transmitted at time slot t is not enough to meet the task requirements, it is considered a transmission failure and the reward value is -1, that is:

[0095] when

[0096] When UAV n selects the channel interfered by the jammer at time slot t, the UAV will receive a negative reward with a reward value of -1. When multiple UAVs select the same channel C at time slot t, j(t), then the reward value of the drone is equal to the negative value of the total number of drones that choose the same channel, that is:

[0097]

[0098] Among them, C n (t) is the channel selected by UAV n at time slot t; is the set of interference channels; M is the number of other drones except drone n; 1{.} means that when the conditions in {·} are met, the value is 1; C j (t) is the channel selected by UAV j in time slot t.

[0099] Since this embodiment introduces the sweep frequency interference mode and the intelligent interference mode, the above formula can be expressed as:

[0100]

[0101] The UAV swarm spectrum resource optimization method based on orthogonal gradient multi-agent deep reinforcement learning proposed in this embodiment aims to optimize the spectrum resources of the UAV swarm. The algorithm diagram is shown in the following figure. Figure 5 As shown in the figure. In the context of Multi-Agent Proximal Policy Optimization (MAPPO), each UAV adopts an Actor-Critic structure. Specifically, for each UAV’s Actor network (i.e., policy network), according to the local observation vector o at time slot t, n (t) Select action a n (t), after which all drones perform their action a n (t), the environment will move to the new state and return the reward Update the local observation vector o n (t+1) and termination signal (When the number of iterations reaches the maximum number of iterations, the termination signal is 1, otherwise, the termination signal is 0. When the termination signal is 1, the current round ends and the next round of training begins.) It will be stored in the experience buffer (ie, experience pool) D. The Critic network (ie, evaluation network) generates a state value V based on the state of all drones at time slots t and t+1 φ (s t ) and V φ (s t+1 ).

[0102] During network initialization, orthogonal initialization is used to provide a better starting point for the policy and evaluation networks, enhancing model stability in the early stages of training. During training, gradient clipping is also introduced to constrain the gradient norms of the policy and evaluation networks to a fixed value, preventing instability caused by excessively large gradients and thus improving training stability.

[0103] In order to improve the efficiency and stability of training, the Generalized Advantage Estimation (GAE) technology is adopted. GAE can control the estimated deviation and variance by introducing the balance parameter λ, thereby achieving a smoother policy update process. It is specifically expressed as:

[0104]

[0105] in, is the advantage function; l is the index of a time step, which represents the influence of the system's behavior and rewards in the next l moments; γ is the discount factor; r t+l is the reward for time slot t+l; V φ (s t+l+1 ) is the state value of the t+1+1 time slot; V φ (s t+l ) is the state value of the t+1 time slot.

[0106] The above formula strikes a balance between the deviation of short-term returns and the variance of long-term returns, thereby taking into account the stability of short-term and long-term returns in the drone swarm, ensuring that the advantage estimate is more stable during the update process of the policy network and the evaluation network.

[0107] The first objective function of the policy network is defined as follows:

[0108]

[0109] Among them, L(θ n ) is the first loss function; is the expected function; ρ t (θ n ) is the importance sampling ratio, New Strategy Network In the local observation vector o n (t) Select action a n The probability of (t), For the old policy network In the local observation vector o n (t) Select action a n (t); clip(.) is the clipping function, which is used to limit the range of policy updates and prevent excessive changes in policy updates; ∈ is the clipping parameter.

[0110] Use gradient descent to minimize L(θ n ), calculate the first loss function L(θ n ) to further update the network parameters θ of the policy network of each drone n .

[0111] For the second loss function of the evaluation network, the cumulative reward r n (t) and state value V φ (s t ), V φ (s t+1 ) calculated:

[0112]

[0113] Among them, L(φ) is the second loss function; V φ (s(t) is the state value of time slot t; V φ (s(t+1)) is the state value of the t+1 time slot.

[0114] The evaluation network's parameters, φ, are iteratively updated via backpropagation to optimize the estimate of the state-value function. Gradient descent is used to minimize L(φ), and the gradient of the second loss function, L(φ), is calculated to further update the evaluation network's parameters. This allows the evaluation network to learn an accurate estimate of the expected reward, providing a reliable benchmark for advantage function calculation during policy network updates.

[0115] Table 1 Training process

[0116]

[0117]

[0118] At this time, in this embodiment, the multi-agent deep reinforcement learning model is trained based on the drone cluster spectrum resource optimization model to obtain a trained multi-agent deep reinforcement learning model.

[0119] Among them, the multi-agent deep reinforcement learning model is trained based on the drone cluster spectrum resource optimization model to obtain a trained multi-agent deep reinforcement learning model, which specifically includes:

[0120] (1) Obtain the local observation vector of each drone in the drone cluster in the current historical time slot. The local observation vector includes the interference channel observed by the drone, the drone's action in the previous historical time slot, and the actions of other drones except the drone in the previous historical time slot. The action includes the channel and transmission power selected by the drone when transmitting data packets in the current historical time slot.

[0121] (2) For each drone, the local observation vector of the drone in the current historical time slot is used as input, and the drone's corresponding policy network is used to determine the drone's action in the current historical time slot. Based on the drone's local observation vector and action in the current historical time slot, the drone's local observation vector in the next historical time slot and the drone's reward in the current historical time slot are determined. The drone's local observation vector, action and reward in the current historical time slot and the drone's local observation vector in the next historical time slot are combined into sample data for the current historical time slot and stored in the experience pool. The policy network is used to determine the action based on the drone cluster spectrum resource optimization model.

[0122] (3) The local observation vectors of all UAVs in the current historical time slot are combined into the state of the current historical time slot. The state of the current historical time slot is used as input, and the evaluation network is used to determine the state value of the current historical time slot. The state value of the current historical time slot is used as the state value corresponding to the sample data of the current historical time slot.

[0123] (4) Determine whether the number of iterations reaches the maximum number of iterations and obtain a first determination result.

[0124] (5) If the first judgment result is no, the number of iterations is increased by 1, and the local observation vector of the drone in the next historical time slot is used as the local observation vector of the drone in the current historical time slot, and the step of "for each drone, the local observation vector of the drone in the current historical time slot is used as input, and the strategy network corresponding to the drone is used to determine the action of the drone in the current historical time slot" is returned.

[0125] (6) If the first judgment result is yes, determine whether the number of rounds reaches the first maximum number of rounds to obtain a second judgment result.

[0126] (7) If the second judgment result is no, the number of rounds is increased by 1, and the process returns to the step of “obtaining the local observation vector of each drone in the drone cluster in the current historical time slot”.

[0127] (8) If the second judgment result is yes, then determine whether the number of sample data in the experience pool reaches a preset number to obtain a third judgment result.

[0128] (9) If the third judgment result is yes, a plurality of sample data are randomly sampled in the experience pool, and the policy network and the evaluation network are updated based on the plurality of sample data and the state value corresponding to each sample data to obtain an updated policy network and an updated evaluation network.

[0129] (10) Determine whether the number of rounds reaches the second maximum number of rounds, and obtain a fourth determination result.

[0130] (11) If the fourth judgment result is no, the updated policy network is used as the policy network for the next round, and the updated evaluation network is used as the evaluation network for the next round, and the process returns to the step of "randomly sampling in the experience pool to obtain multiple sample data".

[0131] (12) If the fourth judgment result is yes, the training is terminated and a trained multi-agent deep reinforcement learning model is obtained. The trained multi-agent deep reinforcement learning model includes the updated policy network and the updated evaluation network of the current round.

[0132] Among them, the policy network and the evaluation network are updated based on multiple sample data and the state value corresponding to each sample data to obtain an updated policy network and an updated evaluation network, specifically including: based on multiple sample data and the state value corresponding to each sample data, calculating the first gradient of the first loss function of the policy network and the second gradient of the second loss function of the evaluation network; performing gradient clipping on the first gradient and the second gradient respectively to obtain a first clipped gradient and a second clipped gradient; using the first clipped gradient to update the policy network to obtain an updated policy network; using the second clipped gradient to update the evaluation network to obtain an updated evaluation network.

[0133] Before obtaining the local observation vector of each drone in the drone cluster in the current historical time slot, the drone cluster spectrum resource optimization method of this embodiment also includes: using the orthogonal initialization method to initialize the policy network and the evaluation network respectively. Orthogonal initialization is a method specifically used for initializing the weights of neural networks. It is based on the concept of orthogonal matrices, that is, the rows or columns of the matrix are orthogonal to each other and normalized. This initialization method helps to maintain the scale of the gradient and prevent gradient explosion or disappearance during deep neural network training.

[0134] Related algorithms for large-scale drone clusters may experience collaborative instability and unavoidable interference. This embodiment is based on the MAPPO algorithm and further adopts orthogonal initialization and gradient clipping methods to prevent the algorithm's gradient explosion, thereby achieving stable operation of the drone cluster.

[0135] In this embodiment, a trained multi-agent deep reinforcement learning model is used to determine the drone cluster spectrum resource optimization plan for each time slot. The drone cluster spectrum resource optimization plan includes the channel and transmission power selected by each drone in the drone cluster when transmitting data packets in the current time slot.

[0136] The trained multi-agent deep reinforcement learning model is used to determine the spectrum resource optimization plan for the drone cluster in each time slot, including:

[0137] (1) Obtain the local observation vector of each UAV in the UAV cluster at the current time slot.

[0138] (2) For each drone, the local observation vector of the drone in the current time slot is used as input, and the drone's corresponding strategy network is used to determine the drone's action in the current time slot. Based on the drone's local observation vector and action in the current time slot, the drone's local observation vector in the next time slot and the drone's reward in the current time slot are determined. The drone's local observation vector, action and reward in the current time slot and the drone's local observation vector in the next time slot are combined into a sample data and stored in the experience pool.

[0139] (3) Determine whether the number of sample data in the experience pool reaches the preset number. If so, use all the sample data in the experience pool to update the policy network to obtain the updated policy network, clear all the sample data in the experience pool, and use the updated policy network as the policy network for the next iteration; if not, use the local observation vector of the drone in the next time slot as the local observation vector of the drone in the current time slot, and return to the step of "for each drone, use the local observation vector of the drone in the current time slot as input, and use the policy network corresponding to the drone to determine the drone's action in the current time slot."

[0140] This example studies the spectrum resource optimization method for drone swarms in an interference environment. This method addresses the problem of scarce spectrum resources and the high dynamic changes in communication frequency demand in an interference environment. It proposes a spectrum resource optimization method for drone swarms based on orthogonal gradient multi-agent deep reinforcement learning, aiming to solve the spectrum resource optimization problem faced by drone swarms. This example verifies the effectiveness of the proposed OG-MADRL algorithm through simulation results. In the OG-MADRL algorithm, the neural network used has two hidden layers and uses tanh as the activation function. For this system environment, the jammer's interference mode is set to sweeping frequency jamming and intelligent jamming, and the number of interference channels in each time slot is set to 3. The simulation experiment scenario is set to the number of drones N = 3 and the number of available channels C = 6. The simulation parameters are shown in Table 2.

[0141] Table 2 Simulation parameters

[0142]

[0143] This example compares three different improved strategies: OG-MADRL algorithm; MAPPO algorithm with orthogonal initialization (abbreviated as Without GC); MAPPO algorithm with gradient clipping (abbreviated as Without OI). At the same time, Multi-Agent Deep Deterministic Policy Gradient (MADDPG) and Multi-Agent Twin Delayed DDPG (MATD3) are used as benchmark algorithms for comparison. The simulation results are shown in Figure 2. Figures 6-11 As shown in Tables 3 and 4.

[0144] Table 3 Evaluation indicators in the frequency sweep interference mode

[0145] Algorithm Name Convergence rounds stability OG-MADRL 20 11.6% WithoutGC 28 13.6% WithoutOI 27 17.2% MADDPG 800 20.6% MATD3 820 21.9%

[0146] Table 4 Evaluation indicators in intelligent jamming mode

[0147] Algorithm Name Convergence rounds stability OG-MADRL 30 13.4% WithoutGC 52 18.8% WithoutOI 89 16.3% MADDPG 693 23.9% MATD3 744 27.9%

[0148] Simulation results show that in both the swept-frequency interference mode and the intelligent interference mode, MADDPG and MATD3 significantly outperform OG-MADRL in key metrics such as convergence speed and training stability. Although algorithms using only gradient clipping and only orthogonal initialization outperform MADDPG and MATD3 to some extent, they still exhibit poor stability and significant fluctuations during training. OG-MADRL successfully addresses this issue by combining the advantages of orthogonal initialization and gradient clipping. In terms of cluster throughput, the OG-MADRL algorithm achieves the best throughput in both interference modes, consistently outperforming other compared methods, demonstrating its superiority in complex interference environments. In terms of spectrum collision rate, the OG-MADRL algorithm achieves the lowest overall collision rate and maintains high stability during training. In terms of transmission latency, the OG-MADRL algorithm rapidly reduces latency to below 20 seconds and remains stable with minimal fluctuations throughout training. The above results demonstrate that OG-MADRL, by combining orthogonal initialization with gradient clipping, significantly improves the algorithm's stability, convergence speed, and robustness in interference environments, effectively meeting dynamically changing communication needs. This not only excels in cluster throughput and spectrum conflict rate, but also demonstrates significant advantages in transmission latency, demonstrating its potential for application in drone cluster communications.

[0149] The proposed method is significantly superior to existing methods in improving cluster throughput and system communication stability. In particular, it has higher adaptability and flexibility when facing highly dynamic changes in communication frequency demand, effectively alleviating the problem of spectrum resource optimization in interference environments, demonstrating its superiority in improving UAV cluster throughput and system transmission rate, as well as its application potential in UAV cluster communications.

[0150] Because the communication network in the designed system scenario model is complex, FDMA communication can improve communication efficiency when spectrum resources are scarce. Furthermore, considering that while UAVs are performing their missions in the air, there may be a small number of obstacles between them and the base station, causing multipath effects, but the primary path loss is still dominated by free-space path loss, a Rician fading model is used for channel modeling. Given that the scenario is in a dynamically interfering environment, each UAV cannot respond to environmental changes based solely on its own local observations. Therefore, a multi-agent deep reinforcement learning algorithm (such as the MAPPO algorithm) is employed in a centralized training and distributed execution framework to effectively integrate global information. Furthermore, considering that multi-agent deep reinforcement learning algorithms may suffer from gradient explosion when dealing with high-dimensional state spaces, a combination of orthogonal initialization and gradient clipping is employed. Orthogonal initialization is used at the beginning of training to obtain a robust starting point, and gradient clipping is used during training to constrain gradients to a certain range, thereby preventing gradient explosion.

[0151] To verify the performance of the method in an interference environment, this embodiment designed a simulation scenario combining sweeping interference and intelligent interference, aiming to simulate a real-world environment with scarce spectrum resources and complex interference. This scenario can more realistically reflect the effectiveness of the method in interference environments, especially in terms of spectrum resource optimization, communication decision-making, and system robustness. In an interference environment with highly dynamic communication frequency demand, drone swarms face challenges such as insufficient spectrum resources and poor network adaptability. To address these issues, the method proposed in this embodiment uses distributed decision optimization to respond to spectrum resource changes in real time and improve network adaptability. To address the common problem of gradient instability in discrete action spaces, a spectrum resource optimization method for drone swarms based on orthogonal gradient multi-agent deep reinforcement learning is proposed. By introducing orthogonal initialization at the beginning of training and optimizing the gradient propagation path, instability is suppressed. At the same time, gradient clipping effectively prevents gradient explosion and vanishing. This method significantly improves the network's convergence speed and training stability in discrete action spaces.

[0152] This application also provides an application scenario that applies the above-mentioned drone cluster spectrum resource optimization method. Specifically, the drone cluster spectrum resource optimization method provided in this embodiment can be applied in drone communication scenarios. The drone communication scenario includes a solution optimization link and a communication link. The solution optimization link is used to generate a drone cluster spectrum resource optimization solution for each time slot, and the communication link is used to enable each drone in the drone cluster to communicate according to the drone cluster spectrum resource optimization solution for each time slot. The drone cluster spectrum resource optimization method provided in this embodiment belongs to the solution optimization link.

[0153] Example 2

[0154] Based on the same inventive concept, the embodiments of the present application also provide a drone swarm spectrum resource optimization device for implementing the aforementioned drone swarm spectrum resource optimization method. The implementation solution provided by this device is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations of one or more drone swarm spectrum resource optimization device embodiments provided below can be found in the above-mentioned limitations of the drone swarm spectrum resource optimization method, and will not be repeated here.

[0155] In an exemplary embodiment, Figure 12 As shown, a UAV swarm spectrum resource optimization device is provided, and the UAV swarm spectrum resource optimization device includes:

[0156] The model construction module M1 is used to construct a spectrum resource optimization model for drone clusters in an environment with dynamic changes in interference and dynamic changes in communication frequency demand; the drone cluster spectrum resource optimization model includes an objective function and constraints, the objective function is the maximum communication rate, and the constraints include channel constraints, transmission power constraints and transmission delay constraints.

[0157] The training module M2 is used to train the multi-agent deep reinforcement learning model based on the drone cluster spectrum resource optimization model to obtain a trained multi-agent deep reinforcement learning model; the multi-agent deep reinforcement learning model includes multiple policy networks and an evaluation network, the number of the policy networks is the same as the number of drones in the drone cluster, and the output end of each of the policy networks is connected to the input end of the evaluation network.

[0158] The optimization module M3 is used to use the trained multi-agent deep reinforcement learning model to determine the drone cluster spectrum resource optimization plan for each time slot; the drone cluster spectrum resource optimization plan includes the channel and transmission power selected by each drone in the drone cluster when transmitting a data packet in the current time slot.

[0159] Example 3

[0160] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 13 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for optimizing spectrum resources of a drone cluster is implemented.

[0161] Those skilled in the art will understand that Figure 13 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0162] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method for optimizing spectrum resources of a drone cluster in Example 1 is implemented.

[0163] Example 4

[0164] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which, when executed by a processor, implements the method for optimizing spectrum resources of a drone cluster in Example 1.

[0165] Example 5

[0166] In an exemplary embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the method for optimizing spectrum resources of a drone cluster in Example 1.

[0167] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0168] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0169] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for optimizing spectrum resources of drone clusters, characterized in that: The method for optimizing spectrum resources of a drone cluster includes: Construct a spectrum resource optimization model for drone swarms in an environment with dynamic changes in interference and communication frequency demand; the spectrum resource optimization model for drone swarms includes an objective function and constraints, the objective function is to maximize the communication rate, and the constraints include channel constraints, transmit power constraints, and transmission delay constraints; A multi-agent deep reinforcement learning model is trained based on the drone swarm spectrum resource optimization model to obtain a trained multi-agent deep reinforcement learning model; the multi-agent deep reinforcement learning model includes multiple policy networks and an evaluation network, the number of the policy networks is the same as the number of drones in the drone swarm, and the output of each policy network is connected to the input of the evaluation network; The trained multi-agent deep reinforcement learning model is used to determine the drone cluster spectrum resource optimization plan for each time slot; the drone cluster spectrum resource optimization plan includes the channel and transmission power selected by each drone in the drone cluster when transmitting a data packet in the current time slot.

2. The method for optimizing spectrum resources of drone clusters according to claim 1, characterized in that: The objective function is: Where N is the number of drones in the drone swarm; Select the communication rate of channel c for UAV n to transmit data packets in time slot t; The constraints are: Among them, C1 and C2 are channel constraints; δ n (t) is a binary variable, representing whether UAV n selects at most one channel to transmit data packets in time slot t; C n (t) is the channel selected by UAV n in time slot t; C is the number of channels; C3 is the transmit power constraint; is the transmission power selected by UAV n when it selects channel c in time slot t; P1 is the first transmission power; P2 is the second transmission power; P3 is the third transmission power; P L is the Lth transmission power; C4 is the transmission delay constraint; T rel (t) is the actual transmission delay of time slot t; T max (t) is the maximum transmission delay of time slot t.

3. The method for optimizing spectrum resources of drone clusters according to claim 1, characterized in that: The multi-agent deep reinforcement learning model is trained based on the UAV cluster spectrum resource optimization model to obtain a trained multi-agent deep reinforcement learning model, specifically including: Obtain a local observation vector for each drone in the drone cluster in the current historical time slot; the local observation vector includes the interference channel observed by the drone, the drone's action in the previous historical time slot, and the actions of other drones other than the drone in the previous historical time slot. The action includes the channel and transmit power selected by the drone when transmitting a data packet in the current historical time slot; For each of the drones, the local observation vector of the drone in the current historical time slot is used as input, and the policy network corresponding to the drone is used to determine the action of the drone in the current historical time slot. Based on the local observation vector and action of the drone in the current historical time slot, the local observation vector of the drone in the next historical time slot and the reward of the drone in the current historical time slot are determined. The local observation vector, action and reward of the drone in the current historical time slot and the local observation vector of the drone in the next historical time slot are combined into sample data of the current historical time slot and stored in the experience pool; the policy network is used to determine the action based on the drone cluster spectrum resource optimization model; The local observation vectors of all the drones in the current historical time slot are combined into the state of the current historical time slot, the state of the current historical time slot is used as input, the state value of the current historical time slot is determined by using the evaluation network, and the state value of the current historical time slot is used as the state value corresponding to the sample data of the current historical time slot; Determine whether the number of iterations reaches the maximum number of iterations, and obtain a first determination result; If the first judgment result is negative, the number of iterations is increased by 1, and the local observation vector of the drone in the next historical time slot is used as the local observation vector of the drone in the current historical time slot. The process returns to the step of "for each drone, using the local observation vector of the drone in the current historical time slot as input, and using the policy network corresponding to the drone to determine the action of the drone in the current historical time slot"; If the first judgment result is yes, then determining whether the number of rounds reaches a first maximum number of rounds to obtain a second judgment result; If the second judgment result is no, the number of rounds is increased by 1, and the process returns to the step of "obtaining the local observation vector of each drone in the drone cluster in the current historical time slot"; If the second judgment result is yes, then judging whether the number of sample data in the experience pool reaches a preset number, and obtaining a third judgment result; If the third judgment result is yes, randomly sampling a plurality of sample data in the experience pool, and updating the policy network and the evaluation network based on the plurality of sample data and the state value corresponding to each of the sample data to obtain an updated policy network and an updated evaluation network; determining whether the number of rounds reaches a second maximum number of rounds, and obtaining a fourth determination result; If the fourth judgment result is no, the updated policy network is used as the policy network for the next round, the updated evaluation network is used as the evaluation network for the next round, and the process returns to the step of "randomly sampling in the experience pool to obtain a plurality of sample data"; If the fourth judgment result is yes, the training is terminated to obtain a trained multi-agent deep reinforcement learning model; the trained multi-agent deep reinforcement learning model includes the updated policy network and updated evaluation network of the current round.

4. The method for optimizing spectrum resources of drone clusters according to claim 3, characterized in that: The reward calculation formula is: in, The reward for drone n to choose channel c at time slot t; Select the communication rate of channel c for drone n in time slot t; The packet size selected by drone n for transmission on channel c in time slot t; T rel (t) is the actual transmission delay of time slot t; T max (t) is the maximum transmission delay of time slot t; C n (t) is the channel selected by UAV n in time slot t; is the set of interference channels; M is the number of other drones except drone n; 1{.} means that when the conditions in {·} are met, the value is 1; C j (t) is the channel selected by UAV j in time slot t.

5. The method for optimizing spectrum resources of drone clusters according to claim 3, characterized in that: The policy network and the evaluation network are updated based on the plurality of sample data and the state value corresponding to each of the sample data to obtain an updated policy network and an updated evaluation network, specifically including: Based on the plurality of sample data and the state value corresponding to each of the sample data, a first gradient of a first loss function of the policy network and a second gradient of a second loss function of the evaluation network are calculated; Performing gradient clipping on the first gradient and the second gradient respectively to obtain a first clipped gradient and a second clipped gradient; Updating the policy network using the first clipped gradient to obtain an updated policy network; The evaluation network is updated using the second clipped gradient to obtain an updated evaluation network.

6. The method for optimizing spectrum resources of drone clusters according to claim 5, characterized in that: Before obtaining the local observation vector of each drone in the drone cluster in the current historical time slot, the drone cluster spectrum resource optimization method further includes: initializing the strategy network and the evaluation network respectively using an orthogonal initialization method.

7. A UAV swarm spectrum resource optimization device, characterized in that: The UAV cluster spectrum resource optimization device includes: A model building module is used to construct a spectrum resource optimization model for drone swarms in an environment with dynamic changes in interference and communication frequency demand; the spectrum resource optimization model for drone swarms includes an objective function and constraints, the objective function is to maximize the communication rate, and the constraints include channel constraints, transmit power constraints, and transmission delay constraints; A training module is configured to train a multi-agent deep reinforcement learning model based on the drone swarm spectrum resource optimization model to obtain a trained multi-agent deep reinforcement learning model; the multi-agent deep reinforcement learning model includes multiple policy networks and an evaluation network, the number of the policy networks being the same as the number of drones in the drone swarm, and the output of each policy network being connected to the input of the evaluation network; An optimization module is used to use the trained multi-agent deep reinforcement learning model to determine the drone cluster spectrum resource optimization plan for each time slot; the drone cluster spectrum resource optimization plan includes the channel and transmission power selected by each drone in the drone cluster when transmitting a data packet in the current time slot.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for optimizing spectrum resources of a drone cluster according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for optimizing spectrum resources of a drone cluster according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for optimizing spectrum resources of a drone cluster according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Distributed communication method and system for low-altitude network

    CN121334879A

  • A distributed communication method and system for low-altitude networks

    CN121334879B

  • Unmanned aerial vehicle cluster communication resource allocation method and system based on 5G slicing technology

    CN121357713A