Multi-unmanned aerial vehicle track and communication frequency optimization method based on block reinforcement learning
By applying a dual-network block joint optimization algorithm with block reinforcement learning in drone communication services, the trajectory and communication frequency of the drone are optimized, and communication efficiency problems in the scarce spectrum resources and interference environments are solved, and efficient and fair communication services are achieved.
Patent Information
- Application Number
- CN202510211506.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-03
AI Technical Summary
In drone communication services, in the environment of scarce spectrum resources and interference, how to efficiently optimize the trajectory and communication frequency of the drone to maximize the fair weighted rate sum of ground users.
Using a dual-network block joint optimization algorithm based on block reinforcement learning, UAV agents are trained to optimize frequency usage and flight actions in complex electromagnetic spectrum environments, improving the performance limit of the algorithm and accelerating the convergence process.
In the environment of scarce spectrum resources and interference, the fair weighting rate sum of drone communication services is improved, efficient trajectory-use frequency joint optimization is achieved, and the adaptability and communication quality of drones are enhanced.
Smart Images

Figure CN120091329A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of UAV communication optimization, and particularly relates to a multi-UAV trajectory and communication frequency optimization method based on block reinforcement learning. Background Art
[0002] Unmanned Aerial Vehicles (UAVs), relying on their autonomy, flexibility, and the characteristics of being able to carry multiple payloads, have achieved deep integration and wide application in many fields and can provide communication services for ground users in various harsh environments. They have fully integrated into many core fields such as military and civilian, becoming a key force in promoting industrial upgrading and social development. From strategic reconnaissance and precision strikes in the military field to logistics distribution, surveying and mapping exploration, agricultural plant protection, and emergency rescue in the civilian field, the application scenarios of UAVs are constantly being refined and expanded, and the importance of their role is also increasing day by day. However, the problems faced by UAVs are also developing in the direction of diversification and complexity. The rise of artificial intelligence technology has brought new ideas and methods to solve the problems faced by UAV communication. Its subordinate branch, reinforcement learning technology, has emerged in the field of UAV communication resource optimization. Compared with traditional convex optimization algorithms and heuristic algorithms, reinforcement learning can better adapt to high-dynamic scenarios related to UAVs and optimize communication resources in real time. In addition, reinforcement learning can flexibly adapt to different types of communication scenarios without being limited to a specific mathematical model.
[0003] With the continuous expansion of the scale of UAV applications, the limited spectrum resources are becoming increasingly tense, and the interference problem between communication systems is also becoming increasingly prominent. The communication service tasks faced by UAVs are also developing in the direction of diversification and complexity. On the one hand, as the available spectrum resources gradually become scarce, UAVs will have to face the "internal and external troubles" situation of coexisting internal interference and external interference when providing communication services for ground users. On the other hand, the dynamicity of the positions of ground users poses higher requirements for the trajectory planning ability of UAVs, and UAVs need to adjust their flight trajectories in real time to maintain the communication quality of ground users. How to efficiently optimize the relevant communication resources of UAVs has become the focus of attention of researchers in recent years. Based on the above problems, the present invention uses block reinforcement learning technology to design a dual-network block joint optimization algorithm for the scenario of UAV serving communication users in an interference environment. When the transmission power of UAVs is limited, it improves the upper limit of the performance of joint optimization, accelerates the convergence process of the algorithm, and maximizes the fairness weighted rate sum of users in the interference environment. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a multi-UAV trajectory and communication frequency optimization method based on block reinforcement learning in view of the deficiencies of the above-mentioned existing technologies. By using the block reinforcement learning technology, a dual-network block joint optimization algorithm is designed for the scenario of UAV serving communication users in an interference environment. When the UAV transmission power is limited, the upper limit of the performance of joint optimization is improved, the convergence process of the algorithm is accelerated, and the fairness weighted rate sum of users in the interference environment is maximized.
[0005] To achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:
[0006] A multi-UAV trajectory and communication frequency optimization method based on block reinforcement learning includes the steps of:
[0007] (1) Establish a system model of a UAV base station serving ground users in a complex electromagnetic spectrum environment. The system model of the UAV base station serving ground users consists of UAVs, ground users, and interference sources;
[0008] (2) Establish a channel model including K UAVs, M ground users, and P available communication frequency bands for UAVs, and then obtain the communication rate of ground users;
[0009] (3) Construct a system performance evaluation index, and establish an optimization problem of maximizing the fairness weighted rate sum of users based on the index;
[0010] (4) Model the optimization problem as a distributed observable Markov decision process, and use the block reinforcement learning method to train the UAV agent.
[0011] To optimize the above technical solution, the specific measures taken also include:
[0012] In step (1), the system model of the UAV base station serving ground users includes K UAVs, M ground users, P available communication frequency bands for UAVs, and 1 interference source at a fixed position; among them, all UAVs fly at a constant altitude, and a single UAV serves multiple users simultaneously through orthogonal frequency division multiplexing. There is no mutual interference between users, and different UAVs with the same frequency will cause mutual interference to non-self-service users. The communication quality of all users served by UAVs with the same frequency as the interference source is affected.
[0013] In step (2), the channel model is specifically: assume that the total duration of UAV flight and providing communication services is T, and this period contains E equal-length time slots, and the bandwidth of each frequency band is B tr , and the communication rate of ground user m at time slot t is: Ra m (t) = B tr log 2 (1 + SINR m (t)), where the signal-to-interference-plus-noise ratio of ground user m is: P tr represents the transmission power of the UAV, N represents the noise in the environment, and p k,m represents the UAV 's service strategy for ground users The specific expression is The air-ground channel model between UAV k and ground user m is: where T 0 =(4πf carr / c) 2 f carr represents the carrier frequency, c represents the speed of light, d k,m represents the distance from UAV k to user m, α represents the path loss constant, P los and P nlos represent the probability values of the line-of-sight link and the non-line-of-sight link respectively, λ los and λ nlos represent the attenuation coefficients of the line-of-sight link and the non-line-of-sight link respectively. The co-channel interference brought by other UAVs operating on the same frequency as the UAV serving the current user m is: where indicates whether UAV j, which does not serve user m, operates on the same frequency as UAV k serving user m and causes co-channel interference to it. Its form is the same as p k,m The interference received by user m from external interference sources is: where ch jam,m (t) represents the channel model between jammer jam and ground user m, and its form is the same as ch k,m (t), P jam represents the transmission power of the interference source, indicates whether the UAV serving ground user m operates on the same frequency as the interference source.
[0014] The system performance evaluation index in step (3) is: where The optimization problem is: where a freq,k represents the frequency selection of the UAV. The UAV can only select one frequency band in a time slot, and a traj,k represents the flight action of the UAV.
[0015] In step (4), the optimization problem is modeled as a distributed partially observable Markov decision process. The frequency - using network and the trajectory network equipped on the UAV have their own independent local state information, action spaces, and reward functions. During the training process of UAV k, the agent contains a tuple: <local state information, action, reward>. The global state information includes the observation information of all agents.
[0016] The local state information of the frequency - using network consists of four parts: the position of the UAV itself, the positions of other UAVs within the sensing range, the frequency bands occupied by the current interference sources, and the frequency bands occupied by other UAVs within the sensing range, which is expressed as: Among them, represents the position of UAV k itself, represents the position of UAV j within the sensing range, f Jammer represents the frequency band occupied by the current interference source, represents the frequency band occupied by UAV j within the sensing range. The action of the frequency - using network means that in each time slot, the UAV can select an available frequency band from P frequency bands. The reward of the frequency - using network indicates that all frequency - using networks share a global reward function: Rew freq = η system - pun 1 - pun 2 where, η system represents the weighted sum rate fairness of ground users, pun 1 represents the penalty term for the UAV and the interference source using the same frequency, pun 2 represents the mutual interference penalty term between UAVs.
[0017] The local observation information O traj,k of the trajectory network consists of four parts: the position of the UAV itself, the positions of other UAVs within the sensing range, the states of ground users, and the coordinates of obstacles, which is expressed as: Among them, represents the state of ground user m, Loca barr represents the coordinates of the obstacle. The state of the ground user S GU includes its own position, communication rate, and service status. The flight actions of the UAV are defined as {forward, backward, fly right, fly left, hover}. The reward of the trajectory network indicates that all frequency - using networks share a global reward function: Rew traj = η system - pun 3 - pun 4 where, pun 3 is the penalty term for the UAV colliding with an obstacle, pun 4 is the penalty term for collisions between UAVs.
[0018] The above-mentioned block-based reinforcement learning method for training the UAV agent is specifically as follows: Two deep reinforcement learning networks are used to respectively make decisions on the frequency usage actions and flight actions of the UAV. Assume that the number of frequency bands that the UAV can select is num freq , and the number of flight actions that can be selected is num traj . In the block-based dual-network structure, each network selects actions in its own action space. In the single-network joint decision-making algorithm using the fused action space, the UAV selects actions in an action space with a dimension of num freq ×num traj . After receiving the sensing information, each network outputs the corresponding frequency usage and flight actions, updates the local state information and reward value of the frequency usage decision network, and then updates the local state information and reward value of the trajectory decision network. The user communication rate included in its local state information and the sum of the user fairness weighted rates included in the reward function will also be affected by the transmission frequency band selected by the frequency usage decision network. Both networks share the global state information through the mixing network in the QMix training framework. After updating the reward function and local state information, each network stores the local state information O(t), the selected action a, the reward value r, and the local state information O(t + 1) at the next moment in its own experience pool. Finally, the two networks respectively draw experiences from their own experience pools to update the parameters.
[0019] The present invention has the following beneficial effects:
[0020] The present invention proposes a multi-UAV trajectory and communication frequency usage optimization method based on block-based reinforcement learning, which uses a frequency usage decision network and a trajectory decision network with their own independent action spaces to output a joint optimization strategy, avoiding the influence of the large-size fused action space on the UAV reinforcement learning training speed. During the training process, each network focuses on capturing the unique state information in various action decision-making processes to improve the performance upper limit of the algorithm. The simulation results show that the method proposed in this paper can enable the UAV to adapt to the spectrum resource-scarce environment filled with time-varying electromagnetic interference, efficiently generate a reliable trajectory-frequency usage joint strategy, and improve the sum of the user fairness weighted rates of ground user communication. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flowchart of the method of the present invention;
[0022] Figure 2 is a schematic diagram of the system model involved in the embodiment of the present invention;
[0023] Figure 3 is a schematic diagram of the block-based dual-network structure involved in the embodiment of the present invention;
[0024] Figure 4 is a schematic diagram of the time slot structure involved in the embodiment of the present invention;
[0025] Figure 5 This is for the comparison of the performance of different algorithms in the swept-frequency interference environment involved in the embodiments of the present invention;
[0026] Figure 6 This is for the comparison of the performance of different algorithms in the probabilistic interference environment involved in the embodiments of the present invention. Specific embodiments
[0027] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0028] Although the steps in the present invention are arranged with reference numerals, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It can be understood that the term "and / or" used herein relates to and covers any and all possible combinations of one or more of the associated listed items.
[0029] This embodiment is based on the system model of the UAV base station serving ground users as shown in Figure 2 In the case of the UAV base station serving ground users system model, Figure 1 The flowchart of the present invention is shown as Figure 1 As shown, the multi-UAV trajectory and communication frequency optimization method based on block reinforcement learning of the present invention includes the steps:
[0030] (1) Establish a system model of a UAV base station serving ground users in a complex electromagnetic spectrum environment. The system model of the UAV base station serving ground users consists of UAVs, ground users and interference sources;
[0031] The system model of the UAV base station serving ground users includes K UAVs, M ground users, P communication frequency bands available for UAVs, and 1 interference source at a fixed position; the sets of UAVs and ground users in the communication system are respectively represented as and The communication frequency band set is Each UAV has no prior knowledge of the channel selection tendency of other UAVs or external interference sources, but can sense whether there is external interference in all channels. All UAVs fly at a constant altitude. A single UAV serves multiple users through orthogonal frequency division multiplexing at the same time. There is no mutual interference between users. Different UAVs with the same frequency will cause mutual interference to non-self-served users. The communication quality of all users served by UAVs with the same frequency as the interference source is affected.
[0032] (2) Establish a channel model including K UAVs, M ground users, and P available communication frequency bands for UAVs, and then obtain the communication rate of ground users;
[0033] Assume that the total duration of UAV flight and providing communication services is T, which contains E equal-length time slots during this period, and the bandwidth of each frequency band is B tr , and the communication rate of ground user m at time slot t is:
[0034] Ra m (t) = B tr log 2 (1 + SINR m (t)),
[0035] where the signal-to-interference-plus-noise ratio of ground user m is P tr represents the transmission power of the UAV, N represents the noise in the environment, and p k,m represents the service strategy of the UAV to ground user
[0036]
[0037] The air-to-ground channel model between UAV k and ground user m is:
[0038]
[0039] where, T 0 =(4πf carr / c) 2 , f carr represents the carrier frequency, c represents the speed of light, d k,m represents the distance from UAV k to user m, α represents the path loss constant, P los and P mlos respectively represent the probability values of the line-of-sight link and the non-line-of-sight link, λ los and λ nlos respectively represent the attenuation coefficients of the line-of-sight link and the non-line-of-sight link.
[0040] The co-channel interference caused by other UAVs operating on the same frequency as the UAV currently serving user m is:
[0041]
[0042] where, indicates whether other UAVs j that do not serve user m are operating on the same frequency as the UAV serving user m and causing co-channel interference to it, and its form is the same as p k,m .
[0043] The interference received by user m from external interference sources is:
[0044]
[0045] where ch jam,m (t) represents the channel model between the jammer jam and the ground user m, and its form is the same as ch k,m (t), P jam represents the transmission power of the interference source, indicates whether the UAV serving the ground user m has the same frequency as the interference source.
[0046] (3) Construct system performance evaluation indicators, and establish an optimization problem of maximizing the weighted rate sum of user fairness based on the indicators;
[0047] To avoid the situation where the UAV completely abandons some users due to interference, the Jain index is considered to characterize the fairness of ground users receiving communication services. The weighted rate sum of user fairness is used as the system performance evaluation indicator:
[0048]
[0049] where
[0050] The ultimate goal is to maximize the weighted rate sum of user fairness in the interference environment by jointly optimizing the flight trajectory of the UAV and the frequency band used under the limited transmission power of the UAV. Therefore, the optimization problem can be expressed in the following form:
[0051]
[0052] a freq,k is the frequency selection of the UAV, and a traj,k is the flight action of the UAV.
[0053] (4) Model the optimization problem as a distributed observable Markov decision process, and use the block reinforcement learning method to train the UAV agent;
[0054] As Figure 3 shown is the block double-network structure. The above optimization problem is modeled as a distributed partially observable Markov decision process. The frequency selection network and the trajectory network equipped with the UAV have their own independent local state information, action space, and reward function. The three most critical tuples of the agent during the training process of UAV k (there is UAV j and ground user m within its sensing range) are: <local state information, action, reward>, and the global state information includes the observation information of all agents.
[0055] The local state information of the frequency - using network consists of four parts: the position of the UAV itself, the positions of other UAVs within the sensing range, the frequency bands occupied by the current interference sources, and the frequency bands occupied by other UAVs within the sensing range, which is expressed as:
[0056]
[0057] The frequency - using network action represents that each UAV can select an available frequency band from P frequency bands in each time slot, and the frequency - using network reward represents that each frequency - using network shares a global reward function:
[0058] Rew freq = η system - pun 1 - pun 2
[0059] Among them, η system represents the weighted rate sum of ground - user fairness, and pun 1 represents the penalty term for the UAV and the interference source using the same frequency, and pun 2 represents the mutual - interference penalty term between UAVs.
[0060] The local observation information of the trajectory network consists of four parts: the position of the UAV itself, the positions of other UAVs within the sensing range, the states of ground users, and the coordinates of obstacles, which is expressed as:
[0061]
[0062] The state of the ground user S GU includes its own position, communication rate, and service - receiving situation. The flight actions of the UAV are defined as {forward, backward, fly right, fly left, hover}. The trajectory - network reward represents that each frequency - using network shares a global reward function:
[0063] Rew traj = η system - pun 3 - pun 4
[0064] Among them, pun 3 is the penalty term for the UAV colliding with obstacles, and pun 4 is the penalty term for UAVs colliding with each other.
[0065] A block - based dual - network model is proposed, which uses two deep reinforcement - learning networks to separately decide the frequency - using actions and flight actions of the UAV. Assume that the number of available frequency bands that the UAV can select is num freq , and the number of available flight actions that the UAV can select is num traj, under the block double - network structure, each network selects actions in its own action space. In the single - network joint decision - making algorithm using the fused action space, the UAV needs to select actions in a huge action space with a dimension of num freq ×num traj . The detailed structure of the block double - network is as shown in Figure 3 . In the figure, the purple arrow part indicates that the frequency - using decision network and the trajectory decision network synchronously sense the spectrum and the real environment in the same time slot. After receiving the sensing information, each network outputs the corresponding frequency - using and flight actions, updates the local state information and the reward value of the frequency - using decision network. Next, as shown by the red arrow and the brown arrow, update the local state information and the reward value of the frequency - using decision network. This update process not only depends on the frequency - using actions output by itself, but also will be affected by the flight actions output by the trajectory network at the same time: the UAV position information in the local observation, the user fairness - weighted rate sum included in the reward function. Subsequently, update the local state information and the reward value of the trajectory decision network. The user communication rate included in its local state information and the user fairness - weighted rate sum included in the reward function will also be affected by the transmission frequency band selected by the frequency - using decision network. Both networks share the global state information through the mixing network in the QMix training framework. After updating the reward function and the local state information, each network stores the current - moment local state information O(t), the selected action a, the reward value r, and the next - moment local state information O(t + 1) into its own experience pool. Finally, the two networks respectively extract experiences from their own experience pools to update the parameters.
[0066] A complete training round includes the whole process of the UAV taking off and completing the downlink transmission service. A training round contains multiple time slots, and the UAV dynamically adjusts its frequency - using and flight actions in each time slot. The whole process of the block double - network running in a time slot is as shown in Figure 4 . It is a schematic diagram of the time - slot structure.
[0067] The simulation analysis is as follows:
[0068] Figure 5 shows the change of the reward value obtained by the UAV using different algorithms with the number of training rounds when the external interference source releases sweep interference. In the sweep - interference mode, the interference source periodically attacks 4 available channels in the current environment in turn. The double - network block joint decision - making algorithm proposed in this paper has the best performance. It equips a dedicated reinforcement learning network for each type of action, enabling the UAV to capture the unique state information in various action - decision - making processes. At the same time, it has the fastest convergence speed in the multi - variable decision - making process, which is beneficial for the UAV to complete the deployment as soon as possible in the interference environment.
[0069] Figure 6It shows the variation of the reward values obtained by the UAV using different algorithms with the number of training rounds when the external interference source releases probability interference, and the number of ground users is 6. In the probability interference mode, the interference source determines the tendency to attack the channel in different time slots according to a specific probability matrix. Compared with the frequency scanning interference mode, the probability interference mode has stronger randomness, and it is more difficult for the UAV to learn its interference law. The performance of the algorithm proposed in this paper is still the best among all algorithms. Compared with the dual-network block decision algorithm and the single-network joint decision algorithm, the performance upper limit is increased by about 10%.
[0070] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0071] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A multi-UAV trajectory and communication frequency optimization method based on block reinforcement learning, characterized in that: Includes steps: (1) Establish a system model of UAV base stations serving ground users in a complex electromagnetic spectrum environment. The system model of UAV base stations serving ground users consists of UAVs, ground users, and interference sources. (2) Establish a channel model including K UAVs, M ground users, and P UAVs’ available communication frequency bands, and then obtain the ground user communication rate; (3) constructing system performance evaluation indicators and establishing user fairness weighted rate and maximization optimization problems based on the indicators; (4) The optimization problem is modeled as a distributed observable Markov decision process, and a block reinforcement learning method is used to train the UAV agent.
2. The multi-UAV trajectory and communication frequency optimization method based on block reinforcement learning according to claim 1 is characterized in that: In step (1), the drone base station serving ground users system model includes K drones, M ground users, P drones’ available communication frequency bands and 1 fixed interference source; all drones maintain a constant altitude flight, a single drone simultaneously serves multiple users through orthogonal frequency division multiplexing, there is no mutual interference between users, different drones with the same frequency will cause mutual interference to users not served by themselves, and the communication quality of all users served by drones with the same frequency as the interference source will be affected.
3. The multi-UAV trajectory and communication frequency optimization method based on block reinforcement learning according to claim 2 is characterized in that: In step (2), the channel model is as follows: Assume that the total duration of the drone’s flight and communication service is T, which contains E time slots of equal length, and the bandwidth of each frequency band is B tr , the communication rate of ground user m in time slot t is: Ra m (t) = B tr log2(1+SINR m (t)), where the signal to interference and noise ratio of ground user m is: P tr represents the transmission power of the drone, N represents the noise in the environment, and p k,m Indicates drone For ground users The service policy is expressed as The air-to-ground channel model between UAV k and ground user m is: Where T0=(4πf carr / c) 2 , f carr represents the carrier frequency, c represents the speed of light, d k,m represents the distance from drone k to user m, α represents the path loss constant, and P los and P nlos Respectively represent the probability values of line-of-sight link and non-line-of-sight link, λ los and λ nlos They represent the attenuation coefficients of line-of-sight links and non-line-of-sight links respectively. The mutual interference caused by other drones with the same frequency as the drone currently serving user m is: in, It indicates whether other drones j that do not serve user m are on the same frequency as drone k serving user m and cause mutual interference to them. Its form is the same as p k,m The same means that the interference of user m by the external interference source is: Among them, ch jam,m (t) represents the channel model between the jammer jam and the ground user m, which is in the same form as ch k,m (t) is the same, P jam Indicates the transmission power of the interference source, Indicates whether the UAV serving ground user m is on the same frequency as the interference source.
4. The multi-UAV trajectory and communication frequency optimization method based on block reinforcement learning according to claim 3 is characterized in that: In step (3), the system performance evaluation index is: in The optimization problem is: Among them, a freq,k Indicates the frequency selection of the drone. The drone can only select one frequency band in a time slot. traj,k Indicates the flight action of the drone.
5. The multi-UAV trajectory and communication frequency optimization method based on block reinforcement learning according to claim 1 is characterized in that: In step (4), the optimization problem is modeled as a distributed partially observable Markov decision process. The frequency network and trajectory network equipped by the drone have their own independent local state information, action space and reward function. During the training process of drone k, the agent contains the tuple: <local state information, action, reward>, and the global state information contains the observation information of all agents.
6. The multi-UAV trajectory and communication frequency optimization method based on block reinforcement learning according to claim 5 is characterized in that: The local status information of the frequency network consists of four parts: the position of the drone itself, the positions of other drones within the sensing range, the frequency band occupied by the current interference source, and the frequency bands occupied by other drones within the sensing range, expressed as: in, represents the position of drone k itself, represents the position of UAV j within the sensing range, f Jammer Represents the frequency band occupied by the current interference source. Represents the frequency band occupied by drone j within the sensing range. The frequency network action indicates that the drone can select an available frequency band from P frequency bands in each time slot. The frequency network reward indicates that each frequency network shares a global reward function: Rew freq =η system -pun1-pun2, where η system represents the fairness weighted rate sum of ground users, pun1 represents the penalty term for the UAV and the interference source to be in the same frequency, and pun2 represents the mutual interference penalty term between UAVs.
7. The multi-UAV trajectory and communication frequency optimization method based on block reinforcement learning according to claim 1 is characterized in that: Trajectory network local observation information O traj,k It consists of four parts: the position of the drone itself, the positions of other drones within the sensing range, the status of ground users, and the coordinates of obstacles, expressed as: in, Represents the state of ground user m, Loca barr Represents the coordinates of the obstacle, the ground user state S GU Including its own position, communication rate and service status, the flight action of the drone is defined as {forward, backward, fly right, fly left, hover}, and the trajectory network reward means that each frequency network shares a global reward function: Rew traj =η system -pun3-pun4, where pun3 is the penalty for collision between drones and obstacles, and pun4 is the penalty for collision between drones.
8. The multi-UAV trajectory and communication frequency optimization method based on block reinforcement learning according to claim 1 is characterized in that: The block reinforcement learning method is used to train the UAV agent as follows: two deep reinforcement learning networks are used to decide the frequency action and flight action of the UAV respectively. Assuming that the number of frequency bands that the UAV can choose is num freq The number of flight actions that can be selected is num traj In the block dual network structure, each network selects actions in its own action space. In the single network joint decision algorithm using the fusion action space, the UAV has a dimension of num freq ×num traj After receiving the perception information, each network outputs the corresponding frequency and flight action, updates the local state information and reward value of the frequency decision network, and then updates the local state information and reward value of the trajectory decision network. The user communication rate contained in the local state information and the user fairness weighted rate and contained in the reward function will also be affected by the transmission frequency band selected by the frequency decision network. Both networks share the global state information through the hybrid network in the QMix training framework. After updating the reward function and local state information, each network stores the local state information O(t), the selected action a, the reward value r and the local state information O(t+1) at the next moment into its own experience pool. Finally, the two networks extract experience from their own experience pools to update parameters.