Unmanned aerial vehicle spectrum anti-interference communication method and system based on mean field reinforcement learning
By applying an average field reinforcement learning method in the UAV communication network, optimizing the drone selection transmission channel, solving the problems of mutual disturbance between UAV users and jammer interference, and improving data throughput and communication quality.
Patent Information
- Application Number
- CN202411966727.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-06-03
AI Technical Summary
In drone communication networks, the prior art is difficult to ensure the drone user's own data throughput when the mutual interference between drone users and malicious interference between jammers exist at the same time, resulting in poor communication quality.
Adopting the anti-jamming communication method of drone spectrum based on average field reinforcement learning, through deep Q network and Boltzmann strategy, the transmission channel selection of the transmitting drone in the drone group is optimized to maximize its own profits and reduce interference.
It improves the data throughput of drone users, improves the communication quality of drone, and reduces communication overhead through distributed multi-agent reinforcement learning theory.
Smart Images

Figure CN120090739A_ABST
Abstract
Description
Background Art
[0002] In recent years, the anti-interference problem of UAV communication networks has attracted great attention. Compared with ground networks, UAVs have the advantages of high flexibility, fast deployment speed, and low cost, and can perform air-to-ground and air-to-air communications through line-of-sight (LoS) propagation. However, LoS links may be subject to more interference threats. On the one hand, there will be co-channel interference among UAV users; on the other hand, UAV communication is more vulnerable to attacks from malicious jammers.
[0003] Traditional anti-interference technologies, such as uncoordinated frequency hopping (UFH) and frequency hopping spread spectrum (FHSS), usually consume spectrum and have low efficiency in complex dynamic spectrum environments. Therefore, most traditional anti-interference methods are based on the assumption of prior knowledge of the jammer, which is difficult to obtain in practice. Reinforcement learning technology enables agents to find the best strategy in an unknown environment, so it has gradually begun to be applied to the field of UAV communication. However, with the in-depth research and the increasing complexity of application scenarios, the reinforcement learning of a single agent no longer meets the requirements. Recently, multi-agent reinforcement learning (MARL) has been applied to multi-UAV communication networks. Although some proposed multi-agent anti-interference schemes have better performance than traditional schemes, the competitive channel access problem among selfish UAV users in the case of anti-interference has not been discussed. In the case of the coexistence of mutual interference among UAV users and malicious interference from jammers, it is difficult to guarantee the data throughput of UAV users themselves, thus affecting the quality of UAV communication. Summary of the Invention
[0004] This specification provides a UAV spectrum anti-interference communication method based on mean-field reinforcement learning to solve the problem in the prior art that it is difficult to guarantee the data throughput of UAV users themselves in the case of the coexistence of mutual interference among UAV users and malicious interference from jammers, resulting in poor UAV communication quality.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] In the first aspect of this specification, a UAV spectrum anti-interference method based on mean-field reinforcement learning is provided, including:
[0007] Step S1, reset the network environment of the UAV swarm and initialize the target parameters of the transmitting UAVs in the UAV swarm;
[0008] Step S2: Obtain the environmental observation values of each transmitting UAV, and input the environmental observation values and the average action of the previous time slot into the evaluation network in the pre-constructed deep Q-network to determine the Q values of each transmitting UAV in the current time slot; combine the Q values, and through the Boltzmann strategy, obtain the target transmission channels corresponding to the optimal actions of each transmitting UAV in the current time slot.
[0009] Step S3: Each transmitting UAV communicates according to the target transmission channel. After the communication is completed, determine the environmental feedback reward of each transmitting UAV, update the environmental observation value at the same time, and calculate the average action of the current time slot.
[0010] Step S4: Combine the updated environmental observation value, the action of the next time slot, and the average action of the current time slot, and through the target network in the deep Q-network, obtain the target Q value; combine the target Q value, calculate the loss value through a preset loss function to update the evaluation network; the target network updates the network parameters according to the preset update step size based on the updated network parameters of the evaluation network.
[0011] Step S5: Determine whether the current time slot reaches the maximum time slot of the UAV swarm in this round. If not, return to Step S2; if so, continue to execute Step S6.
[0012] Step S6: Determine whether the preset loss function converges. If not, return to Step S1; if so, continue to execute Step S7.
[0013] Step S7: Obtain the trained evaluation network, and combine the trained evaluation network to enable each transmitting UAV in the UAV swarm to communicate according to the target transmission channel.
[0014] In some preferred embodiments, the network environment of the UAV swarm includes the channel model of the UAV, the movement model of the UAV, the wireless transmission model of the UAV, the channel model of the jammer, the movement model of the jammer, and the interference model of the jammer.
[0015] In some preferred embodiments, the target parameters of the transmitting UAV include the evaluation network parameters of the transmitting UAV, the target network parameters, the evaluation network learning rate, the update step size of the target network, the Boltzmann coefficient, and the average action.
[0016] In some preferred embodiments, the transmitting UAV determines the average action according to the communication channels selected by other transmitting UAVs around it:
[0017]
[0018] where N is the number of other transmitting UAVs within the interaction range of transmitting UAV i, an It is the communication channel for other transmitting UAVs within the interaction range of the transmitting UAV i.
[0019] In some preferred embodiments, it is characterized in that the environmental observation value The method for obtaining it includes:
[0020]
[0021] In the formula, It is used to characterize whether the communication is successful in the previous time slot. If the communication channel selected by the transmitting UAV in the data transmission stage is interfered by a jammer or occupied by other UAVs, it is a communication failure, represented by 0. If it is not interfered by a jammer or occupied by other UAVs, it is a communication success, represented by 1;
[0022] It is used to characterize a communication channel of the transmitting UAV in the current time slot. If the transmitting UAV detects that channel m is interfered by at least one jammer, then is set to 1. If it is not detected that channel m is interfered, then is set to 0.
[0023] In some preferred embodiments, the Boltzmann strategy is:
[0024]
[0025] In the formula, represents the set of all optional actions of the transmitting UAV i, including the set of all communication channels corresponding to the transmitting UAV.
[0026] In some preferred embodiments, the environmental feedback reward is used to characterize the actual throughput achieved by the transmitting UAV i at time slot t. The environmental feedback reward is defined as D i (t) represents the throughput and satisfies:
[0027]
[0028] In the formula, B is the bandwidth, P i is the transmission power of the transmitting UAV i, d i represents the distance from the transmitting UAV to the receiving UAV, I i (t) is the interference power of other UAVs, J i (t) is the interference power of the jammer, N i is the noise of the channel where the transmitting UAV i is located, T trans is the length of a single time slot; represents the maximum throughput that the transmitting UAV i can achieve.
[0029] In some preferred embodiments, the transmitting UAV updates the network parameters of the evaluation network by minimizing the loss function of the gradient descent algorithm The minimization loss function is defined as:
[0030]
[0031] In the formula, represents the evaluation network, represents the target network, o i and o i ′ respectively represent the adjacent environmental observation values before and after, o i and o i ′ follow the Markov decision process, represents the maximum future reward that the transmitting UAV i can obtain in the next time slot after selecting the channel a i r i represents the reward obtained by the transmitting UAV i in the current time slot. γ ∈ [0, 1] is the discount factor, which is used to indicate the degree of importance that the transmitting UAV attaches to future rewards. When γ is close to 1, it means that the transmitting UAV attaches importance to future rewards, and at this time, the training will tend to maximize long-term rewards; when γ is close to 0, it means that the transmitting UAV attaches importance to immediate rewards, and at this time, the training will tend to maximize short-term rewards.
[0032] In some preferred embodiments, the transmitting UAV updates the parameters of the evaluation network to the target network according to the preset update step τ.
[0033] In the second aspect of this specification, a UAV spectrum anti-jamming communication system based on mean-field reinforcement learning is provided, including:
[0034] An initialization module, which is used to reset the network environment of the UAV swarm and initialize the target parameters of the transmitting UAVs in the UAV swarm;
[0035] A data input module, which is used to obtain the environmental observation values of each transmitting UAV, and input the environmental observation values and the average action of the previous time slot into the evaluation network in the pre-constructed deep Q network to determine the Q value of each transmitting UAV in the current time slot; combining the Q value, through the Boltzmann strategy, obtain the target transmission channel corresponding to the optimal action of each transmitting UAV in the current time slot;
[0036] A data output module, which is used to determine the environmental feedback reward of each transmitting UAV after the communication is completed, update the environmental observation value at the same time, and calculate the average action of the current time slot;
[0037] The data calculation module combines the updated environmental observation value, the action of the next time slot, and the average action of the current time slot to obtain a target Q value through the target network in the deep Q network; combines the target Q value to calculate the loss value through a preset loss function to update the evaluation network; the target network updates the parameters according to the preset update step size based on the updated network parameters of the evaluation network;
[0038] A data iteration module, used to sequentially call the data input module, the data output module and the data calculation module until the current time slot reaches the maximum time slot of the drone group in this round, and when the maximum time slot is reached, sequentially call the initialization module, the data input module, the data output module and the data calculation module until the preset loss function converges;
[0039] The communication module is used to obtain the evaluation network trained by the above module, and combine the trained evaluation network to enable each transmitting end drone in the drone group to communicate according to the target transmission channel.
[0040] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects:
[0041] This method is based on the theory of distributed multi-agent reinforcement learning, which uses less communication overhead in exchange for improved throughput performance; the distributed multi-UAV anti-interference problem is expressed as a partially observable stochastic game (POSG), in which each transmitting UAV will focus on maximizing its own benefits; the mean field theory is introduced to convert POSG into a mean field game, which greatly reduces the complexity of solving the game problem; the Boltzmann strategy is used for decision-making, which increases the robustness of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0043] Figure 1 A flowchart of a UAV spectrum anti-interference communication method based on mean field reinforcement learning provided in an embodiment of this specification;
[0044] Figure 2 A schematic diagram of an anti-interference transmission model for UAV communication provided in an embodiment of this specification;
[0045] Figure 3 A convergence comparison diagram provided in one embodiment of this specification;
[0046] Figure 4A graph showing the variation of average reward with the number of U2U for an embodiment of this specification;
[0047] Figure 5 A graph showing the variation of average reward with the increase in the number of channels for an embodiment of this specification. Detailed implementation manners
[0048] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0049] The following will detail the technical solutions provided in each embodiment of this application with reference to the drawings.
[0050] Please refer to Figure 1 As shown in Figure 1 An anti-jamming communication method for UAV spectrum based on mean-field reinforcement learning is proposed in the first embodiment of this application. The method includes:
[0051] S1. Reset the network environment of the UAV swarm and initialize the target parameters of the transmitting UAVs in the UAV swarm;
[0052] S2. Obtain the environmental observation values of each transmitting UAV, and input the environmental observation values and the average action of the previous time slot into the evaluation network in a pre-constructed deep Q network to determine the Q values of each transmitting UAV in the current time slot; combine the Q values, and obtain the target transmission channels corresponding to the optimal actions of each transmitting UAV in the current time slot through the Boltzmann strategy;
[0053] S3. Each transmitting UAV communicates according to the target transmission channel. After the communication is completed, determine the environmental feedback rewards of each transmitting UAV, update the environmental observation values at the same time, and calculate the average action of the current time slot;
[0054] S4. Combine the updated environmental observation values, the actions of the next time slot, and the average action of the current time slot, and obtain the target Q values through the target network in the deep Q network; combine the target Q values, calculate the loss value through a preset loss function to update the evaluation network; the target network updates the network parameters according to the preset update step size based on the network parameters updated by the evaluation network;
[0055] Step S5. Determine whether the current time slot reaches the maximum time slot of the UAV swarm in this round. If not, return to step S2; if so, continue to execute step S6;
[0056] Step S6: Determine whether the preset loss function converges. If not, return to Step S1; if so, continue to execute Step S7;
[0057] Step S7: Obtain the trained evaluation network, and in combination with the trained evaluation network, enable each transmitting UAV in the UAV group to communicate according to the target transmission channel.
[0058] To better explain this method, in this embodiment, continuous time is first discretized into multiple time slots, and a positive integer t ∈ T = {1, 2,..., T} is used to represent the t-th time slot;
[0059] Meanwhile, assume that there are several UAV pairs in the network, represented by the set to represent.
[0060] Secondly, this embodiment solves the UAV communication interference problem based on the idea of mean field reinforcement learning, and specifically uses a Deep Q-Network (DQN) to implement it. This network mainly includes an Evaluation Network and a Target Network:
[0061] The Evaluation Network is used to estimate the Q-values of all possible actions in the current state and select the optimal action. The parameters of the evaluation network will be updated at each step by the gradient descent method, and the data in the Replay Buffer is used to calculate the loss and backpropagate for optimization.
[0062] The Target Network is used to generate the Target Q-value, which is used to calculate the loss function of the evaluation network. The parameters of the target network are not updated at each step, but are updated only every fixed number of steps (for example, every 1000 steps), specifically by copying the parameters of the evaluation network to the target network.
[0063] In this embodiment, the network environment of the above UAV group includes the channel model of the UAV, the movement model of the UAV, the wireless transmission model of the UAV, the channel model of the jammer, the movement model of the jammer, and the interference model of the jammer.
[0064] Among them, the target parameters of the transmitting UAV include the evaluation network parameters, target network parameters, evaluation network learning rate, update step of the target network, Boltzmann coefficient, and average action of the transmitting UAV.
[0065] In this embodiment, the transmitting UAV determines the average action according to the communication channels selected by other transmitting UAVs around it, satisfying:
[0066]
[0067] where N is the number of other transmitting UAVs within the interaction range of transmitting UAV i, and a n is the communication channel of other transmitting UAVs within the interaction range of transmitting UAV i.
[0068] In the above embodiment, the environmental observation value is used to characterize the communication result of the transmitting UAV in the previous time slot and the communication channel interfered with by the transmitting UAV in the current time slot satisfying:
[0069]
[0070] where is used to characterize whether the communication was successful in the previous time slot. If the communication channel selected by the transmitting UAV during the data transmission stage is interfered with by a jammer or occupied by other UAVs, it is a communication failure, represented by 0. If it is not interfered with by a jammer or occupied by other UAVs, it is a communication success, represented by 1;
[0071] The communication channel interfered with by the transmitting UAV detected in the current time slot satisfying:
[0072]
[0073] where is used to characterize a communication channel of the transmitting UAV in the current time slot. If the transmitting UAV detects that channel m is interfered with by at least one jammer, then is set to 1. If it is not detected that channel m is interfered with, then is set to 0.
[0074] In this embodiment, the transmitting UAV selects the transmission channel for the current time slot according to the output result of the evaluation network and the Boltzmann strategy. Assuming that the evaluation neural network environmental observation value o i and the average action then the Boltzmann strategy satisfies:
[0075]
[0076] where represents the set of all optional actions of transmitting UAV i, including the set of all current communication channels.
[0077] In this embodiment, the environmental feedback reward is used to characterize the actual throughput achieved by the transmitting UAV i at time slot t, and the environmental feedback reward is defined as D i (t) represents the throughput and satisfies:
[0078]
[0079] In the formula, B is the bandwidth, P i is the transmission power of the transmitting UAV i, d i represents the distance from the transmitting UAV to the receiving UAV, I i (t) is the interference power of other UAVs, J i (t) is the interference power of the jammer, N i is the noise of the channel where the transmitting UAV i is located, T trans is the length of a single time slot; represents the maximum throughput that the transmitting UAV i can achieve.
[0080] In this embodiment, the transmitting UAV updates the network parameters of the evaluation network through the loss function minimized by the gradient descent algorithm and the minimized loss function is defined as:
[0081]
[0082] In the formula, represents the evaluation network, represents the target network, o i and o i ′ respectively represent the adjacent environmental observation values before and after, o i and o i ′ follow the Markov decision process, represents the maximum future reward that the transmitting UAV i can obtain in the next time slot after selecting channel a i , r i represents the reward obtained by the transmitting UAV i in the current time slot, γ ∈ [0, 1] is the discount factor, which is used to indicate the degree of importance that the transmitting UAV attaches to future rewards. If γ is close to 1, it means that the transmitting UAV attaches importance to future rewards, and at this time the training will tend to maximize long-term rewards; if γ is close to 0, it means that the transmitting UAV attaches importance to immediate rewards, and at this time the training will tend to maximize short-term rewards.
[0083] In this embodiment, the transmitting UAV updates the parameters of the evaluation network to the target network according to the preset update step τ.
[0084] It is easy to understand that since reinforcement learning is based on the Markov decision process, the transition of the environmental state is continuous. Therefore, if a series of empirical data is directly obtained by sampling in chronological order, the neural network trained with these data is prone to overfitting (because the training samples are not independent).
[0085] Therefore, in this embodiment, an experience pool is designed to store past empirical data. By randomly sampling from the pool for training, it can not only prevent the overfitting of the neural network but also ensure that the samples are independent and identically distributed.
[0086] Specifically, in the above steps of this embodiment, each transmitting UAV packages "environmental observation value, transmission channel of the current time slot, reward feedback from the environment, new environmental observation value, average action of the current time slot", etc. into an empirical data and stores it in the experience pool. If the experience pool is full, the transmitting UAV first extracts multiple empirical data from the pool, updates the evaluation network by minimizing the loss function, and then updates the target network according to the step size. Specifically, the update process can refer to the update process of the above evaluation network and will not be elaborated here.
[0087] In this embodiment, the dual-network structure of the evaluation network plus the target network can prevent overfitting and the oscillation problem during the training process. The transmitting UAV as the intelligent agent copies the parameters of the evaluation network to the target network according to the update step size τ.
[0088] Furthermore, the second embodiment of this application proposes a simulation process for the above method to verify the effectiveness of the above method.
[0089] Refer to Figures 2 to 5 , in this embodiment, Python programming is used for simulation. The specific parameter settings do not affect the generality of the verification results. The algorithms used for comparison with the above method in this embodiment include: IQL algorithm and IQL-soft algorithm.
[0090] Figure 2It is a schematic diagram of the experimental environment. Specifically, there are 100 U2Us randomly distributed at any positions in the set space in the environment, and their positions change over time according to the Gaussian-Markov mobility model. There are a total of 50 available channels in the environment. They have the same bandwidth, and since only line-of-sight link conditions are considered, these channels are homogeneous for the UAVs. Additionally, there are 2 jammers flying in this space according to the Gaussian-Markov mobility model as well. They both adopt the Markov jamming mode, but their channel transition probability matrices are independent of each other. Therefore, it is possible for the two jammers to interfere on the same channel in the same time slot. During the simulation process, the communication task is continuously repeated 200 times, also known as 200 rounds. Each task contains 2000 time slots, and the main difference between each task lies in the different initial positions of the U2Us and the jammers. The main simulation parameters are as follows in the table:
[0091] Parameter Value UAV / jammer transmission power 23 dBm UAV / jammer antenna height 1.5m UAV / jammer antenna gain 3 dB UAV receiver noise figure 9 dB Number of UAV pairs 100 Number of jammers 2 Number of available channels 50 Duration of each time slot 1.18s Transmission time within the time slot 0.98s Jammer frequency sweeping time 2.28s Actual transmission rate of UAV pairs 16 bits / s / Hz Learning rate 0.01 Boltzmann coefficient 0.01 Discount factor 0.95 Target network update step size 100 Experience pool capacity 5000
[0092] From Figure 3 It can be seen that the proposed algorithm is faster than the three baseline algorithms in terms of convergence speed, and the reward value obtained after convergence is also higher than that of the other three algorithms. The algorithm proposed in the present invention is Soft-MFQ, and the algorithm named Soft-MFQproa is a variant of the Soft-MFQ algorithm. The advantage of Soft-MFQ is not only because of the adoption of the Boltzmann strategy and the soft update of the target network, but also because of considering the average actions of the transmitting UAVs among other surrounding U2Us. The concept of average action abstracts the interaction of the U2U's neighbors with it, and at the same time implicitly includes the influence of the entire network on this U2U. All these information helps the U2U's deep learning network to be better trained.
[0093] From Figure 4 It can be seen that the proposed algorithm can better reflect its advantageous role in larger-scale UAV communication networks. The change in the number of U2Us shown in the figure directly reflects the change in the scale of the UAV communication network. It can be seen that the four algorithms have similar performances when the number of U2Us is small, and as the number of U2Us increases, the advantage of the algorithm proposed in this chapter becomes gradually obvious.
[0094] In Figure 5 the experiment shown, there are 100 U2Us in the set environment, and the number of available channels increases from 20. It can be seen from the figure that when the number of channels cannot fully cover the number of U2Us, the growth trend of the average reward obtained by using the algorithm Soft-MFQ proposed in the present invention is between linear growth and exponential growth until the number of channels completely exceeds the number of U2Us; while the average reward of IQL grows slowly and the growth rate shows a downward trend, indicating that U2Us can better coordinate spectrum resources through average actions and avoid the occurrence of co-channel interference.
[0095] The third embodiment of this application proposes a UAV spectrum anti-interference communication system based on mean-field reinforcement learning, including:
[0096] An initialization module, configured to reset the network environment of the UAV swarm and initialize the target parameters of the transmitting UAVs in the UAV swarm;
[0097] A data input module, configured to obtain the environmental observation values of each of the transmitting UAVs, and input the environmental observation values and the average action of the previous time slot into the evaluation network in a pre-constructed deep Q network to determine the Q values of each transmitting UAV in the current time slot; combining the Q values, through the Boltzmann strategy, obtain the target transmission channels corresponding to the optimal actions of each transmitting UAV in the current time slot;
[0098] A data output module, configured to determine the environmental feedback rewards of each transmitting UAV after communication is completed, update the environmental observation values at the same time, and calculate the average action of the current time slot;
[0099] A data calculation module, combining the updated environmental observation values, the actions of the next time slot, and the average action of the current time slot, through the target network in the deep Q network, obtain the target Q values; combining the target Q values, calculate the loss value through a preset loss function to update the evaluation network; the target network updates the network parameters according to a preset update step size based on the updated network parameters of the evaluation network;
[0100] A data iteration module, configured to sequentially call the data input module, the data output module, and the data calculation module until the current time slot reaches the maximum time slot of the UAV swarm in this round, and when the maximum time slot is reached, sequentially call the initialization module, the data input module, the data output module, and the data calculation module until the preset loss function converges;
[0101] A communication module, configured to obtain the trained evaluation network of the above modules, and combine the trained evaluation network to enable each transmitting UAV in the UAV swarm to communicate according to the target transmission channels.
[0102] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and related descriptions of each module in the above-described system can refer to the corresponding processes in the foregoing method examples, and will not be repeated here.
[0103] The above are only the embodiments of this application and are not intended to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.
Claims
1. A UAV spectrum anti-interference communication method based on mean field reinforcement learning, characterized in that: include: Step S1, resetting the network environment of the drone swarm and initializing the target parameters of the transmitting drone in the drone swarm; Step S2, obtaining the environmental observation value of each transmitting drone, and inputting the environmental observation value and the average action of the previous time slot into the evaluation network in the pre-constructed deep Q network to determine the Q value of each transmitting drone in the current time slot; combining the Q value, and using the Boltzmann strategy, obtaining the target transmission channel corresponding to the optimal action of each transmitting drone in the current time slot; Step S3, each transmitting end UAV communicates according to the target transmission channel. After the communication is completed, the environmental feedback reward of each transmitting end UAV is determined, the environmental observation value is updated, and the average action of the current time slot is calculated; Step S4, combining the updated environmental observation value, the action of the next time slot, and the average action of the current time slot, and obtaining a target Q value through the target network in the deep Q network; combining the target Q value, calculating the loss value through a preset loss function to update the evaluation network; the target network updates the parameters according to the preset update step size based on the updated network parameters of the evaluation network; Step S5, determining whether the current time slot reaches the maximum time slot of the drone group in this round, if not, returning to step S2; if yes, continuing to step S6; Step S6, determining whether the preset loss function converges, if not, returning to step S1; if yes, continuing to step S7; Step S7, obtaining a trained evaluation network, and combining the trained evaluation network to enable each transmitting end drone in the drone group to communicate according to the target transmission channel.
2. The UAV spectrum anti-interference communication method based on mean field reinforcement learning according to claim 1 is characterized in that: The network environment of the drone swarm includes a channel model of the drone, a mobile model of the drone, a wireless transmission model of the drone, a channel model of the jammer, a mobile model of the jammer, and an interference model of the jammer.
3. The UAV spectrum anti-interference communication method based on mean field reinforcement learning according to claim 1 is characterized in that: The target parameters of the transmitting UAV include evaluation network parameters, target network parameters, evaluation network learning rate, update step size of the target network, Boltzmann coefficient and average action of the transmitting UAV.
4. The UAV spectrum anti-interference communication method based on mean field reinforcement learning according to claim 1 is characterized in that: The transmitting end drone determines the average action according to the communication channels selected by other transmitting end drones around it: Where N is the number of other transmitting end drones within the interaction range of transmitting end drone i, a n It is the communication channel of other transmitting end UAVs within the interaction range of transmitting end UAV i.
5. The UAV spectrum anti-interference communication method based on mean field reinforcement learning according to claim 1 is characterized in that: The environmental observation value The methods for obtaining it include: In the formula, It is used to indicate whether the communication in the previous time slot is successful. If the communication channel selected by the transmitting drone is interfered by the jammer or occupied by other drones during the data transmission phase, the communication fails and is represented by 0. If it is not interfered by the jammer or occupied by other drones, the communication is successful and is represented by 1. It is used to characterize a communication channel of the transmitting drone in the current time slot. If the transmitting drone detects that the channel m is interfered by at least one jammer, Set to 1, if channel m is not detected to be interfered, Set to 0.
6. The UAV spectrum anti-interference communication method based on mean field reinforcement learning according to claim 1 is characterized in that: The Boltzmann strategy is: In the formula, A i Represents the set of all optional actions of the transmitting UAV i, including the set of all communication channels corresponding to the transmitting UAV.
7. The UAV spectrum anti-interference communication method based on mean field reinforcement learning according to claim 1 is characterized in that: The environmental feedback reward is used to characterize the actual throughput achieved by the transmitting UAV i at time slot t, and the environmental feedback reward is defined as D i (t) represents the throughput, which satisfies: Where B is the bandwidth, P i is the transmitting power of UAV i at the transmitting end, d i Indicates the distance from the transmitting drone to the receiving drone, I i (t) is the interference power of other UAVs, J i (t) is the jammer's interference power, N i is the noise of the channel where the transmitting drone i is located, T trans is the length of a single time slot; It represents the maximum throughput that can be achieved by the transmitting UAV i.
8. The UAV spectrum anti-interference communication method based on mean field reinforcement learning according to claim 1 is characterized in that: The transmitting drone evaluates the network parameters of the network by minimizing the loss function of the gradient descent algorithm. To update, the minimization loss function is defined as: In the formula, represents the evaluation network, represents the target network, o i and i ′ respectively represent the two adjacent environmental observation values, o i and i ′ obeys the Markov decision process, Indicates that the transmitting drone i selects channel a i After that, the maximum future reward that can be obtained in the next time slot is r i represents the reward obtained by the transmitting UAV i in the current time slot, γ∈[0,1] is the discount factor, which is used to indicate the importance of the transmitting UAV to future rewards. If γ is close to 1, it means that the transmitting UAV attaches importance to future rewards. At this time, the training will tend to maximize long-term rewards; if γ is close to 0, it means that the transmitting UAV attaches importance to immediate rewards. At this time, the training will tend to maximize short-term rewards.
9. The UAV spectrum anti-interference communication method based on mean field reinforcement learning according to claim 1 is characterized in that: The transmitting drone updates the parameters of the evaluation network to the target network according to a preset update step size τ.
10. A UAV spectrum anti-interference communication system based on mean field reinforcement learning, characterized in that: include: An initialization module, used to reset the network environment of the drone swarm and initialize the target parameters of the transmitting drone in the drone swarm; A data input module is used to obtain the environmental observation value of each transmitting end drone, and input the environmental observation value and the average action of the previous time slot into the evaluation network in the pre-constructed deep Q network to determine the Q value of each transmitting end drone in the current time slot; combined with the Q value, through the Boltzmann strategy, to obtain the target transmission channel corresponding to the optimal action of each transmitting end drone in the current time slot; The data output module is used to determine the environmental feedback reward of each transmitting drone after the communication is completed, update the environmental observation value, and calculate the average action of the current time slot; The data calculation module combines the updated environmental observation value, the action of the next time slot, and the average action of the current time slot to obtain a target Q value through the target network in the deep Q network; combines the target Q value to calculate the loss value through a preset loss function to update the evaluation network; the target network updates the parameters according to the preset update step size based on the updated network parameters of the evaluation network; A data iteration module, used to sequentially call the data input module, the data output module and the data calculation module until the current time slot reaches the maximum time slot of the drone group in this round, and when the maximum time slot is reached, sequentially call the initialization module, the data input module, the data output module and the data calculation module until the preset loss function converges; The communication module is used to obtain the evaluation network trained by the above module, and combine the trained evaluation network to enable each transmitting end drone in the drone group to communicate according to the target transmission channel.