Interference frequency domain resource scheduling method and device for multi-unmanned aerial vehicle communication
Through an improved multi-agent reinforcement learning algorithm, a distributed frequency domain blocking jamming model and a partially observable Markov game process are established to optimize the credit allocation strategy of multiple jammers. This solves the problems of insufficient robustness of the jamming strategy and low efficiency of reward allocation in the multi-UAV frequency hopping communication system, and achieves efficient jamming frequency domain resource scheduling.
Patent Information
- Application Number
- CN202510777237.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-09
AI Technical Summary
When facing a multi-UAV frequency-hopping communication system, the existing single-agent method has difficulty in fully perceiving and responding to the complex electromagnetic environment, resulting in insufficient robustness of the interference strategy, low credibility of reward allocation, and poor allocation efficiency.
An improved multi-agent reinforcement learning algorithm is used to establish a distributed frequency-domain blocking jamming model. Multiple jammers are regarded as a group of agents, and a partially observable Markov game process is constructed. The jamming frequency-domain resource scheduling strategy is optimized through an improved counterfactual multi-agent policy gradient algorithm and a finite-length elite trajectory experience replay mechanism.
The interference efficiency of multiple jammers against multiple UAV frequency hopping communication systems is improved, the credibility and efficiency of credit allocation are enhanced, and fast and effective interference is achieved.
Smart Images

Figure CN120614022A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication countermeasure technology, and specifically to a method and device for scheduling interference frequency domain resources for multi-UAV communications. Background Art
[0002] In communication confrontation, the selection of jamming strategies is a key link. By optimizing the jamming strategy, jamming resources can be effectively saved while significantly improving the success rate of jamming. Formulating a scientific and reasonable jamming strategy is of great significance to improving the effectiveness of the confrontation. Among the many anti-interference communication technologies, frequency hopping communication has the widest application range. When jamming against the frequency hopping communication of multiple drones, the jammer faces multiple technical challenges. Frequency hopping communication shortens the dwell time of a single frequency by switching carrier frequencies at a high rate, making it difficult for traditional tracking jammers to capture the target frequency in a timely manner. Although traditional blocking jammers can cover a wide frequency band, they require extremely high power to suppress the entire frequency hopping range, which easily exposes the interference source and consumes huge energy, resulting in poor interference flexibility. In addition, drones can dynamically avoid the interference frequency band through spectrum sensing, or use distributed collaboration to adjust the frequency hopping pattern in real time to further compress the interference window.
[0003] With the development of artificial intelligence (AI), research on communication countermeasures has achieved significant breakthroughs after incorporating AI into the technology. Reinforcement learning is a machine learning theory that doesn't rely on prior knowledge. Its core principle is that an intelligent agent learns through dynamic interaction with its environment to maximize its benefits. This approach has broad application potential in a variety of fields, including intelligent decision-making and control, and resource allocation. It is particularly effective in scenarios that require rapid adaptation to complex environments and dynamic decision-making. Reinforcement learning methods are widely used in the core intelligent decision-making modules of communication countermeasures systems, providing efficient and accurate decision-making support for jamming, significantly improving countermeasure effectiveness and response speed.
[0004] Existing blocking jamming methods incorporating reinforcement learning primarily focus on scenarios where a single agent counters a single frequency-hopping communication. In complex electromagnetic environments, the frequency-hopping communication systems of swarm drones are not singular. Single-agent approaches struggle to fully perceive and respond to these complex frequency-hopping communications, resulting in insufficient robustness in jamming strategies. Existing reinforcement learning-based multi-jammer technologies for countering multi-UAV frequency-hopping communication systems suffer from low reliability and inefficient reward distribution. Summary of the Invention
[0005] In response to the problems mentioned in the background technology, the present application provides a method and device for scheduling interference frequency domain resources for multi-UAV communications. By utilizing an improved multi-agent reinforcement learning algorithm, the coordination of multiple jammers can be achieved to quickly and effectively improve the credibility of credit allocation, and improve the efficiency of credit allocation, thereby realizing intelligent scheduling of interference frequency domain resources and achieving rapid and effective interference with the frequency-hopping communications of a group of UAVs.
[0006] This application is implemented through the following technical solutions: A method for scheduling interference frequency domain resources for multi-UAV communications, comprising: Based on the interference-to-signal ratio criterion and combined with the multi-jammer frequency-hopping communication system scenario for multiple UAVs, a distributed frequency domain blocking jamming model is established; Multiple jammers are regarded as a group of intelligent agents, and the process of jamming multiple UAV communications is constructed as a partially observable Markov game process. Based on the credit distribution criterion, a credit distribution model for multiple agents is established. The improved counterfactual multi-agent policy gradient algorithm and the finite-length elite trajectory experience replay mechanism are used to optimize the credit allocation strategy in the process of multiple jammers interfering with multiple UAVs, and the optimal interference frequency domain resource scheduling strategy is obtained and output.
[0007] In some embodiments, establishing a distributed frequency-domain blocking interference model includes: When the interference bandwidth coverage is greater than or equal to the UAV K If the interference covers more than one-third of the target frequency points, the interference is effective; otherwise, the interference is invalid. Among them, drones K The value acquisition methods include: Calculating a first ratio of the jammer transmission power to the signal transmitter power; Calculating a second ratio of the product of the jammer transmitting antenna gain and the signal receiving antenna gain to the product of the signal transmitter antenna gain and the signal receiving antenna gain; calculating a third ratio of the square value of the transmission distance of the communication signal to the square value of the transmission distance of the jammer signal; calculating the product of the first ratio, the second ratio, and the third ratio; Calculate the ratio of the suppression coefficient to the product to obtain the value.
[0008] In some implementations, the interference bandwidth coverage is obtained by: Calculate the signal bandwidth of all target drones in the interference airspace being jammed by a certain jammer; Calculating a ratio of the signal bandwidth to the interference bandwidth of the jammer; Repeat the above process to obtain and accumulate the ratios corresponding to all jammers to obtain the interference bandwidth coverage.
[0009] In some embodiments, the process of interfering with multi-UAV communications is constructed as a partially observable Markov game process, and a multi-agent trust allocation model is established based on a trust allocation criterion, including: The partially observable Markov game is described as a tuple, which includes the system state observed by the jammer group, the joint action of the jammer group, the transition probability of the system state, the joint reward obtained by the jammer group after performing the joint action, the number of jammer agents, the independent observation space of each jammer agent, and the transition probability of the observation space. The system state space is composed of the individual states of all jammer agents and the global environment state; The joint action space consists of the action spaces of all jammer agents; The joint reward function is expressed as the sum of the independent action rewards of all jammer agents and the interaction rewards between all jammer agents; Define the credit distribution function: when performing a joint action under a certain system state, the contribution weight of the action of a single jammer agent to the joint reward, and the sum of the contribution weights of the actions of all jammer agents to the joint reward is equal to 1.
[0010] In some embodiments, the use of an improved counterfactual multi-agent policy gradient algorithm and a finite-length elite trajectory experience replay mechanism to optimize the credit allocation strategy during the multi-jammer jamming multi-UAV process includes: Establish a common trajectory experience pool and an elite trajectory experience pool, initialize the sampling ratio coefficient, initialize the reward threshold, set the storage trajectory length, set the update step size, and set the soft update hyperparameters; Initialize the Critic value evaluation network, Critic target network and N Actor networks, where N is the number of jammers; initialize the maximum rounds, the compensation of each round, the greedy parameter and the greedy attenuation parameter; Get the system status at the current moment of the current round, as well as the actions of all agents at the previous moment; For the current agent, if the greed parameter is greater than the random value, the action of the current agent at the current moment is randomly obtained according to the greedy strategy; otherwise, the observation space of the current agent, the unique hot encoding of the current agent, and the joint action of other agents except itself at the previous moment are obtained, and input into the Actor network to obtain the output, thereby obtaining the action of the current agent at the current moment; repeat this step until all agents are processed and the actions of all agents at the current moment are obtained; According to the current actions of all agents, the total reward is obtained and the system state at the next moment is obtained; Save all information; the saved information includes: the system state at the current moment, the system state at the next moment, the observation space of all agents, the actions of all agents at the current moment, the total reward, and the actions of all agents at the previous moment; If the current moment can divide the stored trajectory length, the information stored in the historical stored trajectory length is pushed into the ordinary trajectory experience pool; if the total reward of a certain step in the historical stored trajectory length is greater than the reward threshold, the information stored in the historical stored trajectory length is pushed into the elite trajectory experience pool; If the current moment can divide the update step size and the number of samples in the two experience pools is sufficient, sampling is performed, and the loss value of the Critic network is calculated, and the advantage function is obtained, and then N The loss value of the Actor network; Update the Critic target network; update the sampling ratio coefficient, update the reward threshold, and update the greed parameter; Repeat the above process to optimize the next moment until the current round of optimization is completed; Repeat the above process to perform the next round of optimization until the maximum round of optimization is completed and the optimization results are output.
[0011] In some embodiments, the critic value evaluation network and the critic target network have the same network architecture and the same initial parameter configuration. The critic value evaluation network is used to calculate the estimated value of the agent's current state in real time, and use this as a basis to guide the action selection strategy; the critic target network is responsible for generating the target value, providing a stable parameter signal for the training process. The update of the critic target network adopts the soft update method, and the update formula is expressed as: ; in, is a soft update hyperparameter, is the weight parameter of the Critic value evaluation network; is the weight parameter of the Critic target network.
[0012] In some embodiments, the loss function of the critic value evaluation network and the critic target network is expressed as: ; ; in, represents the state-action value function, Indicates that the parameter is The state-action value function, represents the weighted sum of the steps, is the discount factor, For the immediate reward at the next moment, is the trajectory length, is the parameter vector of the Critic network; is the regularization hyperparameter; is the number of training samples; The loss function of the Actor network is expressed as: ; ; ; in, Indicates the policy parameters Find the gradient, Indicates that the Actor network output is logarithmic, Representing an agent Current strategy, Representing an agent In the current strategy The following actions are taken, is the Actor network output, Representing an agent In the current strategy Actions taken The probability of Indicates that the status Downsampling Action Combination The state-action pair value, is the parameter vector of the Actor network, Indicates the current strategy The entropy regularization term under .
[0013] In some embodiments, the total reward consists of positive rewards and negative rewards; wherein, the positive reward is calculated by accumulating the positive rewards of all jammers achieving effective interference at the current moment; the negative reward is calculated by accumulating the negative rewards of all jammers failing to cause effective interference to the drone at the current moment.
[0014] In some embodiments, the size of the reward threshold increases as the number of rounds increases; The sampling ratio coefficient decreases as the number of rounds increases; The number of samples obtained by the sampling is equal to the product of the number of samples in the ordinary trajectory experience pool and the sampling ratio coefficient plus the product of the number of samples in the elite trajectory experience pool and (1-sampling ratio coefficient).
[0015] On the other hand, this application also proposes an interference frequency domain resource scheduling device for multi-UAV communication, including: The first modeling unit, based on the interference-to-signal ratio criterion and combined with the multi-jammer multi-UAV frequency hopping communication system scenario, establishes a distributed frequency domain blocking jamming model; The second modeling unit is used to regard multiple jammers as a group of intelligent agents, construct the process of jamming multiple UAV communications as a partially observable Markov game process, and establish a multi-agent trust distribution model based on the trust distribution criterion; In addition, the optimization unit uses the improved counterfactual multi-agent policy gradient algorithm and the finite-length elite trajectory experience replay mechanism to optimize the trust allocation strategy in the process of multiple jammers interfering with multiple UAVs, and obtains and outputs the optimal interference frequency domain resource scheduling strategy.
[0016] This application proposes a method for scheduling interference frequency domain resources for multi-UAV communications. First, a distributed frequency domain blocking interference model of multi-interference machines facing multi-UAV frequency hopping communications based on the interference-to-signal ratio criterion is established; then, based on the distributed frequency domain blocking interference model, multiple jammers are regarded as a group of intelligent agents, and the process of interfering with multi-UAV communications is constructed as a partially observable Markov game process, and based on the credit allocation criterion, a credit allocation model of multiple agents is built; finally, the improved counterfactual multi-agent policy gradient algorithm and the finite-length elite trajectory experience replay mechanism are used to optimize the credit allocation strategy in the process of multi-jammers against multi-UAVs, realize reasonable and effective interference frequency domain resource scheduling, effectively improve the efficiency of credit allocation, solve the problems of low credit allocation and poor allocation efficiency of the previous multi-jammer counter-multi-UAV frequency hopping communication system, and effectively improve the interference efficiency of the multi-jammer counter-multi-UAV frequency hopping communication system; Correspondingly, the interference frequency domain resource scheduling device for multi-UAV communication proposed in this application also has the same technical effects as mentioned above. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the embodiments of the present application, constitute a part of the present application, and do not constitute a limitation of the embodiments of the present application. In the drawings: Figure 1 This is a flow chart of the interference frequency domain resource scheduling method proposed in an embodiment of the present application; Figure 2 This is a diagram of a classic scenario of multiple jammers versus multiple drones; Figure 3 for Figure 2 Schematic diagram of the frequency hopping point distribution of each drone in the shown scenario; Figure 4 for Figure 2 Schematic diagram of each jammer selecting different locations to deploy the jamming strip in the scenario shown; Figure 5 This is a schematic diagram of the overall structure of COMA used in the embodiment of this application; Figure 6This is a schematic diagram of the Critic network and Actor network structure used in the embodiment of this application; Figure 7 This is a schematic diagram of the limited-length elite trajectory experience replay mechanism used in the embodiment of this application; Figure 8 This is a functional block diagram of the interference frequency domain resource scheduling device proposed in an embodiment of the present application; Figure 9 This is a schematic diagram of the system architecture of the interference frequency domain resource scheduling device proposed in an embodiment of the present application; Figure 10 A schematic diagram of an electronic device proposed in an embodiment of the present application; Figure 11 A schematic diagram of a computer-readable storage medium proposed in an embodiment of the present application; Reference numerals and corresponding component names: 200-interference frequency domain resource scheduling device, 201-first modeling unit, 202-second modeling unit, 203-optimization unit, 300-interference frequency domain resource scheduling system, 301-input device, 302-output device, 303-processor A, 304-memory A, 400-electronic device, 410-memory B, 420-processor B, 411-computer program A, 500-computer readable storage medium, 511-computer program B. DETAILED DESCRIPTION
[0018] Hereinafter, the terms "include" or "may include" as used in various embodiments of the present application indicate the presence of an invented function, operation, or element, and do not limit the addition of one or more functions, operations, or elements. In addition, as used in various embodiments of the present application, the terms "include," "have," and their cognates are intended only to indicate specific features, numbers, steps, operations, elements, components, or combinations of the foregoing, and should not be understood as excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing or the possibility of adding one or more features, numbers, steps, operations, elements, components, or combinations of the foregoing.
[0019] In various embodiments of the present application, the expression "or" or "at least one of A or / and B" includes any or all combinations of the words listed simultaneously. For example, the expression "A or B" or "at least one of A or / and B" may include A, may include B, or may include both A and B.
[0020] The expressions (such as "first", "second", etc.) used in the various embodiments of the present application may modify the various constituent elements in the various embodiments, but may not limit the corresponding constituent elements. For example, the above expressions do not limit the order and / or importance of the elements. The above expressions are only used to distinguish one element from other elements. For example, a first user device and a second user device indicate different user devices, although both are user devices. For example, without departing from the scope of the various embodiments of the present application, a first element may be referred to as a second element, and similarly, a second element may also be referred to as a first element.
[0021] It should be noted that when a component is described as being “connected” to another component, the first component may be directly connected to the second component, and a third component may be “connected” between the first and second components. Conversely, when a component is described as being “directly connected” to another component, it can be understood that there is no third component between the first and second components.
[0022] The terms used in the various embodiments of the application are only used to describe the purpose of specific embodiments and are not intended to limit the various embodiments of the application. As used herein, the singular form is intended to also include the plural form, unless the context clearly indicates otherwise. Unless otherwise limited, all terms used here (including technical terms and scientific terms) have the same meaning as the meaning generally understood by those of ordinary skill in the art of the application. The terms (such as the terms defined in the dictionary generally used) will be interpreted as having the same meaning as the contextual meaning in the relevant technical field and will not be interpreted as having idealized meaning or too formal meaning, unless clearly defined in the various embodiments of the application.
[0023] In order to make the objectives, technical solutions and advantages of this application more clear, the present application is further described in detail below in conjunction with examples and drawings. The schematic implementation methods of this application and their descriptions are only used to explain this application and are not intended to limit this application.
[0024] Existing research on blocking jamming methods combined with reinforcement learning has mostly focused on scenarios where a single agent is countering frequency-hopping communications, whereas actual communication environments may involve multiple agents working together. Single-agent methods struggle to effectively coordinate and optimize jamming strategies in multi-agent scenarios, resulting in a decrease in overall jamming effectiveness. Existing methods using multiple jammers to counter multi-UAV frequency-hopping communication systems suffer from low reward allocation reliability and poor allocation efficiency. To address this, the present application proposes a method for scheduling jamming frequency domain resources for multi-UAV communications.
[0025] like Figure 1As shown, the interference frequency domain resource scheduling method proposed in this embodiment includes the following steps: Step 110: Based on the interference-to-signal ratio criterion and in combination with the multi-jammer multi-UAV frequency hopping communication system scenario, a distributed frequency domain blocking jamming model is established; Step 120: Consider multiple jammers as a group of intelligent agents, construct the process of jamming multiple UAV communications as a partially observable Markov game process, and establish a multi-agent trust distribution model based on the trust distribution criterion; Step 130, using the improved counterfactual multi-agent policy gradient algorithm and the finite-length elite trajectory experience replay mechanism to optimize the credit allocation strategy in the process of multiple jammers interfering with multiple UAVs, and obtain and output the optimal interference frequency domain resource scheduling strategy.
[0026] Furthermore, in step 110 of the embodiment of the present application, blocking interference is achieved by implementing suppressive interference within a specific frequency band. As long as the frequency band contains the target frequency and the interference power reaches the required jamming-to-signal ratio (JSR) condition, an effective jamming effect can be achieved. This jamming method mainly relies on comprehensive coverage of the target frequency band and sufficient power output to ensure that the jamming signal can drown out the target signal, thereby achieving the purpose of jamming. Ignoring the polarization loss caused by the different transmitting and receiving antennas, the jamming-to-signal ratio calculation method can be expressed using formula (1).
[0027] (1) in, is the jammer transmission power, is the signal transmitter power; is the product of the jammer's transmitting antenna gain and the signal receiving antenna gain, It is the product of the signal transmitter antenna gain and the signal receiving antenna gain; is the spatial loss of jammer signal transmission, is the spatial loss of communication signal transmission. The signal transmission loss is considered as free space path loss (FSPL), and its calculation method can be expressed by formula (2).
[0028] (2) in, is the spatial transmission loss, is the wavelength, is the frequency, is the propagation distance. The spatial loss of jammer signal transmission and the spatial loss of communication signal transmission can be expressed by the above formula (2).
[0029] Substituting Equation (2) into Equation (1), we can obtain the general calculation method of the interference-to-signal ratio, as shown in Equation (3).
[0030] (3) in, is the transmission distance of the communication signal, is the transmission distance of the jammer signal.
[0031] like Figure 2 The figure shows a classic scenario of multiple jammers against multiple drones. Multiple communication jammers are deployed in a small area with the same jamming airspace. Reconnaissance reveals that there are multiple drones in the jamming airspace and that they communicate using frequency hopping communication. The frequency hopping points used by each drone are not completely independent.
[0032] Due to limited conditions, the reconnaissance and perception aircraft can only obtain the number of drones, their frequency hopping range, and their frequency hopping distribution. Each drone is labeled, for example, Drone 1, Drone 2, Drone 3, and Drone 4. The frequency hopping range is primarily between 0 and 100 MHz, and the communication signal carrier is 2.4 GHz.
[0033] The frequency hopping points of each drone are distributed as follows: Figure 3 Among them, the number of frequency hopping points of drones 4, 3, 2, and 1 is 64, some of which are shared by multiple drones (i.e., the frequencies are shared but time-independent), and some are independent.
[0034] All jammers use broadband blocking jamming. The spectrum components within each spectrum bandwidth are evenly distributed, and the transmission power of each jammer is consistent. The bandwidth of each jammer is fixed and the same. Assume that there are several drones and several jammers in the jamming airspace. The jamming effect is mainly calculated by the interference-to-signal ratio calculation formula. The interference-to-signal ratio can be calculated according to formula (4). When the interference-to-signal ratio exceeds the suppression coefficient , and the interference coverage target frequency is concentrated Only when the frequency points exceed one-third can the jamming measures be effective and successfully block the drone's communication link.
[0035] (4) in, Indicates the number of target drones for interference; Indicates the number of jammers; Indicates the effective jamming power of each jammer on the UAV; Indicates the signal bandwidth of the target being interfered by a certain jammer, Indicates the bandwidth of each frequency hopping point, For the The interference bandwidth of each jammer; The indicator value indicating whether the interference frequency band and the target frequency point are aligned in the frequency domain is expressed by formula (5). When the frequency is If the frequency hopping frequency is within the interference bandwidth, the indication value is 1, otherwise it is 0.
[0036] (5) make , simplifying formula (4), we get formula (6).
[0037] (6) in, Indicates the interference bandwidth coverage.
[0038] Regardless of the form of interference strategy, such as tracking interference, blocking interference, etc., the interference wave propagates over a certain distance and the frequency and power selection of the receiving antenna must be greater than a certain value, so that the interference will be effective. The product of is a fixed value in this scenario, so as long as , it can be considered that the interference bandwidth coverage meets the standard. In this scenario, the four drones The values are 、 、 and The above conditions are combined with the interference coverage target frequency concentration The condition of more than one-third of the frequency points can be obtained as formula (7).
[0039] (7) like Figure 4 The figure shows the jammers selecting different locations to place blocking interference bands. In the 0-100MHz range, the solid red line indicates the situation where a jammer places a blocking interference band at a specific location. In this area, the frequencies are densely arranged, one frequency band contains frequency hopping frequencies of multiple targets, and the frequencies of different targets are intertwined. In this case, selecting different locations to place blocking interference bands will have a significant impact on the allocation of interference frequency domain resources and the overall interference effect. All target frequencies are integrated into a whole for interference planning. By selecting frequency bands containing multiple different targets to implement interference, multiple targets can be effectively interfered with at the same time. This strategy helps to reduce the number of jammers required, reduce the bandwidth required for interference, and achieve optimal allocation of interference frequency domain resources.
[0040] Furthermore, in step 120 of the embodiment of the present application, the partially observable Markov game takes into account the situation in many real-world scenarios where the agent cannot fully observe the entire environment state, which is a further extension of the Markov game. In this case, each jammer agent can only receive a partial observation related to the environment state, which increases the uncertainty in the decision-making process and forces the agent to make decisions based on limited information. The partially observable Markov game can be described as a tuple ,in represents the system state observed by the interference crew; Indicates the joint action of the jammer crew; Represents the transition probability of the system state; Indicates that the jammer crew is performing a joint action Joint rewards obtained after represents the number of jammer agents; It represents the independent observation space of each jammer agent, which is a subset of the system state. For a fully observable environment, ; represents the observation space transition probability.
[0041] Credit assignment is a critical issue in cooperative multi-agent reinforcement learning scenarios. Because all agents share a global reward signal, determining each agent's contribution to the overall task is difficult, leading to challenges in credit assignment. In cooperative scenarios, the core of credit assignment is to quantify the contribution of individual actions to the collective goal, coordinate the strategic choices of agents or participants, and maximize global returns. Compared to game-based scenarios, cooperative credit assignment emphasizes collaborative consistency, goal alignment, and the accumulation of trust through dynamic interactions.
[0042] Considered by A cooperative system composed of intelligent agents, each of which has independent decision-making capabilities and its state is determined by the individual state. and the global environment state Together, the system state space is .in, Describe the characteristics of the external environment, Contains agents internal state.
[0043] Agent Action space It is defined as the set of actions it can perform, allowing discrete actions (or continuous actions. The joint action space of the system is the Cartesian product , any joint action Represents the action combination of all agents.
[0044] The core feature of a cooperative system is the existence of a globally consistent joint reward function , whose output value reflects the contribution of all agent actions to the collective goal. The joint reward function can be decomposed into the synergistic effect of individual actions, as shown in Equation (8).
[0045] (8) in, For intelligent agents Independent action rewards, For intelligent agents i With the agent j The interaction reward reflects the collaborative gain or conflict loss of the action. This function satisfies the non-negativity ( indicating effective cooperation) and additivity (allowing the decomposition of individual and interactive contributions) are the fundamental signals for credit assignment.
[0046] Define the credit assignment function , its physical meaning is: in the state Perform joint actions When the agent Action Joint Rewards The contribution weight of . Mathematically, it satisfies strict normalization conditions, as shown in formula (9).
[0047] (9) This condition ensures that the beliefs form a probability distribution that fully distributes the global reward to the individual actions, avoiding "building leaks" or "double distribution".
[0048] when When , it is a completely dominant scenario. Action When it is the only determinant of the joint reward (e.g., in a single-point control task, only the agent's actions affect the goal completion), the trust function takes the value , the trustworthiness of other agents For example, in a single-robot manipulation task, the precise grasping action of the robot arm dominates the task reward, and the credibility of other auxiliary devices is zero.
[0049] when When , it is a symmetric cooperation scenario. When all agents’ actions contribute equally to the joint reward (such as in a multi-agent formation maintenance task, where each agent’s position adjustment contributes equally to the overall formation stability), the trust function satisfies For any Established. At this point, the system is in a completely fair cooperative state, and the credit distribution does not depend on the difference in individual actions.
[0050] However, in interference scenarios, the credit allocation function is not a stable, fixed function; it changes with the state of the electromagnetic space, exhibiting a certain degree of dynamism. This dynamism stems from the non-stationary nature of the electromagnetic environment and the adversarial nature of strategies. Its core characteristic is that the credit function evolves dynamically with real-time parameters such as spectrum state, interference pattern, and power allocation.
[0051] Furthermore, in step 130 of the embodiment of the present application, based on the above-established distributed frequency-domain blocking jamming model, partially observable Markov game model, and credit allocation model, an improved Counterfactual Multi-Agent Policy Gradient (ICOMA) algorithm and a Finite-length Elite Trajectory Experience Replay Mechanism (FETERM) are used to optimize the credit allocation strategy for multiple jammers against multiple drones. The Counterfactual Baseline is primarily used to address the credit allocation problem in multi-agent reinforcement learning. It simulates the impact of different actions chosen by an agent on the global reward while keeping the actions of other agents fixed, thereby more accurately measuring the contribution of the current action and reducing the variance in policy gradient updates.
[0052] Based on the idea of counterfactual baseline, the embodiment of the present application subtracts the global reward in the current situation from the global reward after replacing the agent action with a “default action”, as shown in formula (10).
[0053] (10) in, Represents a single agent Take joint action The independent reward after the return is actually calculated by the agent Take joint action will take the default action To be better ( ) or worse ( ); Represents the state of the agent group in the environment Take joint action The global reward obtained after Represents the joint action space excluding the current agent The action taken at this moment; Represents the current agent Take the "default action" The joint action space of all agents.
[0054] The specific actions of the above specific agents It is called a counterfactual baseline. Using the Counterfactual Multi-Agent Policy Gradient (COMA) operator, a counterfactual baseline is calculated for each action of each agent.
[0055] Among them, COMA uses a Center-Critic network and multiple Actor networks to implement a "centralized learning, distributed execution" model. That is, during the training phase, global information is input to the central Critic network, and the central Critic network teaches each Actor network how to make action decisions. However, unlike traditional AC algorithms or QMIX algorithms, the Critic network in COMA does not directly estimate the state value function. Or global total return , but to estimate the number of actions that an agent can take in a given state. values, and through these The value completes the calculation of the counterfactual baseline. Optionally, the embodiment of the present application can adopt Figure 5 The overall structure of COMA is shown in Figure 1. The Critic network and Actor network structures used are as follows: Figure 6 shown.
[0056] In the Actor network, it is first processed by the MLP composed of the linear layer, GELU activation layer, and Dropout layer, and then the GRU modeling time series information is input. Its output is processed by the linear layer and SoftMax to output the agent. action strategies to support intelligent agents in learning and making decisions in dynamic environments.
[0057] In the Critic network, features are first extracted through the linear layer + GELU activation, passed through the Dropout layer, and then processed through a multi-layer linear layer + GELU layer structure, and finally the state-action value corresponding to the joint action is output.
[0058] Among them, the input of the Actor network is 3, namely: the current agent Observation , agent One-hot encoding, the joint action of other agents except itself at the previous moment .
[0059] The Critic network has five inputs: the current global state , the current agent Observation , agent The unique hot encoding of all agents’ actions at the last moment , the joint action of other agents except themselves .
[0060] Use the Critic network to calculate , take the “average utility value” of all actions as the action utility value of the “default action”, and the action utility value calculation formula of the default action is shown in formula (11).
[0061] (11) in, The state-action value representing the default action, Representing an agent Current strategy, Representing an agent In the current strategy The following actions are taken, It is the output of the Actor network after passing through the SoftMax layer. Representing an agent In the current strategy Actions taken The probability of Output from the Critic network, Indicates that it is in the state Downsampling Action Combination The value of a state-action pair.
[0062] Will Equivalent to , then formula (10) becomes formula (12).
[0063] (12) in, Represents a separate baseline computed for each agent that utilizes a centralized critic to reason about the agent-only Counterfactuals in which the agent's actions change are learned directly from the agent's experience, without relying on additional simulations, reward models, or user-designed default actions.
[0064] COMA is a policy-based (PB) method. For this purpose, the Actor network update uses the traditional PB update formula, which is shown in Equation (13).
[0065] (13) in, Indicates the policy parameters Find the gradient, which points in the direction of increasing the probability of the current action; Indicates that the value output by the Actor network after Softmax takes the logarithm log, where Obtained using independent return calculations proposed by COMA, using Right now Combining Equation (12) and Equation (13), the loss calculation formula is shown in Equation (14).
[0066] (14) The critic network is updated using the traditional temporal-difference error (TD-Error). TD-Error includes and Two update methods, sampled in this application example The update method of , its loss function is as follows.
[0067] (15) (16) (17) in, represents the state-action value function, Indicates that the parameter is The state-action value function, represents the weighted sum of the steps, is the discount factor, For the immediate reward at the next moment; Indicates from time Start by considering The cumulative return of the step, The state value function is used to measure the jammer arrival state The value of express According to equations (15) to (17), the loss function can be expressed as equation (18).
[0068] (18) In the scenario of multiple jammers against multiple drones, this embodiment of the application proposes an improved Counterfactual Multi-Agent Policy Gradient (ICOMA) algorithm, combining the COMA algorithm with the idea of holistic adversarial game and cooperation, to optimize the credit allocation strategy during the process of multiple jammers interfering with multiple drones. The basic elements required by this algorithm are as follows: (1) Global observation: mainly composed of the frequency sequence similarity coefficient and threat coefficient of the frequency hopping graph, .in, is the frequency sequence similarity coefficient. The frequency sequence of the first received frequency hopping pattern is taken as the starting sequence. The similarity coefficient corresponding to the starting sequence is If it is 0, the frequency sequence similarity coefficient of the subsequent frequency hopping pattern changes The calculation uses the cost time warping (CTW) distance calculation; For the degree of threat coefficient, when the movement of a certain drone changes and the distance from us changes, the closest drone is defined as having the highest threat coefficient. The jammer will prioritize dispatching frequency domain resources to interfere with the drone with the highest threat coefficient.
[0069] (2) Local observation: independent observation by each drone .in, and The definition of is the same as that of the global observation above, The action of the drone at the last moment.
[0070] (3) Action space: The interference bandwidth of each jammer Fixed, in the bandwidth of 0-100M, it is divided into 400 segments, each with a bandwidth of 0.25M. The starting point of the interference band is selected as the action, then the action space Size ,in The interference bandwidth of each jammer is preferably , that is, the interference bandwidth of each jammer is 5.25M.
[0071] (4) Reward function: , The value range is The reward mainly has two parts and composition. Represents positive reward, which is expressed as follows: (19) (20) Represents a negative reward, when If you are unable to effectively interfere with a certain drone, you will receive a reward of -0.05. is a cumulative negative reward, then There are five possible values of , as shown in formula (21).
[0072] (twenty one) It can be understood that formula (21) is only an example and does not limit the value of the negative reward. The value of the negative reward can be selected according to actual needs.
[0073] In order to improve the stability of training, two neural networks with complementary functions are constructed for the Critic network: the value assessment network With the target network The two network architectures are exactly the same and the initial parameter configurations are the same: Value Evaluation Network Used to calculate the estimated value of the current state of the intelligent agent in real time, and use this as a basis to guide the action selection strategy; target network It is responsible for generating the target value and providing a stable reference signal for the training process. represents the weight parameters of the value evaluation network in the nth training round, This corresponds to the weight parameters of the target network at the same time. Through this dual network structure design, the variance problem of value estimation in traditional learning algorithms is effectively alleviated, and the stability and convergence efficiency of the training process are improved. The update sampling soft update method is shown in formula (22).
[0074] (twenty two) in, It is a soft update hyperparameter with a value range of [0,1].
[0075] When training a network, in order to prevent the network from overfitting, the embodiment of the present application uses a method of sampling weight decay to reduce overfitting. There are many weight decay methods. The embodiment of the present application uses a method of sampling L2 norm regularization to reduce overfitting. Regularization is a common means of dealing with overfitting by adding a penalty term to the model loss function to make the trained model parameter value smaller. For vector , its L2 norm is expressed as: (twenty three) In machine learning model training, to prevent overfitting, the L2 norm regularization term is often added to the original loss function to obtain a new loss function: (twenty four) (25) in, 、 are the loss functions of the Actor network and the Critic network with L2 norm regularization added respectively; is the parameter vector of the Actor network; is the parameter vector of the Critic network; is a regularization hyperparameter that controls the strength of the regularization. The larger the value, the stronger the constraints on the model parameters and the greater the effect of preventing overfitting, but it may also lead to underfitting of the model. The smaller it is, the weaker the regularization effect. It represents the number of training samples.
[0076] From the perspective of parameter constraints, L2 norm regularization works by penalizing large parameter values. The existence of this term will cause the parameter vector The values of the individual elements of tend to decrease. In some models, such as neural networks, if parameter values are too large, the model is likely to overfit to the noise in the training data. However, constraining the parameters to a smaller range of values can reduce the complexity of the model and thus enhance its generalization ability. Adding L2 norm regularization can limit the size of these coefficients, making the model smoother.
[0077] In addition, to improve the exploratory nature of training, an entropy regularization term (entropy) is introduced when calculating the loss of the Actor model. The calculation formula is: (26) At this point, the loss of the complete Actor model is: (27) The loss function The optimization of the policy (through the policy gradient term) and exploration (through the entropy regularization term) are balanced. During the training process, by minimizing the loss function, the Actor model can continuously adjust parameters and generate better policies to maximize the long-term cumulative rewards.
[0078] Although the random sampling method used in the traditional experience replay pool can alleviate the data correlation problem, it has many disadvantages. Especially in the multi-agent scenario, the traditional experience replay pool has problems such as poor training stability, high variance, poor gradient estimation quality, and limited convergence speed. In order to adapt to the multi-agent joint interference scenario and promote training stability and rapid convergence, the embodiment of the present application improves the traditional experience replay pool and combines it with the traditional experience replay pool. , proposed a finite-length elite trajectory experience replay mechanism (Finite-length Elite Trajectory Experience Replay Mechanism, FETERM).
[0079] According to formula (15)-formula (18), It is calculated based on the empirical trajectory obtained by sampling, and the final result is The track length set in the embodiment of the present application is There are two track experience pools, one is the ordinary track experience pool, and the other is the elite track experience pool. Set the reward threshold For the common trajectory experience pool, when the number of steps stored in the trajectory reaches When , the track is stored in the ordinary track experience pool; for the elite track experience pool, when a track is In the trajectory of , then the trajectory is stored in the elite trajectory experience pool. The reward threshold The size of becomes larger as the number of rounds increases, and its calculation formula is as follows: (28) in, is the current round number, is the total number of rounds.
[0080] The structure diagram of the finite length elite trajectory experience replay mechanism is as follows Figure 7 shown.
[0081] When the algorithm trains the network, the sample extraction is determined according to the following formula (29).
[0082] (29) (30) in, is the sampling ratio coefficient; is the number of samples drawn; is the number of samples in the common trajectory experience pool; is the number of samples in the elite trajectory experience pool.
[0083] The complete ICOMA algorithm process is as follows: Establish a common trajectory experience pool and an elite trajectory experience pool, initialize the sampling ratio coefficient, set the reward threshold, initialize the storage trajectory length, set the update step size, and set the soft update hyperparameters; Initialize the Critic value evaluation network, Critic target network and N Actor networks, whereN is the number of jammers; initialize the maximum round, the step size of each round, and the greedy parameter , set the greedy decay parameter ; Get the system status of the current round at the current moment , and the actions of all agents at the previous moment ; For the current agent, if the greedy parameter is greater than the random value, the action of the current agent at the current moment is randomly obtained according to the greedy strategy. ; Otherwise, get the current agent's observation , agent The unique hot encoding of the last moment is the joint action of other agents except itself , input into the Actor network to get the output , and then get the action of the current agent at the current moment Repeat this step until all agents are processed and the actions corresponding to the current moment of all agents are obtained. .
[0084] According to the current actions of all agents , get the total reward , and obtain the system status at the next moment ; Save all information as ; If the current moment can divide the stored trajectory length (Right now ), then (i.e. the information saved in the historical trajectory length L) is pushed into the common trajectory experience pool; if the trajectory length If the total reward of a certain step is greater than the reward threshold, Advance in the Elite Track experience pool.
[0085] If the current moment can divide the update step size and the number of samples in the two experience pools is sufficient, then sample according to the sample extraction formula (i.e., formula (29)) and use Calculate the loss value of the Critic network and find the advantage function , and then get N The loss value of the Actor network; according to the target network update formula (i.e., formula (22)), update the Critic target network; update the sampling ratio coefficient and update the reward threshold; according to the formula , update the greedy parameters.
[0086] Repeat the above process to perform optimization at the next moment until the current round of optimization is completed.
[0087] Repeat the above process to perform the next round of optimization until the maximum round of optimization is completed and the optimization results are output.
[0088] The embodiment of the present application also proposes an interference frequency domain resource scheduling device for multi-UAV communication, such as Figure 8 As shown, the interference frequency domain resource scheduling device 200 includes: The first modeling unit 201 establishes a distributed frequency domain blocking interference model based on the interference-to-signal ratio criterion and in combination with a multi-jammer multi-UAV frequency hopping communication system scenario. The interference model establishment process is as described in step 110 above and will not be repeated here.
[0089] The second modeling unit 202 is used to treat multiple jammers as a group of intelligent agents, construct the process of jamming multiple UAV communications as a partially observable Markov game process, and establish a multi-agent trust distribution model based on the trust distribution criterion. The specific process is as described in step 120 above and will not be repeated here.
[0090] Furthermore, optimization unit 203 utilizes an improved counterfactual multi-agent policy gradient algorithm and a finite-length elite trajectory experience replay mechanism to optimize the credit allocation strategy during the multi-jammer jamming of multiple drones, obtaining and outputting the optimal jamming frequency domain resource scheduling strategy. The specific optimization process is as described in step 130 above and will not be repeated here.
[0091] The embodiment of the present application also proposes an interference frequency domain resource scheduling system for multi-UAV communication, such as Figure 9 As shown, the interference frequency domain resource scheduling system 300 proposed in the embodiment of the present application includes: Input device 301, output device 302, processor A303 and memory A304; wherein the number of processor A303 and memory A304 can be one or more, Figure 9 The input device 301, the output device 302, the processor A303 and the memory A304 can be connected by a bus or other means. Figure 9 The bus connection is taken as an example.
[0092] By calling the operation instructions stored in the memory A304, the processor A303 is configured to perform the following steps: Based on the interference-to-signal ratio criterion and combined with the multi-jammer frequency-hopping communication system scenario for multiple UAVs, a distributed frequency domain blocking jamming model is established; Multiple jammers are regarded as a group of intelligent agents, and the process of jamming multiple UAV communications is constructed as a partially observable Markov game process. Based on the credit distribution criterion, a credit distribution model for multiple agents is established. The improved counterfactual multi-agent policy gradient algorithm and the finite-length elite trajectory experience replay mechanism are used to optimize the credit allocation strategy in the process of multiple jammers interfering with multiple UAVs, and the optimal interference frequency domain resource scheduling strategy is obtained and output.
[0093] Optionally, by calling the operation instructions stored in the memory A304, the processor A303 is further configured to execute any implementation method in the corresponding embodiments of the above-mentioned interference frequency domain resource scheduling method.
[0094] In another embodiment, the present application also provides an electronic device, such as Figure 10 As shown, the electronic device 400 includes: a memory B410, a processor B420, and a computer program A411 stored in the memory B410 and executable on the processor B420. When the processor B420 executes the computer program A411, the following steps are implemented: Based on the interference-to-signal ratio criterion and combined with the multi-jammer frequency-hopping communication system scenario for multiple UAVs, a distributed frequency domain blocking jamming model is established; Multiple jammers are regarded as a group of intelligent agents, and the process of jamming multiple UAV communications is constructed as a partially observable Markov game process. Based on the credit distribution criterion, a credit distribution model for multiple agents is established. The improved counterfactual multi-agent policy gradient algorithm and the finite-length elite trajectory experience replay mechanism are used to optimize the credit allocation strategy in the process of multiple jammers interfering with multiple UAVs, and the optimal interference frequency domain resource scheduling strategy is obtained and output.
[0095] Optionally, when the processor B420 executes the computer program A411, it can implement any implementation method corresponding to the above-mentioned interference frequency domain resource scheduling method.
[0096] It should be noted that the electronic device proposed in the embodiment of the present application is a device used to implement the above-mentioned interference frequency domain resource scheduling method. Therefore, based on the above-mentioned interference frequency domain resource scheduling method proposed in the embodiment of the present application, technical personnel in this field can understand the specific implementation methods of the electronic device in the embodiment of the present application and its various variations. Therefore, how the electronic device specifically implements the above-mentioned interference frequency domain resource scheduling method will not be introduced in detail here. As long as the electronic device used by technical personnel in this field to implement the above-mentioned interference frequency domain resource scheduling method falls within the scope of protection to be protected by this application.
[0097] In another embodiment, the present application also provides a computer-readable storage medium, such as Figure 11 As shown, the computer readable storage medium 500 stores a computer program B511. When the computer program B511 is executed by the processor, the following steps are implemented: Based on the interference-to-signal ratio criterion and combined with the multi-jammer frequency-hopping communication system scenario for multiple UAVs, a distributed frequency domain blocking jamming model is established; Multiple jammers are regarded as a group of intelligent agents, and the process of jamming multiple UAV communications is constructed as a partially observable Markov game process. Based on the credit distribution criterion, a credit distribution model for multiple agents is established. The improved counterfactual multi-agent policy gradient algorithm and the finite-length elite trajectory experience replay mechanism are used to optimize the credit allocation strategy in the process of multiple jammers interfering with multiple UAVs, and the optimal interference frequency domain resource scheduling strategy is obtained and output.
[0098] Optionally, when the computer program B511 is executed by a processor, it can implement any implementation method of the embodiments corresponding to the above-mentioned interference frequency domain resource scheduling method.
[0099] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0100] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0101] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0102] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1The function specified in one or more boxes.
[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0104] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application. It should be understood that the above description is only the specific implementation methods of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.
Claims
1. A method for scheduling interference frequency domain resources for multi-UAV communications, characterized in that: include: Based on the interference-to-signal ratio criterion and combined with the multi-jammer frequency-hopping communication system scenario for multiple UAVs, a distributed frequency domain blocking jamming model is established; Multiple jammers are regarded as a group of intelligent agents, and the process of jamming multiple UAV communications is constructed as a partially observable Markov game process. Based on the credit distribution criterion, a credit distribution model for multiple agents is established. The improved counterfactual multi-agent policy gradient algorithm and the finite-length elite trajectory experience replay mechanism are used to optimize the credit allocation strategy in the process of multiple jammers interfering with multiple UAVs, and the optimal interference frequency domain resource scheduling strategy is obtained and output.
2. The method for scheduling interference frequency domain resources for multi-UAV communication according to claim 1, characterized in that: The establishment of a distributed frequency domain blocking interference model includes: When the interference bandwidth coverage is greater than or equal to the UAV K If the interference covers more than one-third of the target frequency points, the interference is effective; otherwise, the interference is invalid. Among them, drones K The value acquisition methods include: Calculating a first ratio of the jammer transmission power to the signal transmitter power; Calculating a second ratio of the product of the jammer transmitting antenna gain and the signal receiving antenna gain to the product of the signal transmitter antenna gain and the signal receiving antenna gain; calculating a third ratio of the square value of the transmission distance of the communication signal to the square value of the transmission distance of the jammer signal; calculating the product of the first ratio, the second ratio, and the third ratio; Calculate the ratio of the suppression coefficient to the product to obtain the value.
3. The method for scheduling interference frequency domain resources for multi-UAV communication according to claim 2, characterized in that: The interference bandwidth coverage is obtained by: Calculate the signal bandwidth of all target drones in the interference airspace being jammed by a certain jammer; Calculating a ratio of the signal bandwidth to the interference bandwidth of the jammer; Repeat the above process to obtain and accumulate the ratios corresponding to all jammers to obtain the interference bandwidth coverage.
4. The interference frequency domain resource scheduling method for multi-UAV communication according to claim 1 is characterized in that: The process of interfering with multi-UAV communications is constructed as a partially observable Markov game process, and a multi-agent trust allocation model is established based on the trust allocation criterion, including: The partially observable Markov game is described as a tuple, which includes the system state observed by the jammer group, the joint action of the jammer group, the transition probability of the system state, the joint reward obtained by the jammer group after performing the joint action, the number of jammer agents, the independent observation space of each jammer agent, and the transition probability of the observation space. The system state space is composed of the individual states of all jammer agents and the global environment state; The joint action space consists of the action spaces of all jammer agents; The joint reward function is expressed as the sum of the independent action rewards of all jammer agents and the interaction rewards between all jammer agents; Define the credit distribution function: when performing a joint action under a certain system state, the contribution weight of the action of a single jammer agent to the joint reward, and the sum of the contribution weights of the actions of all jammer agents to the joint reward is equal to 1.
5. The method for scheduling interference frequency domain resources for multi-UAV communication according to any one of claims 1 to 4, characterized in that: The aforementioned method of optimizing the credit allocation strategy in the process of multiple jammers interfering with multiple UAVs by utilizing the improved counterfactual multi-agent policy gradient algorithm and the finite-length elite trajectory experience replay mechanism includes: Establish a common trajectory experience pool and an elite trajectory experience pool, initialize the sampling ratio coefficient, initialize the reward threshold, set the storage trajectory length, set the update step size, and set the soft update hyperparameters; Initialize the Critic value evaluation network, Critic target network and N Actor networks, where N is the number of jammers; initialize the maximum rounds, the compensation of each round, the greedy parameter and the greedy attenuation parameter; Get the system status at the current moment of the current round, as well as the actions of all agents at the previous moment; For the current agent, if the greed parameter is greater than the random value, the action of the current agent at the current moment is randomly obtained according to the greedy strategy; otherwise, the observation space of the current agent, the unique hot encoding of the current agent, and the joint action of other agents except itself at the previous moment are obtained, and input into the Actor network to obtain the output, thereby obtaining the action of the current agent at the current moment; repeat this step until all agents are processed and the actions of all agents at the current moment are obtained; According to the current actions of all agents, the total reward is obtained and the system state at the next moment is obtained; Save all information; the saved information includes: the system state at the current moment, the system state at the next moment, the observation space of all agents, the actions of all agents at the current moment, the total reward, and the actions of all agents at the previous moment; If the current moment can divide the stored trajectory length, the information stored in the historical stored trajectory length is pushed into the ordinary trajectory experience pool; if the total reward of a certain step in the historical stored trajectory length is greater than the reward threshold, the information stored in the historical stored trajectory length is pushed into the elite trajectory experience pool; If the current moment can divide the update step size and the number of samples in the two experience pools is sufficient, sampling is performed, and the loss value of the Critic network is calculated, and the advantage function is obtained, and then N The loss value of the Actor network; Update the Critic target network; update the sampling ratio coefficient, update the reward threshold, and update the greed parameter; Repeat the above process to perform the optimization at the next moment until the current round of optimization is completed; Repeat the above process to perform the next round of optimization until the maximum round of optimization is completed and the optimization results are output.
6. The method for scheduling interference frequency domain resources for multi-UAV communication according to claim 5, characterized in that: The critic value evaluation network and the critic target network have the same network architecture and the same initial parameter configuration. The critic value evaluation network is used to calculate the estimated value of the agent's current state in real time, which is used as a basis to guide the action selection strategy. The critic target network is responsible for generating the target value and providing a stable parameter signal for the training process. The update of the critic target network adopts the soft update method, and the update formula is expressed as: ; in, is a soft update hyperparameter, is the weight parameter of the Critic value evaluation network; is the weight parameter of the Critic target network.
7. The method for scheduling interference frequency domain resources for multi-UAV communication according to claim 5, characterized in that: The loss function of the Critic value evaluation network and the Critic target network is expressed as: ; ; in, represents the state-action value function, Indicates that the parameter is The state-action value function, represents the weighted sum of the steps, is the discount factor, For the immediate reward at the next moment, is the trajectory length, is the parameter vector of the Critic network; is the regularization hyperparameter; is the number of training samples; The loss function of the Actor network is expressed as: ; ; ; in, Indicates the policy parameters Find the gradient, Indicates that the Actor network output is logarithmic, Representing an agent Current strategy, Representing an agent In the current strategy The following actions are taken, is the Actor network output, Representing an agent In the current strategy Actions taken The probability of Indicates that the status Downsampling Action Combination The state-action pair value, is the parameter vector of the Actor network, Indicates the current strategy The entropy regularization term under .
8. The method for scheduling interference frequency domain resources for multi-UAV communication according to claim 5, characterized in that: The total reward consists of positive rewards and negative rewards; the positive reward is calculated by accumulating the positive rewards of all jammers that achieve effective interference at the current moment; the negative reward is calculated by accumulating the negative rewards of all jammers that cannot effectively interfere with the drone at the current moment.
9. The method for scheduling interference frequency domain resources for multi-UAV communication according to claim 5, characterized in that: The size of the reward threshold increases as the number of rounds increases; The sampling ratio coefficient decreases as the number of rounds increases; The number of samples obtained by the sampling is equal to the product of the number of samples in the ordinary trajectory experience pool and the sampling ratio coefficient plus the product of the number of samples in the elite trajectory experience pool and (1-sampling ratio coefficient).
10. An interference frequency domain resource scheduling device for multi-UAV communication, characterized in that: include: The first modeling unit, based on the interference-to-signal ratio criterion and combined with the multi-jammer multi-UAV frequency hopping communication system scenario, establishes a distributed frequency domain blocking jamming model; The second modeling unit is used to regard multiple jammers as a group of intelligent agents, construct the process of jamming multiple UAV communications as a partially observable Markov game process, and establish a multi-agent trust distribution model based on the trust distribution criterion; In addition, the optimization unit uses the improved counterfactual multi-agent policy gradient algorithm and the finite-length elite trajectory experience replay mechanism to optimize the trust allocation strategy in the process of multiple jammers interfering with multiple UAVs, and obtains and outputs the optimal interference frequency domain resource scheduling strategy.
Citation Information
Cited By
Unmanned aerial vehicle group dynamic interception method and system based on distributed interference array cooperation
CN121239346A