Air-space-ground integrated network anti-interference communication method based on reinforcement learning

By constructing intelligent agents in an integrated air-space-ground network and optimizing communication strategies using a distributed Q-learning algorithm, the problem of insufficient adaptability of traditional anti-interference strategies in highly dynamic environments is solved, and more efficient anti-interference communication is achieved.

CN120935655APending Publication Date: 2025-11-11CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511132881.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing traditional spread spectrum anti-jamming strategies are not adaptable enough to the highly dynamic and adversarial interference environment in integrated air-space-ground networks, and lack coordination and intelligent countermeasure capabilities, making it difficult to cope with intelligent jamming attacks.

Method used

An intelligent agent is constructed, and the agent is trained to learn the optimal communication strategy by combining the average received power of the channel, the transmission channel and the power set. The communication strategy is then optimized by constructing a reward function and a target optimization model.

Benefits of technology

It improves the anti-interference communication real-time performance and effectiveness of the integrated air-space-ground network, reduces computational complexity, enables lightweight deployment, and ensures continuous and stable communication between satellites, drones, and ground users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120935655A_ABST
    Figure CN120935655A_ABST
Patent Text Reader

Abstract

The invention relates to an air-space-ground integrated network anti-interference communication method based on reinforcement learning, and the method comprises the steps: constructing intelligent agents for a communication link between an unmanned plane and a ground user, and a communication link between the unmanned plane and a low-orbit satellite; constructing a state space of the intelligent agent according to the average receiving power of each available channel between the sending end and the receiving end in the current time slot; constructing an action space of the intelligent agent according to the transmission channel set and the transmission power set between the sending end and the receiving end; constructing a reward function of the intelligent agent according to the received signal interference plus noise ratio of the link between the sending end and the receiving end; constructing a target optimization model by taking the maximum accumulated discount rewards of all agents as an optimization target; the intelligent agents are trained by using a distributed Q learning algorithm based on similarity measurement and the constructed target optimization model, so that each intelligent agent learns an optimal communication strategy; and the intelligent agent outputs the optimal transmission channel and transmission power between the sending end and the receiving end through the learned optimal communication strategy according to the observed state. The method can adapt to the environment of dynamic change of an interference source and time varying of network topology under the SAGIN, improves the real-time performance and effectiveness of anti-interference communication, reduces the calculation complexity of reinforcement learning, accelerates learning convergence, achieves lightweight deployment in the SAGIN, and guarantees continuous and stable communication among a satellite, an unmanned aerial vehicle and a ground user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of communication technology, and in particular relates to an anti-interference communication method for an integrated air-space-ground network based on reinforcement learning. Background Technology

[0002] As one of the core architectures of future 6G, the Space-Air-Ground Integrated Network (SAGIN) integrates three heterogeneous subnetworks—space-based, airborne, and terrestrial—to fully leverage the wide coverage, deployment flexibility, and transmission complementarity of each layer, providing users with seamless wide-area interconnection services. However, communication between users, drones, and satellites in SAGIN primarily relies on line-of-sight (LoS) transmission, operating in a complex electromagnetic environment. This makes the links vulnerable to malicious interference and attacks. Specifically, malicious jammers can transmit jamming signals into wireless channels, preventing legitimate access to the wireless medium or disrupting signal reception on specific channels. With increasingly complex network topologies and electromagnetic environments, and the continuous development of jamming technologies, malicious interference has become a serious threat to SAGIN networks.

[0003] Traditional spread spectrum anti-jamming strategies, such as Frequency Hopping Spread Spectrum (FHSS) and Direct Sequence Spread Spectrum (DSSS), significantly reduce the power spectral density of a signal by spreading its energy across the frequency band. They effectively resist narrowband interference thanks to the low detectability afforded by spread spectrum and specific processing mechanisms at the receiver. However, due to the high dependence of spread spectrum anti-jamming on precise synchronization and pseudo-random sequences, it is severely inadequate in adapting to the high latency caused by high-speed node movement, link switching, and rapid follow-up jamming attacks from intelligent jammers in SAGIN. It struggles to cope with the highly dynamic and adversarial jamming environment of SAGIN. Furthermore, against intelligent jamming attacks, traditional single-domain anti-jamming defenses are limited in scope, relatively static in strategy, and lack coordination and intelligent countermeasure capabilities, resulting in significant limitations. Summary of the Invention

[0004] To address the problems existing in the background art, one aspect of the present invention provides a reinforcement learning-based anti-jamming communication method for an integrated air-space-ground network. The integrated air-space-ground network includes a drone, ground users, low-Earth orbit satellites, an airborne jammer, and a ground-based jammer. The airborne jammer interferes with the communication link between the drone and the low-Earth orbit satellites, and the ground-based jammer interferes with the communication link between the drone and the ground users. The communication method includes:

[0005] S1: Construct intelligent agents for the communication links between drones and ground users, and for the communication links between drones and low-Earth orbit satellites;

[0006] S2: Construct the state space of the agent based on the average received power of each available channel between the transmitter and receiver in the current time slot; where drones, ground users, and low-Earth orbit satellites can all serve as transmitters or receivers;

[0007] S3: Construct the action space of the agent based on the set of transmission channels and the set of transmission power between the sender and receiver;

[0008] S4: Construct the agent's reward function based on the interference plus noise ratio of the received signal in the link between the transmitter and receiver;

[0009] S5: Construct a target optimization model with the goal of maximizing the cumulative discount reward of all agents;

[0010] S6: Train the agents using a distributed Q-learning algorithm based on similarity metrics and a constructed objective optimization model so that each agent learns the optimal communication strategy.

[0011] S7: The agent outputs the optimal transmission channel and transmission power between the sender and receiver based on the observed state and the learned optimal communication strategy.

[0012] Another aspect of the present invention provides a reinforcement learning-based integrated air-space-ground network anti-jamming communication system, characterized in that the system includes a memory and a processor; the memory is used to store an application program; the processor is used to run the application program and execute the reinforcement learning-based integrated air-space-ground network anti-jamming communication method as described in any one of claims 1 to 6.

[0013] A computer storage medium, characterized in that the computer storage medium stores a remote monitoring program, which, when executed by a processor, implements a reinforcement learning-based anti-interference communication method for integrated air-space-ground networks as described in any one of claims 1 to 6.

[0014] The present invention has at least the following beneficial effects

[0015] This invention constructs intelligent agents for communication links between UAVs and ground users, and between UAVs and low-Earth orbit satellites. It constructs a state space by combining the average received power of the channel, an action space by combining the transmission channel and power set, and a reward function by combining the received signal interference plus noise ratio. With the goal of maximizing the cumulative discount reward, it uses a distributed Q-learning algorithm based on similarity metric to train the intelligent agents to learn the optimal communication strategy. This can adapt to the environment of dynamic changes in interference sources and time-varying network topology in the integrated air-space-ground network, improve the real-time performance and effectiveness of anti-interference communication, reduce the computational complexity of reinforcement learning, accelerate learning convergence, and achieve lightweight deployment, thereby ensuring continuous and stable communication between satellites, UAVs, and ground users. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0017] Figure 2 This is a schematic diagram of the communication framework of the present invention. Detailed Implementation

[0018] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0019] Please see Figure 1 and Figure 2 One aspect of the present invention provides a reinforcement learning-based anti-jamming communication method for an integrated air-space-ground network. The integrated air-space-ground network includes a drone, ground users, low-Earth orbit satellites, an airborne jammer, and a ground-based jammer. The airborne jammer jams the communication link between the drone and the low-Earth orbit satellites, and the ground-based jammer jams the communication link between the drone and the ground users. The communication method includes:

[0020] S1: Construct intelligent agents for the communication links between drones and ground users, and for the communication links between drones and low-Earth orbit satellites;

[0021] Example 1: In this example, we consider a... A distributed wireless network consisting of transmitter-receiver pairs, including a low-Earth orbit satellite, a single-antenna UAV, and... One ground user G, one intelligent ground jammer J l A smart airborne jammer J h Dividing the transmission time T into D equal-length time slots is represented as: Ground users, drones, and low-Earth orbit satellites can all send and receive data as nodes. An intelligent agent is constructed for each connection between a drone and a low-Earth orbit satellite, and between a ground user and a drone. This network includes... One communication link An intelligent agent can be represented as All transmit and receive pairs are capable of time-division communication within bandwidth B; and possess dynamic spectrum access and power control capabilities; the available channel set for the ground-UAV link is... , This indicates the center frequency of the corresponding available channel, and the set of available power is... , The bandwidth of each non-overlapping channel is The available channel set for UAV-satellite links is as follows: , This represents the center frequency of the corresponding available channel, and the set of available power is... , The bandwidth of each non-overlapping channel is Ground-based and airborne jammers can dynamically adjust their jamming power and channel between different time slots. The ground jamming power set and jamming channel set are respectively represented as follows: , The power set of the airborne jammer and the jamming channel set are respectively , .

[0022] S2: Construct the state space of the agent based on the average received power of each available channel between the transmitter and receiver in the current time slot; where drones, ground users, and low-Earth orbit satellites can all serve as transmitters or receivers;

[0023] Preferably, the state space of the constructed agent includes:

[0024] S21: Constructing a channel gain model between UAVs and low-Earth orbit satellites based on the effects of UAV beamforming gain, spatial loss, and small-scale fading:

[0025] S22: Construct a channel gain model between the UAV and ground users based on line-of-sight loss and non-line-of-sight loss as defined in 3GPP;

[0026] S23: Calculate the power spectral density function between the UAV and the low-Earth orbit satellite based on the channel gain model between the UAV and the low-Earth orbit satellite and the channel gain of the airborne jammer;

[0027] S24: Calculate the power spectral density function between the UAV and the ground user based on the channel gain model between the UAV and the ground user and the channel gain of the ground jammer;

[0028] S25: Discretize the power spectral density function between the UAV and the ground user, and calculate the average received power of each available channel for the UAV and the ground user in the current time slot;

[0029] S26: Discretize the power spectral density function between the UAV and the low-Earth orbit satellite, and calculate the average received power of each available channel of the UAV and the low-Earth orbit satellite in the current time slot;

[0030] S27: Construct the state space of the agent based on the average received power of each available channel of the UAV and ground user in the current time slot and the average received power of each available channel of the UAV and low-orbit satellite in the current time slot.

[0031] Example 2 further extends Example 1. In this example, a UAV is used as the transmitter and a low-Earth orbit satellite as the receiver. The channel gain model between the UAV and the low-Earth orbit satellite is as follows:

[0032]

[0033]

[0034]

[0035]

[0036] in, Indicates in time slot drones and low-orbit satellites Channel gain of the link between them; The wavelength representing the carrier frequency; Represents the speed of light; Indicates the carrier frequency; This represents the straight-line distance between the drone u and the low-orbit satellite s; Represents pi; Indicates in time slot drones and low-orbit satellites The small-scale fading of the links between them can be modeled using a correlated shadowed Rice distribution; This indicates the maximum gain of the drone's transmitting antenna; and These are first-order and third-order Bessel functions, respectively; Indicates the angle of deviation from the direction of maximum gain of the UAV antenna; It is represented as the first positive root of the first-order Bessel function; The 3 dB beamwidth represents the UAV antenna; k = 2π / λ represents the wavenumber; and a represents the antenna radius. Similarly, based on the same principle, a low-Earth orbit satellite can be used as the transmitter, and the UAV as the receiver to construct a channel gain model from the low-Earth orbit satellite to the UAV. This will not be elaborated in detail in this embodiment.

[0037] Example 3 further extends Example 1. In this example, a ground user is used as the transmitter and the drone as the receiver. The channel gain model between the ground user and the drone is described below:

[0038]

[0039]

[0040]

[0041]

[0042]

[0043] in, Indicates in time slot drones and the Channel gain between ground users; Indicates in time slot Drones and the Small-scale fading coefficients between individual ground users; Indicates the path loss conversion factor; Indicates the non-line-of-sight probability; Indicates the probability of line-of-sight distance; Distance of the breakpoint representing the probability of sight distance; The attenuation coefficient; It is an exponential function; Indicates in time slot drones and the The horizontal two-dimensional distance between ground users; The maximum value function is represented by A, B, C, and D, which represent the LoS path loss model parameters defined in 3GPP version 15; E, F, and G represent the NLoS path loss model parameters defined in 3GPP version 15. Indicates the drone's flight altitude; Indicates in time slot drones and the The distance between ground users; This represents the carrier frequency. Similarly, based on the same principle, a channel gain model from the UAV to the ground user can be constructed, with the UAV acting as the transmitter and the ground user as the receiver. This will not be elaborated in detail in this embodiment.

[0044] Example 4: This example further extends Example 2. In this example, a UAV is used as the transmitter and a low-Earth orbit satellite as the receiver. The power spectral density function between the UAV and the low-Earth orbit satellite is calculated as follows:

[0045]

[0046] in, This indicates that drones and low-orbit satellites are in time slots. Mid-moment The power spectral density function of the received signal; Represents any frequency point in the frequency domain; Indicates in time slot drones and low-orbit satellites Channel gain of the link between them; This indicates the transmission center frequency of the drone and low-orbit satellite at time 𝑡; The power spectral density function representing a bandpass signal; Indicates time slot T n The number of channels that are interfered with; Indicates the channel gain of the airborne jammer; The power spectral density function representing the interfering bandpass signal; Let represent the channel frequency of the i-th interference signal at time t; This represents the power spectral density function of the noise. Similarly, based on the same principle, a power spectral density function from a low-Earth orbit satellite to a drone can be constructed, using a low-Earth orbit satellite as the transmitter and a drone as the receiver. This will not be elaborated in detail in this embodiment.

[0047] Example 5, this example further extends Example 3. In this example, a ground user is used as the transmitter and a drone as the receiver. The power spectral density function between the ground user and the drone is calculated:

[0048]

[0049] in, Indicating drones and ground users In the time slot Mid-moment The power spectral density function of the received signal; Indicates the number of ground users; Indicates in time slot drones and the Channel gain between ground users; The power spectral density function representing a bandpass signal; Indicates drones and the The transmission center frequency of a ground user at time 𝑡; This represents any frequency point in the frequency domain; Indicates time slot T n The number of channels that are interfered with; Indicates the channel gain of the ground jammer; The power spectral density function representing the interfering bandpass signal; Let represent the channel frequency of the i-th interference signal at time t; This represents the power spectral density function of the noise. Similarly, based on the same principle, the power spectral density function from the UAV to the ground user can be constructed, with the UAV acting as the transmitter and the ground user as the receiver. This will not be elaborated in detail in this embodiment.

[0050] Example 6 further extends Example 5. In this example, a ground user is used as the transmitter and a drone as the receiver. The power spectral density function between the drone and the ground user is discretized, and the average received power of each available channel of the drone and the ground user at each moment in the current time slot is calculated.

[0051]

[0052] in, The first refers to drones and ground users. Available channels in the current time slot Mid-moment The average received power; This represents the ratio of channel bandwidth to spectrum sensing resolution. Indicates the spectrum sensing resolution; This indicates the initial sensing frequency. Similarly, in this embodiment, the average received power of each available channel of the UAV and the low-Earth orbit satellite at each moment in the current time slot can be calculated based on the same principle, which will not be elaborated in detail in this embodiment.

[0053] Example 7: This example further extends Example 6. Taking a ground user as the transmitter and a drone as the receiver, the average received power of each available channel of the drone and the ground user in the current time slot is calculated based on the average received power of each available channel in the current time slot at each moment.

[0054]

[0055] in, The first refers to drones and ground users. Available channels in the current time slot The average received power; This indicates the number of spectrum sensing operations within each time slot; The first refers to drones and ground users. Available channels in the current time slot Mid-moment The average received power; where, Ts The time slot length is denoted as . Similarly, in this embodiment, the average received power of each available channel of the UAV and the low-Earth orbit satellite in the current time slot can be calculated based on the same principle, which will not be elaborated in detail in this embodiment.

[0056] Example 8 is a further extension of Example 7. In this example, the state space of the agent is constructed as follows:

[0057]

[0058]

[0059]

[0060] in, The first refers to drones and ground users. Available channels in the current time slot The average received power; This refers to the intelligent agents between drones and ground users in time slots. The observed state; Represents the agent time slot between drones and low-Earth orbit satellites The observed state; This represents the set of states between the drone and the ground user; This represents the set of states between drones and low-orbit satellites; It represents the set of states of all agents.

[0061] In this embodiment, step S1 constructs intelligent agents for the communication links between the UAV and ground users, and between the UAV and low-Earth orbit satellites, respectively. This allows for targeted responses to interference sources (ground jammers and airborne jammers) on different links, enabling independent optimization of anti-interference strategies for each link. Step S2 constructs a state space based on the average received power of the current time slot of each available channel between the transmitter and receiver, and includes the cases of UAVs, ground users, and low-Earth orbit satellites as transceivers. This allows the intelligent agent to accurately perceive the real-time status and interference level of each channel. The combination of these two aspects provides a foundation for the intelligent agent to learn the optimal communication strategy in accordance with the dynamic characteristics of the integrated air-space-ground network, which helps improve the adaptability to dynamic interference environments of different links and the accuracy of state perception.

[0062] S3: Construct the action space of the agent based on the set of transmission channels and the set of transmission power between the sender and receiver;

[0063] Example 9 further extends the examples 1-8. In this example, the state space of the agent is constructed as follows: the state space of the agent in the current time slot. Selected transmission channel and transmission power.

[0064] The action is represented as: Selecting a transmission channel. and power The link corresponding to agent m in the time slot The action is defined as:

[0065]

[0066]

[0067] in, Represents the action space of link m, and the joint actions of all agents. It is the set of actions of all links in the network.

[0068] In this embodiment, step S3 constructs the action space of the agent based on the set of transmission channels and the set of transmission power between the sending end and the receiving end. This clarifies the transmission channels and transmission power that the agent can choose during communication, providing the agent with a specific and operable decision range. This allows the agent to flexibly select appropriate transmission parameters based on the actual channel state and interference situation, thereby better adapting to the environment of dynamic changes in interference sources and time-varying network topology in the integrated air-space-ground network. This lays the foundation for subsequent learning of the optimal communication strategy and helps improve the flexibility and effectiveness of anti-interference communication.

[0069] S4: Construct the agent's reward function based on the interference plus noise ratio of the received signal in the link between the transmitter and receiver;

[0070] Preferably, the reward function for constructing the intelligent agent includes:

[0071] S41: Calculate the interference-to-noise ratio of the received signal at the receiver between the UAV and the low-Earth orbit satellite based on the channel gain model between the UAV and the low-Earth orbit satellite, and the channel gain model of the airborne jammer.

[0072] S42: Calculate the received signal interference plus noise ratio between the UAV and the ground user based on the channel gain model between the UAV and the ground user, and the channel gain model of the ground jammer;

[0073] S43: Based on the interference plus noise ratio of the received signal at the receiving end, construct the reward function of the agent with the criterion that the smaller the transmission power between the transmitting end and the receiving end, the greater the reward value.

[0074] Example 10, this example further extends the examples 1-9, the reward function for constructing the agent includes:

[0075]

[0076]

[0077]

[0078]

[0079]

[0080] in, This indicates that all agents are in the time slot. Observed environmental conditions Next action The sum of rewards; M represents the number of agents; This indicates that the m-th agent is in time slot Observed environmental conditions Next action The reward; Indicates an indicator function, if ,but ,on the contrary ; This indicates the set threshold. Indicates the power discount factor; This represents the transmission power selected by the m-th agent; Indicates the maximum power; This represents the transmission power of agent m at time t; Represents intelligent agents The channel gain of the corresponding link; Indicates the bandwidth of the drone-satellite link; Represents intelligent agents The power spectral density function of the bandpass signal in the corresponding link; Indicates the bandwidth of the user-drone link; SJNR represents the received signal of the UAV-satellite link; This represents the interference plus noise ratio of the received signal in the ground user-UAV link.

[0081] In this embodiment, step S4 constructs a reward function for the agent based on the interference-to-noise ratio of the received signal between the UAV and the low-orbit satellite, and between the UAV and the ground user, with the criterion that the smaller the transmission power, the greater the reward value. This enables the agent to balance communication quality (reflected by the interference-to-noise ratio of the received signal) and transmission efficiency (encouraging low-power transmission through the reward mechanism) during the learning process. This guides the agent to learn strategies that ensure effective communication while reducing energy consumption, adapting to the dynamic interference environment of the integrated air-space-ground network and improving the efficiency and rationality of anti-interference communication.

[0082] S5: Construct a target optimization model with the goal of maximizing the cumulative discount reward of all agents;

[0083] Preferably, the objective optimization model includes: obtaining the optimal joint policy for all agents. To optimize the objective, the strategy Let M represent the probability distribution of agents choosing joint action a in state s, where M represents the number of agents. It is an independent policy of agent m; using the standard Q-learning method, by executing From the current time slot Starting from state s, construct a target optimization model by maximizing the cumulative reward.

[0084] Example 11, this example further extends the examples 1-10, the target optimization model includes:

[0085]

[0086]

[0087] in, Represents the objective optimization model; Represents the strategy function; Expressing expectations; Represents positive infinity; Indicates the discount factor; Indicates from time slot The number of subsequent steps after the initial step; Indicates the agent in the time slot Observed state Next action The rewards received; Indicates time slot The environmental conditions; Indicates time slot Joint actions; Represents the joint Q-function; Indicates time slot The global reward; M represents the number of agents; This indicates that the link corresponding to agent m is in the time slot. The reward; This indicates that agent m observes the state. Next action The value of.

[0088] In this embodiment, step S5 constructs a target optimization model with the goal of maximizing the cumulative discount reward of all agents. By focusing on obtaining the optimal joint strategy of all agents, it considers both the reward of the current time slot and the cumulative reward of the future, which can avoid the global suboptimal problem caused by the optimization of a single agent, guide the agents to learn collaboratively, and thus adapt to the dynamic characteristics of the integrated air-space-ground network. This provides a clear direction for the training of agents and helps them learn strategies that enable the overall communication system to maintain high efficiency, stability and anti-interference performance in the long term.

[0089] S6: Train the agents using a distributed Q-learning algorithm based on similarity metrics and a constructed objective optimization model so that each agent learns the optimal communication strategy.

[0090] Preferably, training the agent using a distributed Q-learning algorithm based on similarity metrics and a constructed target optimization model includes:

[0091] S61: Initialize the Q table, where row indices represent states, column indices represent actions, and values ​​in the Q table represent the Q values ​​corresponding to the state-action pairs.

[0092] S62: Each agent stores its observed historical states into a statistical state set. And it continues to expand as transmission progresses;

[0093] S63: Each agent will update its current state and statistical state set. The similarity between historical states is calculated, and historical states with a similarity greater than a set threshold are used as estimated states. ;

[0094] S64: Pair the estimated state with the actions in the action space to construct estimated state-action pairs, and update the Q-table based on the similarity between the estimated state-action pairs and the previous time slot state-action pairs. The update strategy is as follows:

[0095]

[0096] in, This indicates the updated Q table. Indicates the previous time slot Observed environment Next action The Q value obtained later; Indicates the learning rate. Indicates the discount factor; This indicates that an action is being performed in the current state S. The reward value obtained later; This indicates taking the maximum value; Indicates that the next state has been observed. Next action The Q value obtained later; Represents the estimated set of state-action pairs; This indicates an estimated state-action pair;

[0097] S65: Based on the Q-table update strategy and the target optimization model, update the agent's policy function. The update strategy is as follows:

[0098]

[0099] in, This represents the agent's optimal strategy; Indicates the maximum value; Represents the action space; This indicates that agent m observes the state. Next action The value of.

[0100] Example 12 is a further extension of Examples 1-10. In this example, each agent calculates its current state and statistical state set. The similarity calculation of historical states and the estimation of the similarity between state-action pairs and the previous time slot state-action pairs include:

[0101]

[0102]

[0103]

[0104] in, Indicates the similarity between environmental action pairs; Indicates the weighting coefficient; Indicates the current environmental state; Indicates the next environmental state; Represents the environmental state space; Indicates an indicator function; and This indicates the action of selecting the transmission channel corresponding to the current state and the next state; and This indicates the action for power selection corresponding to the current state and the next state; Indicates the similarity between states; Indicates the similarity between actions; This indicates the similarity between state-action pairs.

[0105] In this embodiment, step S6 uses a distributed Q-learning algorithm based on similarity measurement and a target optimization model to train the agent. By initializing the Q-table, storing historical states, calculating state similarity to obtain estimated states, and combining estimated state actions to update the Q-table and policy function, the historical state information can be fully utilized to reduce redundant calculations and lower the computational complexity of reinforcement learning. At the same time, it accelerates learning convergence, enabling each agent to efficiently and collaboratively learn the optimal communication strategy adapted to the dynamic interference environment and time-varying topology of the integrated air-space-ground network in a distributed scenario. This provides support for achieving lightweight deployment and ensuring continuous and stable communication.

[0106] S7: The agent outputs the optimal transmission channel and transmission power between the sender and receiver based on the observed state and the learned optimal communication strategy.

[0107] Another aspect of the present invention provides a reinforcement learning-based integrated air-space-ground network anti-jamming communication system, characterized in that the system includes a memory and a processor; the memory is used to store an application program; the processor is used to run the application program and execute the reinforcement learning-based integrated air-space-ground network anti-jamming communication method as described in any one of claims 1 to 6.

[0108] A computer storage medium, characterized in that the computer storage medium stores a remote monitoring program, which, when executed by a processor, implements a reinforcement learning-based anti-interference communication method for integrated air-space-ground networks as described in any one of claims 1 to 6.

[0109] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0110] This invention constructs intelligent agents for communication links between UAVs and ground users, and between UAVs and low-Earth orbit satellites. It constructs a state space by combining the average received power of the channel, an action space by combining the transmission channel and power set, and a reward function by combining the received signal interference plus noise ratio. With the goal of maximizing the cumulative discount reward, it uses a distributed Q-learning algorithm based on similarity metric to train the intelligent agents to learn the optimal communication strategy. This can adapt to the environment of dynamic changes in interference sources and time-varying network topology in the integrated air-space-ground network, improve the real-time performance and effectiveness of anti-interference communication, reduce the computational complexity of reinforcement learning, accelerate learning convergence, and achieve lightweight deployment, thereby ensuring continuous and stable communication between satellites, UAVs, and ground users.

[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A reinforcement learning-based anti-jamming communication method for an integrated air-space-ground network, wherein the integrated air-space-ground network includes a UAV, ground users, low-Earth orbit satellites, an airborne jammer, and a ground jammer, wherein the airborne jammer jams the communication link between the UAV and the low-Earth orbit satellites, and the ground jammer jams the communication link between the UAV and the ground users, characterized in that, The communication method includes: S1: Construct intelligent agents for the communication links between drones and ground users, and for the communication links between drones and low-Earth orbit satellites; S2: Construct the state space of the agent based on the average received power of each available channel between the transmitter and receiver in the current time slot; where drones, ground users, and low-Earth orbit satellites can all serve as transmitters or receivers; S3: Construct the action space of the agent based on the set of transmission channels and the set of transmission power between the sender and receiver; S4: Construct the agent's reward function based on the interference plus noise ratio of the received signal in the link between the transmitter and receiver; S5: Construct a target optimization model with the goal of maximizing the cumulative discount reward of all agents; S6: Train the agents using a distributed Q-learning algorithm based on similarity metrics and a constructed objective optimization model so that each agent learns the optimal communication strategy. S7: The agent outputs the optimal transmission channel and transmission power between the sender and receiver based on the observed state and the learned optimal communication strategy.

2. The anti-interference communication method for an integrated air-space-ground network based on reinforcement learning according to claim 1, characterized in that, The state space of the constructed intelligent agent includes: S21: Constructing a channel gain model between UAVs and low-Earth orbit satellites based on the effects of UAV beamforming gain, spatial loss, and small-scale fading: S22: Construct a channel gain model between the UAV and ground users based on line-of-sight loss and non-line-of-sight loss as defined in 3GPP; S23: Calculate the power spectral density function between the UAV and the low-Earth orbit satellite based on the channel gain model between the UAV and the low-Earth orbit satellite and the channel gain of the airborne jammer; S24: Calculate the power spectral density function between the UAV and the ground user based on the channel gain model between the UAV and the ground user and the channel gain of the ground jammer; S25: Discretize the power spectral density function between the UAV and the ground user, and calculate the average received power of each available channel for the UAV and the ground user in the current time slot; S26: Discretize the power spectral density function between the UAV and the low-Earth orbit satellite, and calculate the average received power of each available channel of the UAV and the low-Earth orbit satellite in the current time slot; S27: Construct the state space of the agent based on the average received power of each available channel of the UAV and ground user in the current time slot and the average received power of each available channel of the UAV and low-orbit satellite in the current time slot.

3. The anti-interference communication method for an integrated air-space-ground network based on reinforcement learning according to claim 1, characterized in that, The action space for constructing the intelligent agent includes: the intelligent agent's action space in the current time slot. Selected transmission channel and transmission power.

4. The anti-interference communication method for an integrated air-space-ground network based on reinforcement learning according to claim 1, characterized in that, The reward function for constructing the intelligent agent includes: S41: Calculate the interference-to-noise ratio of the received signal at the receiver between the UAV and the low-Earth orbit satellite based on the channel gain model between the UAV and the low-Earth orbit satellite, and the channel gain model of the airborne jammer. S42: Calculate the received signal interference plus noise ratio between the UAV and the ground user based on the channel gain model between the UAV and the ground user, and the channel gain model of the ground jammer; S43: Based on the interference plus noise ratio of the received signal at the receiving end, construct the reward function of the agent with the criterion that the smaller the transmission power between the transmitting end and the receiving end, the greater the reward value.

5. The anti-interference communication method for an integrated air-space-ground network based on reinforcement learning according to claim 1, characterized in that, The objective optimization model includes: obtaining the optimal joint strategy for all agents. To optimize the objective, the strategy Let M represent the probability distribution of agents choosing joint action a in state s, where M represents the number of agents. It is an independent policy of agent m; using the standard Q-learning method, by executing From the current time slot Starting from state s, construct a target optimization model by maximizing the cumulative reward.

6. The anti-interference communication method for an integrated air-space-ground network based on reinforcement learning according to claim 1, characterized in that, The process of training the agent using a distributed Q-learning algorithm based on similarity metrics and a constructed target optimization model includes: S61: Initialize the Q table, where row indices represent states, column indices represent actions, and values ​​in the Q table represent the Q values ​​corresponding to the state-action pairs. S62: Each agent stores its observed historical states into a statistical state set. And it continues to expand as transmission progresses; S63: Each agent will update its current state and statistical state set. The similarity between historical states is calculated, and historical states with a similarity greater than a set threshold are used as estimated states. ; S64: Pair the estimated state with the actions in the action space to construct estimated state-action pairs, and update the Q-table based on the similarity between the estimated state-action pairs and the previous time slot state-action pairs. The update strategy is as follows: in, This indicates the updated Q table. Indicates the previous time slot Observed environment Next action The Q value obtained later; Indicates the learning rate. Indicates the discount factor; This indicates that an action is being performed in the current state S. The reward value obtained later; This indicates taking the maximum value; Indicates that the next state has been observed. Next action The Q value obtained later; Represents the estimated set of state-action pairs; This indicates an estimated state-action pair; S65: Based on the Q-table update strategy and the target optimization model, update the agent's policy function. The update strategy is as follows: in, This represents the agent's optimal strategy; Indicates the maximum value; Represents the action space; This indicates that agent m observes the state. Next action The value of.

7. A space-air-ground integrated network anti-jamming communication system based on reinforcement learning, characterized in that, The system includes a memory and a processor; the memory is used to store an application program; the processor is used to run the application program and execute the reinforcement learning-based air-space-ground integrated network anti-interference communication method as described in any one of claims 1 to 6.

8. A computer storage medium, characterized in that, The computer storage medium stores a remote monitoring program, which, when executed by the processor, implements a reinforcement learning-based anti-interference communication method for integrated air-space-ground networks as described in any one of claims 1 to 6.