An asymptotically optimal distributed multi-agent DRL anti-interference method

By modeling the anti-interference problem as a local interactive Markov game and employing the distributed multi-agent DRL algorithm, the dynamic and resource-constrained problems of anti-interference methods in LEO satellite networks are solved, and asymptotically optimal anti-interference performance is achieved.

CN120357955BActive Publication Date: 2025-10-31PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510812282.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-31
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

Traditional anti-jamming methods are difficult to adapt to the dynamic nature and resource constraints of LEO satellite networks, and the complex electromagnetic environment increases the design difficulty, making it difficult to deploy existing methods on LEO satellites.

Method used

The anti-interference problem is modeled as a local interactive Markov game. A distributed multi-agent DRL algorithm is adopted, and an asymptotically optimal distributed multi-agent DRL anti-interference algorithm (DMADRLA) is designed through an architecture of offline training and online execution to autonomously optimize the frequency-power allocation strategy.

Benefits of technology

It effectively balances the training cost and optimization performance of the anti-interference model, achieving good anti-interference effect and is suitable for resource-constrained LEO satellite networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120357955B_ABST
    Figure CN120357955B_ABST
Patent Text Reader

Abstract

This invention relates to the field of satellite communication technology, specifically disclosing an asymptotically optimal distributed multi-agent DRL anti-interference method. It includes: Step S01: Constructing an anti-interference scenario for the satellite communication downlink, providing antenna and signal models for the interfering satellite, the communication satellite, and the ground user, obtaining the received signal-to-noise ratio (SNR) of the ground user, calculating ground user satisfaction, and constructing an optimization problem and constraints for anti-interference decision-making based on the ground user satisfaction; Step S02: Modeling the anti-interference decision-making optimization problem as a locally interactive Markov game model, and providing a reward design; Step S03: Using the distributed multi-agent DRL anti-interference algorithm, obtaining the equilibrium solution of the locally interactive Markov game model, enabling the satellite to autonomously acquire an anti-interference strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of satellite communication technology, and specifically to an asymptotically optimal distributed multi-agent DRL anti-interference method. Background Technology

[0002] Low Earth Orbit (LEO) satellite constellations have attracted widespread attention due to their advantages such as low latency, high data transmission rates, and good ground coverage, and their rapid large-scale deployment is triggering a transformation in global communication architecture. However, the exposed nature of satellites and the openness of wireless channels make satellite communications vulnerable to electromagnetic interference attacks. Therefore, researching anti-jamming technologies for LEO satellite communication systems is of great significance for improving communication reliability.

[0003] However, due to the dynamic nature and limited resources of satellite networks, traditional anti-jamming methods are difficult to utilize directly, facing three main challenges. First, time-varying satellite-to-ground network topology. The continuous link reconfiguration caused by the high-speed movement of LEO satellites makes traditional anti-jamming methods inadequate. Simultaneously, the complexity of the network topology within the constellation further exacerbates the need for timely anti-jamming algorithms. Second, resource constraints. Limited by the computing and storage capabilities of the onboard platform, traditional high-complexity algorithms are difficult to deploy on LEO satellites, and anti-jamming strategy design must adhere to the Pareto optimality criterion of "algorithm efficiency - resource consumption." Third, complex electromagnetic environment. In addition to external malicious interference attacks, the LEO constellation also faces internal beam interference at the same frequency, and these interference relationships dynamically adjust with changes in satellite orbital positions, increasing the design difficulty of anti-jamming methods in satellite communication systems. Summary of the Invention

[0004] To address the aforementioned problems, this invention aims to provide an asymptotically optimal distributed multi-agent DRL anti-jamming method. For the dynamic anti-jamming scenario of the LEO satellite constellation downlink, the anti-jamming problem is modeled as a locally interactive Markov game, and it is proven that this game belongs to the exact potential energy game (EPG) and has at least one Nash equilibrium (NE). Secondly, to obtain the equilibrium solution, an asymptotically optimal distributed multi-agent DRL anti-jamming algorithm (DMADRLA) based on an "offline training-online execution" architecture is proposed. Finally, simulations verify that the proposed DMADRLA scheme can effectively balance the training cost and optimization performance of the anti-jamming model, achieving good anti-jamming performance.

[0005] To achieve the objective of this invention, the technical solution is: an asymptotically optimal distributed multi-agent DRL anti-interference method, comprising:

[0006] Step S01: Construct an anti-interference scenario for the satellite communication downlink, provide antenna and signal models for the interfering satellite, communication satellite, and ground users, obtain the received signal-to-noise ratio of the ground users, calculate the ground user satisfaction, and construct an optimization problem and constraints for anti-interference decision based on the ground user satisfaction.

[0007] Step S02: Model the anti-interference decision optimization problem as a local interactive Markov game model and provide a reward design;

[0008] Step S03: Using the distributed multi-agent DRL anti-interference algorithm, the equilibrium solution of the local interactive Markov game model is obtained, enabling the satellite to autonomously acquire an anti-interference strategy.

[0009] Preferably, in step S01, the antenna model is used to obtain the transmitting antenna gain of the jamming satellite, the transmitting antenna gain of the communication satellite, and the receiving antenna gain of the ground user.

[0010] Preferably, in step S01, the signal model is used to calculate the received signal-to-noise ratio of the ground user based on the external interference from interfering satellites and the internal co-channel interference from adjacent communication satellites.

[0011] Preferably, in step S01, the signal-to-noise ratio received by the ground user is:

[0012]

[0013] Where, p n (t) represents link l in time slot t. n Satellite launch power, h n (t) represents link l in time slot t. n Channel gain, For co-channel interference between adjacent satellites, For interference signals, N n (t) represents the environmental noise power, p u (t) represents link l in time slot t. u Satellite launch power, For time slot t link Channel gain, To determine the link l in time slot t n and l u Whether to select the same channel function, f n (t) represents link l in time slot t. n Satellite communication channels.

[0014] Preferably, in step S01, the formula for calculating the ground user satisfaction is:

[0015]

[0016] The achievable capacity is: r n (f n (t),p n (t))=B f log2(1+γ n (f n (t),p n (t)));

[0017] Among them, B f Where is the channel bandwidth, and c is the sensitivity of the user to demand. For predefined threshold, f n (t) represents link l in time slot t. n Satellite communication channels.

[0018] Preferably, in step S01:

[0019] The optimization problem is expressed as:

[0020] (P):

[0021] The constraints include:

[0022] γ∈(0,1), f n (t)∈F C p n (t)∈P C ;

[0023] Among them, s n (f n (t),p n (t) represents ground users Satisfaction, γ is the time discount factor, p n (t) represents link l in time slot t. n Satellite launch power, f n (t) represents link l in time slot t. n satellite communication channels, P n (t) represents link l in time slot t. n power cost, F C For the set of available channels, P C This is a set of transmit power.

[0024] Preferably, in step S02, the reward design takes into account the satisfaction of individuals and neighbors, and integrates the weighted anti-interference effect of the current time slot and the historical time slot, so that all participants maximize the anti-interference effect of the current time slot and the cumulative anti-interference effect of the entire system cycle.

[0025] Furthermore, the cumulative anti-interference effect uses cumulative long-term reward as the utility function for each participant, and the utility function is expressed as:

[0026]

[0027] Where a represents the strategies of all participants, γ represents the time discount factor, and s represents the strategies of all participants. n (f n (t),p n (t) represents ground users satisfaction, s m (f m (t),p m (t) represents ground users satisfaction, p n (t) represents link l in time slot t. n Satellite launch power, P n (t) represents link l in time slot t. n Power cost, p m (t) represents link l in time slot t. m Satellite launch power, P m (t) represents link l in time slot t. m The power cost, β is the weight of the neighbor reward, For participants to gather, For the number of participants, Let n be the set of neighboring users of participant n.

[0028] Preferably, in step S03, the anti-interference algorithm of the distributed multi-agent DRL includes an offline training phase, which includes:

[0029] Agent network initialization and experience collection: Initialize agent network parameters, use ε-greedy exploration strategy to interact with the environment, collect agent state and decision information under neighbor relationship constraints, and store it in the experience replay buffer;

[0030] Target optimization and error balancing: Batch sampling data from the experience replay buffer, calculation of the target Q value, and optimization strategy through utility function in combination with time series difference error;

[0031] Parameter update and collaborative training: Parameters are updated through gradient descent and shared among neighboring agents. The target network parameters are synchronized periodically, and the optimized parameters are uploaded to the corresponding communication satellite.

[0032] Preferably, in step S03, the anti-interference algorithm of the distributed multi-agent DRL further includes an online execution phase, which includes:

[0033] Pre-trained parameter loading and real-time decision execution: The agent loads the optimized parameters, selects the best strategy based on local observations, and stores the experience data in the local buffer after executing the action;

[0034] Incremental parameter update and local collaborative optimization: At fixed intervals, small batches of data are sampled from the buffer to calculate the target value. Based on the temporal difference error, the network parameters are updated with a constrained step size, and the updated parameters are shared with neighboring agents to achieve distributed collaborative optimization.

[0035] The beneficial effects of the above technical solution are as follows:

[0036] This invention provides an asymptotically optimal distributed multi-agent DRL anti-jamming method for the dynamic anti-jamming scenario of the LEO satellite constellation downlink. It models the anti-jamming problem as a locally interactive Markov game and proves that this game belongs to the exact potential energy game (EPG) and has at least one Nash equilibrium (NE). Secondly, to obtain the equilibrium solution, an asymptotically optimal distributed multi-agent DRL anti-jamming algorithm (DMADRLA) based on an "offline training-online execution" architecture is proposed. Finally, simulations verify that the proposed DMADRLA scheme can effectively balance the training cost and optimization performance of the anti-jamming model, achieving good anti-jamming performance. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 A flowchart of an asymptotically optimal distributed multi-agent DRL anti-interference method provided for one embodiment of the present invention;

[0039] Figure 2 An anti-interference communication scenario diagram in a low-Earth orbit constellation provided as an embodiment of the anti-interference method of the present invention;

[0040] Figure 3 A downlink interference model diagram of an anti-interference method provided in an embodiment of the present invention;

[0041] Figure 4 A comparison chart of convergence results of an anti-interference method provided in an embodiment of the present invention;

[0042] Figure 5 A comparative graph showing the impact of changes in traffic demand on the average satisfaction level of an anti-interference method provided in an embodiment of the present invention. Detailed Implementation

[0043] The embodiments of this application will be described in further detail below. Obviously, the described embodiments are only a part of the embodiments of this application, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.

[0044] The terms “first,” “second,” etc. (if applicable) in the specification and claims are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion, such as a process, method, system, product, or apparatus that comprises a series of steps or units, not necessarily limited to those explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0045] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0046] Example 1

[0047] One embodiment of the present invention provides an asymptotically optimal distributed multi-agent DRL anti-interference method flow as follows: Figure 1 As shown, the process includes: Step S01: Constructing an anti-interference scenario for the satellite communication downlink, providing antenna and signal models for the interfering satellite, communication satellite, and ground users, obtaining the received signal-to-noise ratio (SNR) for ground users, calculating ground user satisfaction, and constructing an optimization problem and constraints for anti-interference decision-making based on the ground user satisfaction; Step S02: Modeling the anti-interference decision-making optimization problem as a locally interactive Markov game model and providing a reward design; Step S03: Using a distributed multi-agent DRL anti-interference algorithm to obtain the equilibrium solution of the locally interactive Markov game model, enabling the satellite to autonomously acquire an anti-interference strategy.

[0048] 1. Model building and optimization problem

[0049] Constructing an anti-interference communication scenario model

[0050] like Figure 2As shown, the anti-jamming communication scenario model constructed in this invention includes two independently operating Walker-configured LEO constellations deployed at different orbital altitudes: a malicious jamming constellation and a communication constellation. Within each constellation, adjacent satellites exchange information via inter-satellite links. The jamming satellites in the malicious jamming constellation employ narrow-beam, high-gain parabolic antennas to implement random-frequency, high-power suppression jamming on the downlink of ground users within ground region S. Frequency switching follows a Markov probability matrix P. The jamming strategy is planned by the ground control center and distributed to the jamming satellites through a ground gateway. The communication satellites in the communication constellation are equipped with multi-beam phased array antennas capable of dynamically adjusting frequency and power, thereby avoiding interference and ensuring that the communication quality of ground users is not affected.

[0051] The set of ground users is represented as LU = {lu1, lu2, ..., lu...} N The set of LEO communication constellations is represented as LEO. C ={leo c1 leo c2 ,...,leo cK Each communication satellite can form multiple independent beams, and each beam has a set of available channels called F. C ={f1,f2,…,f M The channel bandwidth is Bf, and the antenna transmit power set is P. C ={p1,p2,...,p L The interfering constellation is represented as LEO. J ={leo j1 leo j2 ,...,leo jK Its antenna transmit power is fixed at p. J Among them, the interfering satellites closest to the center of region S use the same frequency set F. J =F C To interfere. n (t) represents the ground user lu in time slot t. n The established downlink To interfere with satellites in time slot t for ground users lu n Interference links, For link l in time slot t u (t) For ground users lu n Co-channel interference.

[0052] The time axis is divided into equal time slots T = {t | t = 1, 2, ..., T}. Within each time slot, the transmitting antenna of the interfering satellite switches the interfering channel according to a Markov probability matrix. The communication satellite dynamically selects the optimal frequency-power combination based on real-time interference sensing and historical data analysis, thereby mitigating the adverse effects caused by the interfering satellite. The anti-interference method provided by this invention focuses on strategy optimization for LEO communication satellites under external malicious interference. By dynamically optimizing the frequency-power anti-interference strategy, it reduces the impact of interference and improves the average satisfaction of ground users.

[0053] Constructing an antenna model

[0054] The jamming satellite is equipped with a high-gain narrow-beam antenna conforming to the ITU-R F.699-8 standard. The gain of the jamming satellite's transmitting antenna is expressed as follows:

[0055]

[0056] in, To maliciously interfere with the link between the beam center and the interfered ground user off-axis angle between For antenna half-power beamwidth, To interfere with the maximum gain of the satellite antenna, D J Let λ be the diameter of the parabolic antenna, η = 0.55 be the aperture efficiency, λ be the wavelength, and the interference beam power be fixed at p. J .

[0057] For communication satellites, according to ITU-R S.1528, the transmit antenna gain of a communication satellite is expressed as follows:

[0058]

[0059] in, For interference links and downlink l u The off-axis angle between (t) and θ b For antenna half-power beamwidth, θ b =0.5θ 3dB , For the maximum gain of the communication satellite antenna, L s = -6.75dB is the intersection of the main beam and the near-end sidelobe mask, L F =0dBi represents the far-end sidelobe level, Y = 1.5θ b .

[0060] For terrestrial users, according to ITU-R S.465, the receiving antenna gain is expressed as:

[0061]

[0062] in, For interference links and downlink l n The off-axis angle between (t) For the maximum gain of the ground user antenna, D U Let η be the equivalent circular diameter of the ground user antenna, η = 0.55 be the aperture efficiency, and λ be the wavelength.

[0063] Constructing a signal model

[0064] Downlink interference model such as Figure 3 As shown, in time slot t, ground user lu n It is subject to external interference from interfering satellites and internal co-channel interference from neighboring communication satellites.

[0065] In response to external malicious interference, when the downlink l n When subjected to interference attacks, the interference signal Represented as:

[0066]

[0067] in, For interference signals, p j (t) represents the transmission power of the interference beam in time slot t. For link Channel gain, To determine the interference link and the communication link n Whether to select the same channel function To interfere with satellites and ground users n The distance between them, where λ is the wavelength. To maliciously interfere with satellites for ground users n Transmit antenna gain in direction For ground users lu n The gain of the receiving antenna in the direction of malicious interference with satellites.

[0068] Regarding internal co-channel interference, when multiple satellite links share the same channel, co-channel interference will occur. (Ground user lu) n The cumulative co-channel interference received is expressed as:

[0069]

[0070] Where, p u (t) represents the transmit power of the satellite beams adjacent to time slot t. For link Channel gain, To determine the interference link and communication link ln Whether to select the same channel function

[0071] For link distance, Interfering with communication satellites at ground users lu n Transmit antenna gain in direction For ground users lu n Receiver antenna gain in the direction of interfering communication satellites.

[0072] Therefore, ground user lu n The received signal-to-noise ratio (SNR) at that location can be expressed as:

[0073]

[0074] Where, p n (t) represents link l in time slot t. n Satellite launch power, h n (t) represents link l n Channel gain, This represents the maximum gain of the satellite transmitting antenna. For the user's receiving antenna maximum gain, ||d n (t)|| represents link l n Distance N n (t) represents the environmental noise power.

[0075] The achievable capacity is: r n (f n (t),p n (t))=B f log2(1+γ n (f n (t),p n (t)))

[0076] Among them, B f This refers to the channel bandwidth.

[0077] Satisfaction s n (f n (t),p n (t) is used to measure the arrival rate r. n (f n (t),p n (t)):

[0078]

[0079] Among them, parameter c controls the sensitivity of ground users to demand. For predefined thresholds.

[0080] The predefined threshold reflects the QoS requirements of ground users; a value close to 1 indicates near-satisfactory service, while a value close to 0 indicates unsatisfactory service.

[0081] Construct optimization problem

[0082] The optimization objective is to determine the optimal joint frequency and power allocation anti-interference scheme for the LEO satellite downlink, in order to minimize the impact of external malicious interference within time period T, i.e., maximize average ground user satisfaction.

[0083] The optimization problem is represented as:

[0084] (P):

[0085] Subject to the following constraints:

[0086] Time range: finite or infinite.

[0087] γ∈(0,1): Time discount factor.

[0088] s n (·): Ground user satisfaction function.

[0089] Decision variable: f n (t)∈F C For time slot t link l n satellite communication channel, p n (t)∈P C For time slot t link l n The satellite launch power is allocated to ensure that the frequency band and power constraints are maintained.

[0090] The optimization problem (P) balances ground user satisfaction s n (f n (t),p n (t) and power cost P n (t), where power cost is the normalized value of satellite transmit power relative to its maximum power, a dimensionless scalar ranging from 0 to 1, and a time discount factor γ is introduced to achieve consistent anti-interference performance. The above problem is a multivariate time-series decision problem, requiring a time-series optimization framework to realize the cumulative effect of variable decisions over consecutive time slots.

[0091] 2. Constructing an anti-interference game model

[0092] This invention provides a theoretical framework for the joint spectrum-power optimization anti-interference problem. The optimization problem (P) is modeled as a locally interactive Markov game (LIMG) model, which belongs to the exact potential energy game (EPG) and has a Nash equilibrium (NE).

[0093] Locally Interactive Markov Game Model (LIMG)

[0094] The interference-resistant game model is defined as a seven-tuple. in:

[0095] For state space, Including one's own satisfaction S n (t-1), Neighbor satisfaction Current interference state I j (t), the current total co-channel interference state I co (t).

[0096] Let n be the set of participants, where each participant n corresponds to a downlink l. n .

[0097] For joint action space, F C P is the set of available channels. C This refers to the set of transmit power. A single action is...

[0098] For a set of neighbors, it is defined as

[0099] The state transition function is denoted as F(s). ′ ∣s,a)=P(s(t+1)=s′∣s(t)=s,a(t)=a).

[0100] Let be the reward function, denoted as β is the weight of the neighbor reward.

[0101] γ is the time discount factor, 0 < γ < 1.

[0102] Design reward mechanism

[0103] The reward mechanism is designed to balance the satisfaction of individual users and their neighbors. The reward for participant n in time slot t is expressed as:

[0104]

[0105] Using cumulative long-term rewards as the utility function for each participant, it is defined as:

[0106]

[0107] Among them, a n={a1(t),a2(t),...,a N The utility function (t) integrates the weighted anti-interference effect of the current time slot and the historical time slot, ensuring that all participants maximize the anti-interference utility of the current time slot while also maximizing the cumulative anti-interference effect over the entire system cycle.

[0108] Equilibrium Analysis

[0109] By constructing the potential energy function:

[0110]

[0111] It was also proven that the constructed locally interactive Markov game is an exact potential game, satisfying:

[0112]

[0113] Guarantee the game There exists at least one pure policy Nash equilibrium (NE), i.e., a joint policy equilibrium. satisfy:

[0114] By establishing the equivalence between problem P and the potential function Φ, it is proven that the optimal solution to problem P is a game theory problem. The pure policy NE is used. To obtain the equilibrium solution of the game, a robust algorithm for asymptotically optimal distributed multi-agent DRL is proposed.

[0115] 3. Anti-interference algorithm for distributed multi-agent DRL (DMDRLA)

[0116] Based on the Locally Interactive Markov Game (LIMG) model, a DMDRLA algorithm is proposed, combining offline ground training and on-board incremental learning. It achieves this through a three-layer design: ground user neighborhood constraints, satellite-ground model synchronization, and autonomous online adaptation, making it well-suited for dynamic LEO satellite network topologies.

[0117] Offline training phase

[0118] In the offline training phase, Algorithm 1 employs a distributed multi-agent DRL architecture to train the neural network offline. The training process includes three steps: agent network initialization and experience collection, objective optimization and error balancing, and parameter update and collaborative training. First, the agent network is initialized, and an ε-greedy exploration strategy is used to collect agent state and decision information, which is stored in an experience replay buffer. The decision-making process is constrained by the neighbor relationships between agents within the satellite network. Second, batch sampling, objective value calculation, and local utility optimization are performed to balance temporal difference errors. Finally, network parameters are updated through gradient descent, and parameters are shared among neighboring agents. Periodic objective function synchronization ensures training stability. After training is completed, the optimized model parameters are uploaded to the corresponding LEO satellite for autonomous operation.

[0119] Algorithm 1, the interference-resistant distributed training in the LEO constellation, includes the following steps:

[0120] Step 1: Initialize the main network for all agents The parameter θ is used to create an experience playback buffer. Synchronize the target network parameters θ′←θ.

[0121] Step 2: Initialize the environment state s0 in each round and execute T-step sequential interactions; each agent explores randomly with probability ∈, otherwise selects an action. Execute joint operations Observation Rewards and new state s t+1 .

[0122] Step 3: Each agent bases its actions on neighbor information M. n Calculate the local reward r t n ;Will Store in the experience replay buffer

[0123] Step 4: Each round from Sample a batch of data of size b. n Calculate the target Q value Minimize loss function For θ n Perform gradient descent to update the main network parameters.

[0124] Step 5: The agent shares the updated parameter θ with its neighbors m∈Mn. n .

[0125] Step 6: At fixed intervals, soft update the target network by mixing coefficient τ = 0.01: θ′←τθ+(1-τ)θ′.

[0126] Step 7: Calculate the final optimized main network parameters θ * Uploaded to the satellite intelligent agent, supporting online anti-interference decision-making.

[0127] This algorithm utilizes the computing power of ground user terminals to independently complete model training, avoiding the computational burden on low-Earth orbit satellites. By restricting interactions between agents to a predefined neighborhood range, redundant coordination costs are significantly reduced. A parameter interaction mechanism between neighborhood nodes replaces the centralized network, achieving ordered collaboration in a distributed architecture. Within the EPG framework, it is proven that local optimization in a distributed multi-agent system can converge to the global optimum.

[0128] Online execution phase

[0129] During the online execution phase, the agent achieves fully decentralized real-time decision-making through a pre-trained neural network. As described in Algorithm 2, the agent first loads the pre-trained network parameters to operate independently without centralized coordination. In each time slot, the agent observes its own and its neighbors' states, selects a joint spectrum-power strategy that maximizes the function of its neural network, and executes these actions in the environment. Furthermore, the algorithm integrates an incremental information update mechanism, allowing the agent to periodically optimize its network parameters based on newly observed states and rewards without requiring complete retraining.

[0130] The incremental update for real-time decision-making and anti-interference in Algorithm 2 LEO constellation includes:

[0131] Step 1: Load all agent networks Training parameters θ *

[0132] Step 2: Each agent observes its local state. Select Action Execute action Afterwards, obtain the reward r t n and new status Experience Stored in local buffer This is used for subsequent incremental learning.

[0133] Step 3: Trigger incremental parameter updates at fixed time intervals from the local buffer. Extract small batches of empirical data and calculate the target value for each empirical (s,a,r,s'). Calculate the gradient of the loss function using time-series difference error. With constrained step size α t Update parameters Then the agent only shares the updated parameters with its neighbors m∈Mn.

[0134] In this algorithm, each agent makes independent decisions based solely on local observation information, avoiding the communication and computational overhead of centralized decision-making. An incremental parameter update mechanism ensures the adaptability of the pre-trained neural network in dynamic satellite network environments. These two technologies together enhance the real-time response capability of the space-ground collaborative system.

[0135] Convergence proof

[0136] The proposed anti-interference algorithm exhibits good convergence and asymptotic optimality, and it almost always converges to a neighborhood-constrained NE. Under the established EPG framework, the NE approximates the global optimum of the potential function, which can be proven through the compressibility of the local Bellman operator, the convergence analysis of stochastic approximation in distributed updates, the neighborhood-constrained Nash equilibrium mechanism based on parameter sharing, and the relationship analysis with the global optimum solution under the EPG framework.

[0137] Simulation results of the asymptotically optimal distributed multi-agent DRL anti-interference method provided by this invention

[0138] This invention presents a simulation analysis of the performance of the proposed asymptotically optimal distributed multi-agent DRL algorithm. The simulation involves 50 ground users randomly distributed within a potential interference area S of 500km × 500km, with the satellite downlink subjected to suppressive interference from an external malicious jamming constellation. The relevant simulation parameters are shown in Table 1.

[0139] Table 1 Simulation Parameters

[0140]

[0141] Figure 4 Convergence performance analyses are presented for three different anti-interference methods: Centralized Training with Global Information (CTAGI), Independent Decision Training without Information Interaction (IIIDT), and the proposed method (DMDRLA). CTAGI exhibits superior convergence performance due to its global information awareness capability, but its practical application is constrained by satellite computing resources. IIIDT fails to formulate a coordinated anti-interference strategy, leading to a significant decrease in convergence performance. The proposed method achieves approximately 91% of the convergence performance of CTAGI while requiring only about 8% of the link overhead, demonstrating that the proposed method maintains its coordination advantage through local information interaction and validating the effectiveness of DMDRLA.

[0142] Figure 5The impact of ground user traffic demands on the anti-interference performance of the algorithms was compared and analyzed. CTAGI achieved the best anti-interference performance, DMDRLA was slightly lower, and IIIDT, due to its completely independent training framework, had the lowest anti-interference performance. When LU traffic demand exceeded 80Mbps, IIIDT's anti-interference performance significantly decreased, while DMDRLA's performance remained relatively stable. Experiments confirmed that DMDRLA can find a balance between neural network performance optimization and training overhead, making it suitable for anti-interference scenarios in resource-constrained LEO satellite networks.

[0143] In summary, this invention proposes an asymptotically optimal distributed multi-agent DRL anti-jamming method for LEO satellite communication systems. First, the anti-jamming problem is modeled as a locally interactive Markov game, and it is proven to be an exact potential game with at least one pure policy NE. Next, based on an "offline training-online execution" architecture, a DMDRLA algorithm is proposed. Finally, simulation results verify that the proposed algorithm can effectively balance the network's anti-jamming performance and training overhead, enabling the satellite to autonomously optimize its frequency-power allocation strategy and exhibiting good anti-jamming performance.

[0144] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. For those skilled in the art, other variations or modifications can be made based on the above description. It is impossible to exhaustively list all the implementation methods here. All obvious variations or modifications derived from the technical solutions of the present invention are still within the protection scope of the present invention.

Claims

1. An asymptotically optimal distributed multi-agent DRL anti-interference method, characterized in that, include: Step S01: Construct an anti-interference scenario for the satellite communication downlink, provide antenna and signal models for the interfering satellite, communication satellite, and ground users, obtain the received signal-to-noise ratio of the ground users, calculate the ground user satisfaction, and construct an optimization problem and constraints for anti-interference decision based on the ground user satisfaction. Step S02: Model the anti-interference decision optimization problem as a locally interactive Markov game and provide a reward design. The reward design takes into account the satisfaction of individuals and neighbors, and integrates the weighted anti-interference effect of the current time slot and historical time slots, so that all participants maximize the anti-interference utility of the current time slot while also maximizing the cumulative anti-interference effect over the entire system cycle. The cumulative anti-interference effect uses the cumulative long-term reward as the utility function of each participant, and the utility function is expressed as: Where a represents the strategies of all participants, γ represents the time discount factor, and s represents the strategies of all participants. n (f n (t),p n (t) represents ground user lμ n satisfaction, s m (f m (t),p m (t) represents ground user lμ m satisfaction, f n (t) represents link l in time slot t. n satellite communication channels, p n (t) represents link l in time slot t. n Satellite launch power, P n (t) represents link l in time slot t. n power cost, f m (t) represents link l in time slot t. m satellite communication channels, p m (t) represents link l in time slot t. m Satellite launch power, P m (t) represents link l in time slot t. m The power cost, β is the weight of the neighbor reward, For participants to gather, For the number of participants, Let n be the set of neighboring users of participant n; Step S03: Using the distributed multi-agent DRL anti-interference algorithm, the equilibrium solution of the local interactive Markov game model is obtained, enabling the satellite to autonomously acquire an anti-interference strategy.

2. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 1, characterized in that, In step S01, the antenna model is used to obtain the transmitting antenna gain of the jamming satellite, the transmitting antenna gain of the communication satellite, and the receiving antenna gain of the ground user.

3. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 1, characterized in that, In step S01, the signal model is used to calculate the received signal-to-noise ratio of the ground user based on the external interference from interfering satellites and the internal co-channel interference from adjacent communication satellites.

4. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 1, characterized in that, In step S01, the received signal-to-noise ratio of the ground user is: Where, p n (t) represents link l in time slot t. n Satellite launch power, h n (t) represents link l in time slot t. n Channel gain, For co-channel interference between adjacent satellites, For interference signals, N n (t) represents the environmental noise power, p u (t) represents link l in time slot t. u Satellite launch power, For time slot t link Channel gain, link For link l u Corresponding satellite to ground user lu n Links between To determine the link l in time slot t n and l u Whether to select the same channel function, f n (t) represents link l in time slot t. n Satellite communication channels.

5. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 4, characterized in that, In step S01, the formula for calculating the ground user satisfaction is: The achievable capacity is: r n (f n (t),p n (t))=B f log2(1+γ n (f n (t),p n (t))); Among them, B f Where is the channel bandwidth, and c is the sensitivity of the user to demand. For predefined thresholds.

6. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 1, characterized in that, In step S01: The optimization problem is expressed as: The constraints include: γ∈(0,1)、f n (t)∈F C 、p n (t)∈P C ; Among them, s n (f n (t),p n (t) represents ground user lμ n Satisfaction, γ is the time discount factor, p n (t) represents link l in time slot t. n satellite launch power, F C For the set of available channels, P C This is a set of transmit power.

7. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 1, characterized in that, In step S03, the anti-interference algorithm of the distributed multi-agent DRL includes an offline training phase, which includes: Agent network initialization and experience collection: Initialize agent network parameters, use ε-greedy exploration strategy to interact with the environment, collect agent state and decision information under neighbor relationship constraints, and store it in the experience replay buffer; Target optimization and error balancing: Batch sampling data from the experience replay buffer, calculation of the target Q value, and optimization strategy through utility function in combination with time series difference error; Parameter update and collaborative training: Parameters are updated through gradient descent and shared among neighboring agents. The target network parameters are synchronized periodically, and the optimized parameters are uploaded to the corresponding communication satellite.

8. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 7, characterized in that, In step S03, the anti-interference algorithm of the distributed multi-agent DRL further includes an online execution phase, which includes: Pre-trained parameter loading and real-time decision execution: The agent loads the optimized parameters, selects the best strategy based on local observations, and stores the experience data in the local buffer after executing the action; Incremental parameter update and local collaborative optimization: At fixed intervals, small batches of data are sampled from the buffer to calculate the target value. Based on the temporal difference error, the network parameters are updated with a constrained step size, and the updated parameters are shared with neighboring agents to achieve distributed collaborative optimization.

Citation Information

Patent Citations

  • LEO satellite access switching algorithm based on multi-agent deep cycle Q network

    CN117040588A

  • Ocean communication on-demand service coverage method and device based on low orbit satellite hopping beam

    CN119232236A