Aasymptotically optimal distributed multi-agent DRL anti-interference method

By modeling the anti-interference problem as a local interactive Markov game and using a distributed multi-agent DRL algorithm, the problems of dynamicity and resource limitation in the LEO satellite communication system are solved, and the frequency-power allocation is achieved independently optimized, which improves the anti-interference performance and communication reliability.

CN120357955AActive Publication Date: 2025-07-22PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510812282.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-22
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

Traditional anti-interference methods are difficult to adapt to the dynamic nature, resource limitations and complex electromagnetic environment of LEO satellite communication systems. Especially under the limited network topology and resource of time-varying satellites, it is difficult to effectively deal with external malicious interference and internal synchronous interference.

Method used

The anti-interference problem is modeled as a local interactive Markov game, and a distributed multi-agent deep reinforcement learning (DRL) algorithm is adopted. Through offline training and online execution architecture, the asymptotic optimal distributed multi-agent DRL anti-interference algorithm (DMADRLA) is designed to independently optimize the frequency-power distribution strategy.

Benefits of technology

In the LEO satellite constellation downlink, effective anti-interference performance is achieved, training costs and optimization performance are balanced, dynamic topological changes are adapted to improve the satisfaction of ground users and communication reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120357955A_ABST
    Figure CN120357955A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of satellite communication, and particularly discloses an asymptotically optimal distributed multi-agent DRL anti-interference method. Comprising the following steps: S01, constructing an anti-interference scene of a satellite communication downlink, giving antenna models and signal models of an interference satellite, a communication satellite and a ground user, obtaining a receiving signal-to-noise ratio of the ground user for calculating the satisfaction degree of the ground user, and constructing an optimization problem and constraint of an anti-interference decision according to the satisfaction degree of the ground user; step S02, modeling the anti-interference decision optimization problem into a local interaction Markov game model, and giving award design; and S03, solving an equilibrium solution of the local interaction Markov game model by adopting a distributed multi-agent DRL anti-interference algorithm, so that the satellite autonomously obtains an anti-interference strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of satellite communication, and in particular to an asymptotically optimal distributed multi-agent DRL anti-jamming method. Background Art

[0002] Low Earth Orbit (LEO) satellite constellations have attracted wide attention from all walks of life due to their advantages such as low latency, high data transmission rate, and good ground coverage ability. Their rapid large-scale deployment is triggering a revolution in the global communication architecture. However, due to the exposure of satellites and the openness of wireless channels, satellite communication is vulnerable to electromagnetic interference attacks. Therefore, studying anti-jamming technologies for LEO satellite communication systems is of great significance for improving communication reliability.

[0003] However, due to the dynamics and resource limitations of satellite networks, traditional anti-jamming methods are difficult to directly apply, mainly facing the following three challenges. First, time-varying satellite-ground network topology. The continuous link reconstruction caused by the high-speed movement of LEO satellites makes it difficult for traditional anti-jamming methods to adapt. At the same time, the complexity of the network topology in the constellation further exacerbates the need for the timeliness of anti-jamming algorithms. Second, resource limitations. Limited by the computing power and storage capacity of the on-board platform, traditional high-complexity algorithms are difficult to deploy on LEO satellites, and anti-jamming strategy design must adhere to the Pareto optimal criterion of "algorithm efficiency-resource consumption". Third, complex electromagnetic environment. In addition to external malicious interference attacks, LEO constellations also face co-channel interference between internal beams, and these interference relationships are dynamically adjusted with the change of satellite orbital positions, increasing the design difficulty of anti-jamming methods in satellite communication systems. Summary of the Invention

[0004] Aiming at the above problems, the purpose of the present invention is to provide an asymptotically optimal distributed multi-agent DRL anti-jamming method. For the dynamic anti-jamming scenario of the downlink of LEO satellite constellations, the anti-jamming problem is modeled as a local interaction Markov game, and it is proved that this game belongs to an exact potential game (EPG) and at least one Nash equilibrium (NE) exists. Secondly, to obtain the equilibrium solution, an asymptotically optimal distributed multi-agent DRL anti-jamming algorithm (DMADRLA) is proposed based on the "offline training-online execution" architecture. Finally, simulations verify that the proposed DMDRLA scheme can effectively balance the training cost and optimization performance of the anti-jamming model and can obtain good anti-jamming performance.

[0005] To achieve the purpose of the present invention, the technical solution of the present invention is: an asymptotically optimal distributed multi-agent DRL anti-jamming method, including: Step S01: Construct an anti-jamming scenario for the satellite communication downlink, give the antenna models and signal models of the interfering satellite, communication satellite, and ground user, obtain the received signal-to-noise ratio of the ground user for calculating the ground user satisfaction, and construct an optimization problem and constraints for anti-jamming decision-making according to the ground user satisfaction. Step S02: Model the anti-jamming decision optimization problem as a local interactive Markov game model and give the reward design. Step S03: Use the anti-jamming algorithm of distributed multi-agent DRL to obtain the equilibrium solution of the local interactive Markov game model, so that the satellite can autonomously obtain the anti-jamming strategy.

[0006] Preferably, in step S01, the antenna model is used to obtain the transmitting antenna gain of the interfering satellite, the transmitting antenna gain of the communication satellite, and the receiving antenna gain of the ground user.

[0007] Preferably, in step S01, the signal model is used to calculate the received signal-to-noise ratio of the ground user according to the external interference from the interfering satellite and the internal co-channel interference from the adjacent communication satellite received by the ground user.

[0008] Preferably, in step S01, the received signal-to-noise ratio of the ground user is: ; ; where is the satellite transmission power in time slot t link , is the channel gain in time slot t link , is the co-channel interference of adjacent satellites, is the interference signal, is the environmental noise power, is the satellite transmission power in time slot t link , is the channel gain in time slot t link , is a function for judging whether the same channel is selected in time slot t link and .

[0009] Preferably, in step S01, the calculation formula of the ground user satisfaction is: ; where $\rho$ is the received signal-to-noise ratio of ground users, $c$ is used to control the sensitivity of users to demands, is a predefined threshold.

[0010] Preferably, in step S01: The optimization problem is expressed as: ; The constraints include: , , , ; Among them, is the satisfaction of ground user , is the time discount factor, is the satellite transmission power of time slot t link , is the satellite transmission power of time slot t link 's satellite communication channel, is the set of available channels, is the set of transmission powers.

[0011] Preferably, in step S02, the reward design takes into account the satisfaction of both individuals and neighbors, and comprehensively considers the weighted anti-interference effects of the current time slot and historical time slots, so that all participants can maximize the anti-interference utility of the current time slot while also maximizing the cumulative anti-interference effect throughout the system cycle.

[0012] Furthermore, the cumulative anti-interference effect uses the cumulative long-term reward as the utility function for each participant, and the utility function is expressed as: ; Among them, is the strategy of all participants, is the time discount factor, is the satisfaction of ground user , is the satisfaction of ground user , is the satellite transmission power of time slot t link , is the satellite transmission power of time slot t link 's, $\beta$ is the weight of neighbor rewards, $N$ is the set of participants, is participant n 's set of neighbor users.

[0013] Preferably, in step S03, the anti-interference algorithm of the distributed multi-agent DRL includes an offline training phase, and the offline training phase includes: Agent network initialization and experience collection: Initialize the agent network parameters, use the ε-greedy exploration strategy for environmental interaction, collect agent states and decision information under the constraint of neighbor relationships, and store them in the experience replay buffer; Target optimization and error balance: Batch sample data from the experience replay buffer, calculate the target Q value, combine the temporal difference error, and optimize the strategy through the utility function; Parameter update and collaborative training: Update the parameters through gradient descent, share the parameters between adjacent agents, synchronize the target network parameters regularly, and upload the optimized parameters to the corresponding communication satellite.

[0014] Preferably, in step S03, the anti-interference algorithm of the distributed multi-agent DRL further includes an online execution phase, and the online execution phase includes: Pre-trained parameter loading and real-time decision execution: The agent loads the optimized parameters, selects the best strategy based on local observations, and stores the experience data in the local buffer after executing the action; Incremental parameter update and local collaborative optimization: Sample a small batch of data from the buffer every fixed period to calculate the target value, update the network parameters with a constrained step size based on the temporal difference error, and share the updated parameters with adjacent agents to achieve distributed collaborative optimization.

[0015] Advantages of the above technical solutions: The asymptotically optimal distributed multi-agent DRL anti-interference method provided by the present invention models the anti-interference problem as a local interaction Markov game for the dynamic anti-interference scenario of the LEO satellite constellation downlink, and proves that this game belongs to an exact potential game (EPG) and at least one Nash equilibrium (NE) exists. Secondly, to obtain the equilibrium solution, an asymptotically optimal distributed multi-agent DRL anti-interference algorithm (DMADRLA) is proposed based on the "offline training - online execution" architecture. Finally, simulations verify that the proposed DMDRLA scheme can effectively balance the training cost and optimization performance of the anti-interference model and can obtain good anti-interference performance. Description of the Drawings

[0016] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0017] Figure 1 Flowchart of the asymptotically optimal distributed multi-agent DRL anti-interference method provided by an embodiment of the present invention; Figure 2 Anti-interference communication scenario diagram in a low-earth orbit constellation of the anti-interference method provided by an embodiment of the present invention; Figure 3 Downlink interference model diagram of the anti-interference method provided by an embodiment of the present invention; Figure 4 Convergence result comparison diagram of the anti-interference method provided by an embodiment of the present invention; Figure 5 Comparison diagram of the influence of traffic demand changes on the average satisfaction degree of the algorithm of the anti-interference method provided by an embodiment of the present invention. Detailed implementation manners

[0018] The following further describes in detail the implementation manners of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0019] The terms "first", "second", etc. (if any) in the specification and claims are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0020] It should be understood that the term " / and / " used herein is only a relational expression describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0021] Embodiment 1 An embodiment of the present invention provides a flowchart of an asymptotically optimal distributed multi-agent DRL anti-interference method as Figure 1As shown in the figure, it includes: Step S01: Construct an anti-jamming scenario for the satellite communication downlink, give the antenna models and signal models of the interfering satellite, communication satellite, and ground user, obtain the received signal-to-noise ratio of the ground user, which is used to calculate the ground user satisfaction, and construct an optimization problem and constraints for anti-jamming decision-making according to the ground user satisfaction; Step S02: Model the anti-jamming decision optimization problem as a local interactive Markov game model and give the reward design; Step S03: Use the anti-jamming algorithm of distributed multi-agent DRL to obtain the equilibrium solution of the local interactive Markov game model, so that the satellite can autonomously obtain the anti-jamming strategy.

[0022] 1 Constructing models and optimization problems Constructing an anti-jamming communication scenario model As Figure 2 shown, the anti-jamming communication scenario model constructed by the present invention in the low-earth orbit constellation includes two independently operating Walker-configured LEO constellations deployed at different orbital altitudes, namely, a malicious interference constellation and a communication constellation. Inside each constellation, adjacent satellites exchange information through inter-satellite links. The interfering satellites in the malicious interference constellation use narrow-beam high-gain parabolic antennas to implement random-frequency high-power suppression interference on the downlink of ground users within the ground area S, and the frequency switching follows the Markov probability matrix P. The interference strategy is planned by the ground control center and assigned to the interfering satellites through the ground gateway. The communication satellites in the communication constellation are equipped with multi-beam phased array antennas that can dynamically adjust the frequency and power, so as to avoid interference and ensure that the communication quality of ground users is not affected.

[0023] The set of ground users is denoted as The set of LEO communication constellations is denoted as Each communication satellite can form multiple independent beams, and each beam has an available channel set of The channel bandwidth is Bf, and the set of antenna transmission powers is . The interference constellation is denoted as Its antenna transmission power is fixed at . Among them, the interfering satellite closest to the center of the area S uses the same frequency set for interference. is the downlink established for the ground user at time slot t, is the interference link of the interfering satellite to the ground user at time slot t, is the co-channel interference of the link to the ground user at time slot t.

[0024] The time axis is divided into equal time slots . In each time slot, the transmitting antenna of the interfering satellite switches the interference channel according to the Markov probability matrix, and the communication satellite dynamically analyzes and selects the optimal frequency-power combination based on real-time interference perception and historical data, so as to mitigate the adverse effects caused by the interfering satellite. The anti-jamming method provided by the present invention focuses on the strategy optimization of LEO communication satellites under external malicious interference. By dynamically optimizing the frequency-power anti-jamming strategy, the impact of interference is reduced and the average satisfaction of ground users is improved.

[0025] Construct an antenna model For the interfering satellite, it is equipped with a high-gain narrow-beam antenna that complies with the ITU-R F.699-8 standard. The transmitting antenna gain of the interfering satellite is expressed as: ; ; Among them, is the off-axis angle between the center of the malicious interference beam and the link of the ground user to be interfered ; is the antenna half-power beam width, is the maximum gain of the interfering satellite antenna, is the diameter of the parabolic antenna, is the aperture efficiency, is the wavelength, and the power of the interference beam is fixed at .

[0026] For the communication satellite, according to ITU-R S.1528, the transmitting antenna gain of the communication satellite is expressed as: ; ; ; Among them, is the off-axis angle between the interference link and the downlink ; is the antenna half-power beam width, , is the maximum gain of the communication satellite antenna, is the intersection point of the main beam and the proximal sidelobe mask, is the far sidelobe level, .

[0027] For the ground user, according to ITU-R S.465, the receiving antenna gain is expressed as: ; ; ; Among them, is the off-axis angle between the interference link and the downlink ; is the maximum gain of the ground user antenna, is the circular equivalent diameter of the ground user antenna, is the aperture efficiency, is the wavelength.

[0028] Construct a signal model The downlink interference model is as Figure 3 shown. At time slot t, the ground user is subject to external interference from interfering satellites and internal co-channel interference from adjacent communication satellites.

[0029] For external malicious interference, when the downlink is under interference attack, the interference signal is expressed as: .

[0030] ; ; where is the interference signal, is the transmission power of the interference beam at time slot t, is the channel gain of link ; is a function to determine whether the interference link and the communication link select the same channel, is the distance between the interfering satellite and the ground user ; is the wavelength, is the transmitting antenna gain of the malicious interfering satellite in the direction of the ground user ; is the receiving antenna gain of the ground user in the direction of the malicious interfering satellite.

[0031] For internal co-channel interference, when multiple satellite links share the same channel, co-frequency interference will occur. The cumulative co-frequency interference received by the ground user is expressed as: .

[0032] ; where is the transmission power of the adjacent satellite beam at time slot t, is the channel gain of link ; is to determine the interference link and communication link function for whether to select the same channel, for the link distance, transmit antenna gain of the communication satellite causing interference in the direction of the ground user ; for the ground user receive antenna gain in the direction of the interfering communication satellite.

[0033] Therefore, the received signal-to-noise ratio (SNR) at the ground user can be expressed as: ; ; where, is the satellite transmit power of the link at time slot t ; for the link channel gain, is the maximum gain of the satellite transmit antenna, is the maximum gain of the user receive antenna, for the link distance, is the environmental noise power.

[0034] The achievable capacity is: ; where, is the channel bandwidth.

[0035] The arrival rate is measured by the satisfaction degree : ; where the parameter c controls the sensitivity of the ground user to the demand, is the predefined threshold.

[0036] The predefined threshold reflects the QoS requirements of the ground user. Being close to 1 means a service close to satisfaction, and being close to 0 means a service that does not meet the requirements.

[0037] Construct the optimization problem The optimization goal is to determine the optimal joint frequency and power allocation anti-interference scheme for the LEO satellite downlink, in order to minimize the impact of interference, i.e., maximize the average ground user satisfaction degree, under external malicious interference during the time period T.

[0038] The optimization problem is expressed as: ; Subject to the following constraints: : Time range, finite or infinite.

[0039] : Time discount factor.

[0040] : Ground user satisfaction function.

[0041] Decision variables: is the satellite transmit power for time slot t link and is the satellite communication channel for time slot t link to ensure that the allocation remains within the orthogonal frequency band and power constraints.

[0042] The optimization problem (P) balances the ground user satisfaction and the power cost , and introduces the time discount factor γ to obtain a consistent anti-interference effect. The above problem is a multi-variable sequential decision problem, which requires a sequential optimization framework to achieve the cumulative effect of variable decisions over consecutive time slots.

[0043] 2 Construction of anti-interference game model The present invention provides a theoretical framework for jointly optimizing spectrum-power anti-interference problems. The optimization problem (P) is modeled as a local interaction Markov game (LIMG) model, which belongs to the exact potential game (EPG) and has a Nash equilibrium (NE).

[0044] Local interaction Markov game model (LIMG) The anti-interference game model is defined as a seven-tuple , where: is the state space, , including its own satisfaction , the satisfaction of neighbors , the current interference state I j (t), the current total co-channel interference state I co (t).

[0045] N is the set of participants, where each participant n corresponds to a downlink l n .

[0046] A is the joint action space, , is the set of available channels, is the set of transmit powers. A single action is .

[0047] M is the set of neighbors, defined as 。

[0048] F is the state transition function, denoted as 。

[0049] R is the reward function, denoted as , where β is the weight of the neighbor reward.

[0050] is the time discount factor, 。

[0051] Design the reward mechanism The design of the reward mechanism takes into account the satisfaction of both individual and neighbor users. The reward of participant n at time slot t is denoted as: ; Adopt the cumulative long-term reward as the utility function of each participant, defined as: ; where, , this utility function synthesizes the weighted anti-interference effects of the current time slot and historical time slots, ensuring that all participants maximize the anti-interference utility of the current time slot while also maximizing the cumulative anti-interference effect within the entire system cycle.

[0052] Equilibrium analysis By constructing the potential function: ; And it is proved that the constructed local interaction Markov game belongs to an exact potential game, satisfying: ; Guarantee that the game has at least one pure strategy Nash equilibrium (NE), that is, the joint strategy satisfies: 。

[0053] From the mutual equivalence relationship between problem P and the potential function , it is proved that the optimal solution of problem P is the pure strategy NE of the game . To obtain the equilibrium solution of the game, an asymptotically optimal distributed multi-agent DRL anti-interference algorithm is proposed.

[0054] 3 Distributed multi-agent DRL anti-interference algorithm (DMDRLA) Based on the local interaction Markov game model (LIMG model), a DMDRLA algorithm is proposed by combining ground offline training and on-board incremental learning. It is implemented through a three-layer design: ground user neighborhood constraint, space-ground model synchronization, and autonomous online adaptation, and can be well applied to the dynamic LEO satellite network topology.

[0055] Offline training phase In the offline training phase, Algorithm 1 uses a distributed multi-agent DRL architecture to perform offline training on the neural network. The training process includes three steps: agent network initialization and experience collection, target optimization and error balancing, and parameter update and collaborative training. First, initialize the agent network, use the ε-greedy exploration strategy to collect agent states and decision-making information, and store them in the experience replay buffer. The decision-making process is constrained by the neighbor relationship among the internal agents of the satellite network. Second, perform batch sampling, target value calculation, and local utility optimization to balance the temporal difference error. Finally, update the network parameters through gradient descent, share the parameters among adjacent agents, and ensure the stability of training through periodic target function synchronization. After training is completed, upload the optimized model parameters to the corresponding LEO satellite to support autonomous operation.

[0056] Algorithm 1 Distributed training for anti-interference in the LEO constellation, including the following steps: Step 1: Initialize the parameters θ of all agent main networks, create an experience replay buffer D, and synchronize the target network parameters . .

[0057] Step 2: Initialize the environmental state s0 at each episode, perform T-step temporal interaction; each agent randomly explores with probability ϵ, otherwise selects an action ; execute the joint action , observe the reward and the new state .

[0058] Step 3: Each agent calculates the local reward n based on the neighbor information M ; store in the experience replay buffer D.

[0059] Step 4: Sample a batch of data B of size b from D n in each round, calculate the target Q value , minimize the loss function , and perform gradient descent on θ n to update the main network parameters.

[0060] Step 5: The agent shares the updated parameters with its neighbors m ∈ Mn θ n .

[0061] Step 6: Every fixed period, softly update the target network according to the mixing coefficient = 0.01: .

[0062] Step 7: Upload the finally optimized main network parameters to the satellite agent to support online anti-interference decision-making.

[0063] This algorithm uses the computing power of ground user terminals to independently complete model training, avoiding low-orbit satellites from bearing the computing load. By restricting the interaction between agents within a predefined neighbor range, the redundant coordination cost is significantly reduced. A parameter interaction mechanism between neighboring nodes is adopted to replace the centralized network, realizing orderly cooperation under a distributed architecture. It is proved in the EPG framework that the local optimization of distributed multi-agents can converge to the global optimal solution.

[0064] Online execution phase In the online execution phase, agents achieve fully decentralized real-time decision-making through pre-trained neural networks. As described in Algorithm 2, agents first load the pre-trained network parameters to achieve independent operation without centralized coordination. In each time slot, an agent observes its own and its neighbors' states, selects a joint spectrum-power policy that maximizes its neural network function, and executes these actions in the environment. In addition, this algorithm integrates an incremental information update mechanism, and agents periodically optimize their network parameters according to newly observed states and rewards without the need for complete retraining.

[0065] Algorithm 2 Incremental update for real-time decision-making anti-interference in LEO constellations, including:[[]] Step 1: Load the training parameters of all agent networks .

[0066] Step 2: Each agent observes its local state , selects an action , executes the action , obtains a reward and a new state , and stores the experience in the local buffer for subsequent incremental learning.

[0067] Step 3: Trigger incremental parameter updates every fixed time step, extract a small batch of experience data from the local buffer , calculate the target value of each experience , calculate the gradient of the loss function through the temporal difference error , update the parameter with a constrained step size , and then the agent only shares the updated parameter with neighbor m ∈ Mn.

[0068] In this algorithm, each agent makes independent decisions based only on local observation information, avoiding the communication and computational overheads brought by centralized decision-making; through the parameter incremental update mechanism, the self-adaptability of the pre-trained neural network in the dynamic satellite network environment is ensured. These two technologies jointly improve the real-time response ability of the satellite-ground collaborative system.

[0069] Proof of Convergence The proposed anti-jamming algorithm exhibits good convergence and asymptotic optimality, and it almost surely converges to a neighborhood-constrained NE. Under the established EPG framework, this NE approximates the global optimum of the potential function, which can be proved through the contraction property of the local Bellman operator, the convergence analysis of stochastic approximation in distributed updates, the neighborhood-constrained Nash equilibrium mechanism based on parameter sharing, and the analysis of the relationship with the global optimal solution under the EPG framework.

[0070] Simulation Results of the Asymptotically Optimal Distributed Multi-Agent DRL Anti-Jamming Method Provided by the Present Invention The present invention conducts a simulation analysis on the performance of the proposed asymptotically optimal distributed multi-agent DRL algorithm. The simulation involves 50 ground users randomly distributed within the potential interference area S of 500km×500km, and the satellite downlink is subject to jamming suppression from an external malicious interference constellation. The relevant simulation parameter settings are shown in Table 1.

[0071] Table 1 Simulation Parameters parameter interference constellation communication constellation orbital altitude 1000km 550km number of available channels 8 8 transmit power 45dBW 7 - 15dBW frequency range 10.7 - 12.7GHz 10.7 - 12.7GHz frequency - sweeping mode Markov mode adaptive channel noise - -110dBm Figure 4 The convergence performance analysis of three different anti-jamming methods, namely the centralized training method based on global information (CTAGI), the independent decision training method without information interaction (IIIDT), and the proposed method (DMDRLA), is given. CTAGI exhibits superior convergence performance due to its global information perception ability, but its practical application is restricted by satellite computing resources. IIIDT fails to formulate a coordinated anti-jamming strategy, resulting in a significant decline in convergence performance. The proposed method achieves approximately 91% of the convergence performance of CTAGI while only requiring approximately 8% of the link overhead of CTAGI, indicating that the proposed method maintains a coordinated advantage through local information interaction, verifying the effectiveness of DMDRLA.

[0072] Figure 5The influence of the traffic demand of ground users on the anti-interference performance of the algorithm is analyzed comparatively. Among them, CTAGI achieves the optimal anti-interference performance, DMDRLA is slightly lower than CTAGI, and IIIDT has the lowest anti-interference performance due to its completely independent training framework. When the LU traffic demand exceeds 80 Mbps, the anti-interference performance of IIIDT drops significantly, while the performance of DMDRLA remains relatively stable. Experiments confirm that DMDRLA can find a balance between the performance optimization of the neural network and the training overhead, making it applicable to the anti-interference scenario in resource-constrained LEO satellite networks.

[0073] In summary, the present invention proposes an asymptotically optimal distributed multi-agent DRL anti-interference method in a LEO satellite communication system. First, the anti-interference problem is modeled as a local interaction Markov game, which is proved to be an exact potential game and there is at least one pure strategy NE. Then, based on the "offline training - online execution" architecture, a DMDRLA algorithm is proposed. Finally, the simulation results verify that the proposed algorithm can effectively balance the anti-interference performance and training overhead of the network, enabling the satellite to autonomously optimize the frequency-power allocation strategy and having good anti-interference performance.

[0074] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is impossible to enumerate all the implementation manners here. Any obvious changes or modifications derived from the technical solutions of the present invention still fall within the protection scope of the present invention.

Claims

1. An asymptotically optimal distributed multi-agent DRL anti-interference method, characterized in that Including: Step S01: Construct an anti-jamming scenario for the satellite communication downlink, give the antenna models and signal models of the interfering satellite, communication satellite and ground user, obtain the received signal-to-noise ratio of the ground user for calculating the ground user satisfaction, and construct an optimization problem and constraints for anti-jamming decision-making according to the ground user satisfaction. Step S02: Model the anti-jamming decision optimization problem as a local interactive Markov game model and give the reward design. Step S03: Use the anti-jamming algorithm of distributed multi-agent DRL to obtain the equilibrium solution of the local interactive Markov game model, so that the satellite autonomously obtains the anti-jamming strategy.

2. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 1, characterized in that In step S01, the antenna model is used to obtain the transmitting antenna gain of the interfering satellite, the transmitting antenna gain of the communication satellite and the receiving antenna gain of the ground user.

3. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 1, characterized in that In step S01, the signal model is used to calculate the received signal-to-noise ratio of the ground user according to the external interference from the interfering satellite and the internal co-channel interference from the adjacent communication satellite received by the ground user.

4. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 1, wherein In step S01, the received signal-to-noise ratio of the ground user is: ; ; Among them, is the time slot t link satellite transmission power, is the time slot t link channel gain, is the co-channel interference of adjacent satellites, is the interference signal, is the ambient noise power, is the time slot t link satellite transmission power, is the time slot t link channel gain, is a function for judging whether to select the same channel in the time slot t link and or not.

5. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 1, characterized in that In step S01, the calculation formula for the ground user satisfaction is: ; Among them, is the received signal-to-noise ratio of the ground user, c is to control the sensitivity of the user to the demand, is a predefined threshold.

6. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 1, wherein In step S01: The optimization problem is expressed as: ; The constraints include: 、 、 、 ; wherein, is the satisfaction of ground users , is the time discount factor is the time slot t link is the satellite transmission power is the time slot t link is the satellite communication channel is the set of available channels is the set of transmission powers 7. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 1, wherein In step S02, the reward design takes into account the satisfaction of both individuals and neighbors, and synthesizes the weighted anti-jamming effects of the current time slot and historical time slots, so that all participants maximize the anti-jamming utility of the current time slot while also maximizing the cumulative anti-jamming effect within the entire system cycle.

8. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 7, wherein The cumulative anti-jamming effect uses the cumulative long-term reward as the utility function for each participant, and the utility function is expressed as: ; Among them, is the strategy of all participants, is the time discount factor, is the satisfaction of terrestrial user . is the satisfaction of terrestrial user . is the time slot t link satellite transmission power, is the time slot t link satellite transmission power, β is the weight of neighbor reward, N is the set of participants, is the participant n set of neighbor users.

9. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 1, wherein In step S03, the anti-jamming algorithm of distributed multi-agent DRL includes an offline training stage, and the offline training stage includes: Agent network initialization and experience collection: Initialize the agent network parameters, use the ε-greedy exploration strategy for environment interaction, collect agent state and decision information under the neighbor relationship constraint, and store them in the experience replay buffer. Objective optimization and error balance: Batch sample data from the experience replay buffer, calculate the target Q value, combine the temporal difference error, and optimize the strategy through the utility function. Parameter update and collaborative training: Update the parameters through gradient descent, share the parameters between adjacent agents, synchronize the target network parameters regularly, and upload the optimized parameters to the corresponding communication satellite.

10. The asymptotically optimal distributed multi-agent DRL anti-interference method according to claim 9, characterized in that In step S03, the anti-jamming algorithm of distributed multi-agent DRL also includes an online execution stage, and the online execution stage includes: Pre-trained parameter loading and real-time decision execution: The agent loads the optimized parameters, selects the best strategy based on local observations, and stores the experience data in the local buffer after executing the action. Incremental parameter update and local collaborative optimization: Sample a small batch of data from the buffer every fixed period to calculate the target value, update the network parameters based on the temporal difference error with a constrained step size, and share the updated parameters with adjacent agents to achieve distributed collaborative optimization.

Citation Information

Patent Citations

  • LEO satellite access switching algorithm based on multi-agent deep cycle Q network

    CN117040588A

  • Ocean communication on-demand service coverage method and device based on low orbit satellite hopping beam

    CN119232236A

  • Deep reinforcement learning-based random access method for low earth orbit satellite network and terminal for the operation

    US20230189353A1

  • Digital twin-based deduction and optimization method and system for intelligent reflecting surface communication system

    US20250175216A1