Robust anti-interference spectrum access method facing unknown intelligent interference

By establishing a partially observable adversarial team stochastic game model on the intelligent decision-making devices of base stations and jammers, and combining virtual training and online adversarial learning, the problem of intelligent communication adversarial under unknown intelligent interference is solved, and the anti-interference capability and robustness of wireless communication systems are improved.

CN121750125APending Publication Date: 2026-03-27ARMY ENG UNIV OF PLA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-08
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies are ineffective in dealing with intelligent communication countermeasures when facing unknown intelligent interference, and their scalability is weak, making them unable to adapt to complex scenarios with multiple communication devices.

Method used

Intelligent decision-making equipment using base stations and jammers continuously senses the spectrum. It models the system using a partially observable adversarial team stochastic game model (POATSG), combines virtual training and online adversarial learning frameworks, utilizes virtual jamming for strategy optimization, and designs synchronous update and parameter reset mechanisms to enhance the anti-jamming capabilities of the communicating parties.

Benefits of technology

It improves the anti-interference performance of wireless communication systems, enhances robustness against unknown intelligent interference, reduces end-to-end latency and hop count, and increases network transmission rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121750125A_ABST
    Figure CN121750125A_ABST
Patent Text Reader

Abstract

The invention discloses an unknown intelligent interference-oriented robust anti-interference spectrum access method, which comprises the following steps of: 1, continuously sensing the whole communication frequency band by intelligent decision-making equipment at a base station end, and judging whether the communication is successful or not by a base station through receiving and demodulating data; 2, continuously sensing the whole frequency band at the jammer, deciding a channel to be interfered at the beginning of the time slot, releasing an interference signal, and evaluating an interference effect at the end of the time slot; 3, modeling an intelligent frequency spectrum confrontation problem into a partially observable antagonistic team random game model; 4, adopting a learning framework of virtual training and online confrontation, setting that virtual interference adopts a continuous learning mode, and continuously searching weak points of a communication party strategy; 5, virtual training is started; and step 6, designing a parameter resetting mechanism of virtual interference. According to the invention, the network transmission rate and the anti-interference performance are improved, and the end-to-end time delay and hop count are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication technology, and in particular relates to a robust anti-interference spectrum access method for dealing with unknown intelligent interference. Background Technology

[0002] Due to the openness of the electromagnetic spectrum, wireless communication is highly vulnerable to external interference attacks. This inherent vulnerability has spurred the development of anti-jamming communication technologies. In recent years, with the rapid development of artificial intelligence (AI) technology, intelligent decision-making techniques such as reinforcement learning (RL) and deep reinforcement learning (DRL) have been increasingly applied to the research field of anti-jamming communication, empowering anti-jamming decision-making and achieving remarkable success. AI technology has not only improved the capabilities of anti-jamming communication but has also promoted the rapid development of jamming techniques. The increasingly powerful jamming capabilities not only pose a significant threat to wireless communication systems but also create a new challenge for anti-jamming communication research: how to ensure that the communicating party gains the upper hand in the intelligent confrontation against jamming?

[0003] Some research has begun to study the problem of intelligent spectrum countermeasures. However, most studies assume that the communicating party has prior knowledge of the jammer's information, for example: (Li W, Xu Y, Chen J, et al. Know thy enemy: Anopponent modeling-based anti-intelligent jamming strategy beyond equilibrium solutions[J]. IEEE Wireless Communications Letters, 2022, 12(2): 217-221.) assuming that the jammer's action space is known, (Yuan H, Chen J, Li W, et al. Opponent awareness-based anti-intelligent jamming channel access scheme: A deep reinforcement learning perspective[J]. IEEE Internet of Things Journal, 2024, 11(7), 11202-11216.) assuming that the jammer's neural network architecture is known, etc. Obviously, these strong assumptions are difficult to hold in actual anti-jamming communication scenarios. The existence of these assumptions also limits the applicability of existing work.

[0004] However, the aforementioned studies require prior knowledge of the interference's action space, which is still far from practical application. Furthermore, these studies only consider one-to-one intelligent communication countermeasures, limiting the scenarios and restricting the scalability of the methods. With the increasing number of intelligent frequency-using devices, researching intelligent communication countermeasures that incorporate more communication devices is imperative. Summary of the Invention

[0005] The purpose of this invention is to solve the problems mentioned in the background art and to propose a robust anti-interference spectrum access method for unknown intelligent interference.

[0006] To achieve the objective of this invention, a robust anti-interference spectrum access method for dealing with unknown intelligent interference is provided, the method comprising:

[0007] Step 1: The intelligent decision-making device at the base station continuously senses the entire communication frequency band and decides the corresponding communication channel for each user at the beginning of the communication time slot. The decision result is fed back to each user through the downlink. The base station determines whether the communication is successful by receiving and demodulating the data.

[0008] make Let these represent the decision-making time, feedback time, reception time, and benefit calculation time, respectively. This represents the time spent receiving feedback and the time spent transmitting data, assuming that the time spent on decision-making, feedback, and calculating the benefit is much shorter than the time spent receiving data. And the time spent receiving feedback is much shorter than the time spent transmitting it. Therefore, it can be ignored;

[0009] At time t, let and Representing users respectively The channel gain from the base station and from the jammer to the base station, let , and Let PSD represent the power spectral density functions of the user signal, interference signal, and noise, respectively; and let PSD represent the power spectral density function of the wireless signal received at the base station, considering the presence of all signals. Defined as:

[0010] (1);

[0011] in, and Let represent the center frequencies of the i-th user signal and the k-th interference signal, respectively; and These represent the antenna transmit gains of the user and the jammer, respectively. This indicates the antenna receiving gain of the base station;

[0012] exist At any given time, the spectrum vector sensed by the base station is defined as:

[0013] (2);

[0014] in, This represents the number of samples for spectrum sensing, where B is the bandwidth of the entire frequency band. Frequency resolution for spectrum sensing; The x-th spectrum sample value is defined as:

[0015] (3);

[0016] in, , Indicates the starting frequency for spectrum sensing;

[0017] user The signal-to-jamming plus noise ratio (SJNR) of the transmitted signal propagating to the base station is defined as:

[0018] (4);

[0019] in, Indicates the bandwidth of the user signal; let This represents the signal-to-interference-plus-noise ratio (SIR) threshold required for a base station to correctly demodulate the signal. Indicates user Channel gain from the base station and from the jammer to the base station, user The normalized communications throughput (NCT), i.e. the number of successful communications, is defined as:

[0020] (5);

[0021] in, For indicator functions, when When true ,otherwise .

[0022] Step 2: The intelligent decision-making device at the jammer continuously senses the entire frequency band and decides which channel to jam at the beginning of the time slot; the jammer then releases the jamming signal on the corresponding channel and evaluates the jamming effect at the end of the time slot.

[0023] make These represent the decision-making time, disturbance time, and benefit calculation time, respectively. It is assumed that the decision-making time and benefit calculation time are much smaller than the disturbance time. Therefore, the time spent on decision-making and the time spent calculating benefits can be ignored;

[0024] Unlike the equipment deployment method of the communication party, the jammer's intelligent decision-making equipment is deployed at the jammer. While the intelligent decision-making equipment continuously senses the spectrum, the jammer needs to release jamming signals, thus the jammer faces the problem of self-interference at both the transmitter and receiver. To better focus on uncovering the frequency usage patterns of the communication signal and thus improve the jamming effect, the jammer needs to remove the jamming signal from the sensing results of the intelligent decision-making equipment. Since the jamming signal released by the jammer is known to the intelligent decision-making equipment, self-interference cancellation (SIC) technology is used to eliminate the influence of the jamming signal. In summary, the sensing results of the intelligent decision-making equipment only contain user signals and noise signals, without interference signals; At time t, the signal power spectral density function at the intelligent decision-making device is:

[0025] (6);

[0026] in, This indicates the receiving antenna gain of the interfering party's intelligent decision-making device. Indicates the user at time t Channel gain at the jammer;

[0027] The spectrum vector sensed by the intelligent decision-making device is defined as:

[0028] (7);

[0029] in, This indicates the number of sampling points for spectrum sensing. The frequency resolution for spectrum sensing. The x-th spectrum sample value is defined as:

[0030] (8)

[0031] The interfering party makes intelligent interference decisions based on the perception results.

[0032] Step 3: Based on the characteristics of the above models, the intelligent spectrum adversarial problem is modeled as a partially observable adversarial team stochastic game model (POATSG).

[0033] Partially observable adversarial team stochastic game models are represented as follows: The definition of each element is as follows:

[0034] : Set of environmental states; In the intelligent spectrum countermeasure problem, the environment refers to the electromagnetic spectrum environment, and the environmental state refers to the spectrum state;

[0035] The observation set of the communicating party; observation refers to the result obtained by the communicating party through spectrum sensing of the spectrum state; in the modeling process; the spectrum sensing result of the communicating party in each time slot is directly defined as the observation of the spectrum state, but this definition method contains less information. This invention defines the spectrum waterfall (SW) as the observation of the communicating party; the spectrum waterfall is a sequence composed of spectrum vectors sensed at the current and historical times, and its mathematical expression is:

[0036] (9);

[0037] in, This refers to the spectrum vector perceived by the defined communication party. This refers to the spectrum vectors perceived by the communicating party at the first n time points. The length of the sequence is indicated by the above definition. As can be seen from the above definition, the spectrum waterfall plot contains information in three dimensions: time, frequency, and power. It can not only reflect the changing pattern of the spectrum state, but also provide sufficient information for the communication party to infer the changing pattern of the interference signal.

[0038] The interference party's observation set; unlike the communication party's observation definition, the interference party's observations include two heterogeneous data sets; the interference party's spectrum sensing results exclude interference signals; in fact, directly removing interference signals leads to the interference party losing some information, such as the correlation between the actions of both parties. From the perspective of information integrity, additional design is needed to supplement the action information lost by the interference party. Specifically, the interference party's observations are defined as:

[0039] (10);

[0040] in, This refers to the sequence of historical spectrum vectors perceived by the interfering party, defined as:

[0041] (11);

[0042] The length of the corresponding sequence; These are additional action data, also presented in the form of historical sequences, and defined as follows:

[0043] (12);

[0044] in, This indicates that the action sequence is One-Hot encoded; This refers to the action of the interfering party at time t, the specific definition of which is given below.

[0045] These are the observation functions of the communicating party and the interfering party, respectively. In the problem under study, the observation functions of both parties are implicit. Due to the different geographical locations of the two parties, their observations of the spectrum state are naturally different under the influence of the wireless signal propagation law.

[0046] The set of actions of the communicating parties; the actions of the communicating parties are defined as the combination of channels used by the two users. Each user's selectable channel comes from the set of available channels: C represents the number of available channels; furthermore, the channels for two users within each time slot should be different to avoid internal interference.

[0047] The set of actions taken by the interfering party; the actions of the interfering party are defined as the channel combination in which the interfering signal is located. Similarly, the channel occupied by the interference signal is different in each time slot;

[0048] The revenue function of the communicating party; the revenue of the communicating party is defined as the sum of the revenues of the two communicating users:

[0049] (13);

[0050] For each user, the benefit criterion is to reflect the communication quality and cost as much as possible; specifically, the benefit for each user is defined as:

[0051] (14);

[0052] in, That is, the defined user The normalized throughput obtained This is an indicator function. and These represent the weighting factors corresponding to successful and failed communication, respectively.

[0053] : The profit function of the interfering party; Since the opposing parties have conflicting interests and completely opposite goals, consider the scenario of a zero-sum game, that is, the profit of the interfering party is always the opposite of the profit of the communicating party, and the sum of the profits of both parties is 0.

[0054] P: State transition function; In the problem under consideration, the state transition function is unknown to both parties, so it is not specifically defined here.

[0055] Discount factor; Used to depict combat scenarios with infinite duration;

[0056] make Indicates user The strategy, From the observation space To the action space The mapping; given observations, through Calculate a value in the action space The probability distribution on; let This indicates the joint strategy of the communicating parties, making The utility function, representing the expected long-term cumulative discount benefit of the communicating party, is expressed as:

[0057] (15);

[0058] in, This represents the interfering party's strategy; the communication party's optimization objective is to optimize its strategy. The utility function, expressed as:

[0059] (16);

[0060] Since the interfering party's intention is to disrupt the communicating party's communication as much as possible, which is completely in conflict with the communicating party's interests, its optimization objective is exactly the opposite of the communicating party's; that is, the interfering party attempts to optimize its strategy. To minimize the utility function of the communicating parties; considering the zero-sum game characteristics of both sides, the intelligent spectrum adversarial problem is expressed as:

[0061] (17);

[0062] Solution to the intelligent spectrum adversarial problem This corresponds to the equilibrium strategy in the game theory model (POATSG). The intelligent spectrum adversarial problem involves two decision-making entities, namely, the two adversaries. It is important to note that this invention does not simultaneously optimize the strategies of both parties. This invention studies the problem from the perspective of the communicating party, aiming to solve an anti-interference strategy for the communicating party. The interfering party is introduced as an attack method to quantitatively evaluate the anti-interference communication performance of the communicating party under its interference attack.

[0063] Step 4: To address the issue of unknown intelligent interference information, a learning framework of virtual training and online adversarial is adopted. The virtual interference is set to continuously learn and continuously search for weaknesses in the communication party's strategy, thereby forcing the communication party to continuously improve its anti-interference strategy.

[0064] The virtual interference is configured as follows:

[0065] Observation: The observation of virtual interference is exactly the same as that of the communication party; ensure that the virtual interference party has enough information, at least no less than the information held by the communication party;

[0066] Action: The virtual jamming selects two channels in each time slot to release virtual jamming signals. The action is represented as follows. The channels available for virtual interference are all derived from the set of available channels. C represents the number of available channels; furthermore, the two channels used for interference are different in each time slot; it should be noted that the virtual interference signals are created in the intelligent decision-making device of the communicating party, so they are directly added to the sensing results of the communicating party and do not need to pass through the wireless channel.

[0067] Benefits: Virtual interference serves as a hypothetical enemy introduced by the communicating party.

[0068] The objective is completely opposite to that of the communicating party, that is, the benefit of virtual interference is the opposite of the benefit of the communicating party, and the sum of the benefits of both parties is 0;

[0069] Intelligent decision-making algorithm: Since the parameters of the real jammer cannot be obtained, the communicating party configures the virtual jamming based on its own settings. Therefore, the intelligent decision-making algorithm and neural network architecture of the virtual jamming are exactly the same as those of the communicating party.

[0070] Step 5: Divide the entire virtual training process into several update cycles, start virtual training, and have the communicating party and the virtual interference observe and perform their respective actions to obtain their respective benefits. Then, at the end of the update cycle, perform a policy update.

[0071] To simulate the communication process between the user and the base station, it is assumed that the probability distribution of the channel gain between the user and the base station is known. Since the virtual interference signal is directly added to the sensing results of the communicating party, it is not necessary to consider the channel gain between the virtual interference and the base station.

[0072] Once the above preparations are completed, the virtual training process can begin. During this period, the communicating party and the virtual interference will respectively conduct observations... Make their own actions The virtual environment generates new spectrum states based on the actions of both parties, allowing the communicating parties to obtain new observations. The communication provider calculates its revenue based on the signal demodulation at the base station. Benefits of virtual interference This is the inverse of the communication party's benefit; at the end of each time slot, empirical data is obtained: And store the experience data in the experience replay pool;

[0073] Before updating the strategy, the empirical data needs further processing. Since the virtual training process is designed from a global perspective, the original empirical data contains global information, i.e., data from both sides. When the communicating party and the virtual interference update their strategies, the original data needs to be split into data formats that each party can access. Specifically, the data format used by the communicating party to update its strategy is: The data format used for the virtual interference update strategy is: After the split is completed, both parties will update their strategies using their respective experience data.

[0074] However, if both sides update their strategies independently, the environment becomes unstable, making it difficult for either side to learn a good strategy. To address this environmental instability, this invention designs a synchronous update mechanism. Specifically, the entire virtual training process is divided into several update cycles. During each update cycle, the policies of both parties remain unchanged, with updates performed only at the end of the cycle. This design ensures that the environment remains steady-state within each update cycle. The entire virtual training process repeats this synchronous update mechanism until the policies of both parties converge.

[0075] Step 6: To prevent overfitting between the two strategies that may occur during virtual training, a parameter reset mechanism for virtual interference is designed. By restarting the learning process of virtual interference, the dead loop of overfitting between the two strategies can be broken.

[0076] To add a parameter reset period to the virtual interference. At the end of each parameter reset cycle, with probability Update neural network parameters with probability Reset the neural network parameters;

[0077] Reset refers to reinitializing the neural network parameters and restarting the learning process; the parameter reset period is... With update cycle The relationships between them are integer multiples, that is:

[0078] (18);

[0079] By resetting the parameters and restarting the learning process of the virtual interference, two things can be achieved. First, it prevents the virtual interference from overfitting to the communication party's strategy. Relearning from scratch allows the virtual interference to rediscover weaknesses in the communication party's strategy, enabling targeted interference attacks and improving the quality of its attacks. Second, it prevents the communication party from overfitting to the virtual interference's strategy. As adversaries, a more diverse and powerful virtual interference can force the communication party to further refine its strategy and improve its robustness. By periodically resetting the parameters of the virtual interference with probability, the vicious cycle of mutual entanglement and overfitting between the two parties' strategies is broken, improving the quality of adversarial training and allowing both parties to converge to a better equilibrium solution, thus enhancing the robustness of the communication party's strategy.

[0080] The execution process is as follows:

[0081] Step 61, Initialization Process:

[0082] Initialize the parameters of the communication decision network 1. Evaluate the parameters of the network Experience Data Pool Parameters of the decision network for virtual interference 1. Evaluate the parameters of the network Experience Data Pool ;

[0083] Set the update interval for the communicating party. Reset cycle of virtual interference Reset probability Learning rates of decision-making networks and evaluation networks Learning parameters Weighting factor ;

[0084] Step 62, enter the loop. ;

[0085] The communicating party obtains current spectrum observations ;

[0086] According to the policy Select Action Virtual interference selects actions based on the strategy. Both parties perform their respective actions;

[0087] The communicating party obtains the spectrum observation at the next moment. And calculate the revenue. Gaining benefits through virtual interference ;

[0088] The communication party and the virtual interference will respectively use empirical data and Stored in their respective experience data pools , ;

[0089] The process begins; if it is the end of an update cycle, the neural network parameters need to be updated.

[0090] Virtual interference with probability Reset network parameters with probability Update network parameters;

[0091] The communicating parties synchronously update the neural network parameters;

[0092] Clear the experience data pool of communication parties and virtual interference. , .

[0093] Compared with existing technologies, the significant advancement of this invention lies in its ability to establish routes in interference-prone environments without central control or relying on network-wide topology information. Instead, it relies solely on link status and geographical location information. Nodes dynamically perceive the spectrum environment to identify spectrum vulnerabilities and combine this with the distance from their neighbors to the destination to find suitable next-hop nodes and communication channels. Finally, nodes combine these two aspects to perform joint actions and utilize Q-learning to optimize the selection of next-hop nodes and channels. Ultimately, this method has been verified to improve network transmission rates and anti-interference performance, while reducing end-to-end latency and hop count.

[0094] To more clearly illustrate the functional characteristics and structural parameters of the present invention, further explanation is provided below in conjunction with the accompanying drawings and specific embodiments. Attached Figure Description

[0095] Figure 1 This application provides a typical wireless communication confrontation scenario diagram consisting of red and blue teams;

[0096] Figure 2 This is a simulation scene diagram provided in this application;

[0097] Figure 3 This is the spectral waterfall diagram provided in this application;

[0098] Figure 4 This is an anti-interference effect diagram of the jammer provided in this application when using the PPO algorithm;

[0099] Figure 5 This is a diagram showing the anti-interference effect of this application when facing different interference algorithms;

[0100] Figure 6 This application provides an anti-interference effect diagram when facing different numbers of interference signals;

[0101] Figure 7This is an ablation experiment diagram of the parameter reset mechanism provided in this application. Detailed Implementation

[0102] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0103] For ease of description and illustration, the robust anti-interference spectrum access method proposed in this invention will be abbreviated as "adv," short for adversarial fictitious training. Three comparison methods were considered in the simulation experiments:

[0104] ind-1 / 4: In the proposed method, the communicating party and the virtual interference are updated synchronously during virtual training. The comparative method is an independently updated version of the proposed method. In the comparative method, the update cycle lengths of the communicating party and the virtual interference are equal, but the parameter update times are staggered by 1 / 4 update cycle.

[0105] ind-1 / 2: This comparison method is also an independently updated version. In this comparison method, the update cycle of the communicating party and the virtual interference is the same, but the update times are staggered by 1 / 2 cycle.

[0106] RDM is an abbreviation for random. This comparative method refers to the random selection of perturbation actions in the virtual adversarial training process.

[0107] Figure 1 This invention presents a typical wireless communication adversarial scenario involving two opposing sides, red and blue. The red side acts as the jammer, comprising an intelligent jammer that runs an intelligent jamming algorithm to release jamming signals. Information such as the intelligent jamming algorithm and the number of jamming signals is unknown to the blue side. The blue side acts as the communicator, comprising a base station and two users. The two users are deployed in designated areas, continuously collecting data and transmitting it back to the base station via wireless communication. To achieve anti-jamming communication under conditions of unknown jamming information, a robust anti-jamming spectrum access method for unknown intelligent jamming is proposed. This method does not require prior knowledge of the jammer's information; it constructs virtual jamming based solely on the communicator's own configuration. Through adversarial training against the virtual jamming and periodically resetting the virtual jamming's parameters, robust anti-jamming communication can be achieved against intelligent attacks from different types of jamming.

[0108] Figure 2This is the simulation scenario of the present invention. The location of the communication base station is fixed at (0m, 0m), and the locations of the two communication users are set at (0m, 500m) and (0m, -500m) respectively. The location of the jammer is set at (1000m, 0m). The units of the above distances are meters. Path loss factor in the channel model. The antenna gain of all devices in the scenario is set to: dB. Consider a 20MHz frequency band, evenly divided into 5 orthogonal channels, each with a bandwidth of dB. MHz. Both the communicating party's and the interfering party's intelligent decision-making devices are based on... The spectral resolution is kHz, and the spectrum is sensed once per millisecond. Each sensed spectral vector contains... 200 sample values. Communication time slot length set to 5ms. Transmit power of communication and interference signals set to 0.1W and 10W respectively. Noise power spectral density set to -174dB / Hz. Both user and interference baseband signals are generated using raised cosine filters with a roll-off factor of 0.5. The signal-to-interference-plus-noise ratio (SNR) threshold for correctly demodulated signals at the base station is set to... dB.

[0109] Figure 3 The graph shows a 200ms spectrum waterfall plot. In each time slot, the communicating party needs to select two channels from five for communication; therefore, the action space of the communicating party contains a total of 10 actions. The weighting factors for successful and failed communication are as follows: and The relevant parameters of the PPO algorithm are set as follows: and The parameters of the neural networks were initialized using Orthogonal and optimized using the Adam optimizer. The learning rates for the decision network and the evaluation network were set to [values ​​to be filled in]. and Furthermore, advantage function normalization and payoff normalization were employed during training. All neural networks were implemented using the PyTorch 1.12.0 framework on an Nvidia RTX 4080 graphics card. All simulation experiments were independent, repeated experiments with five random number seeds.

[0110] Figure 4The paper presents the results of testing the anti-jamming performance of the proposed method and three comparative methods, considering the PPO algorithm as the first consideration. Detailed ANCT results are provided. The results of the proposed method are represented by a red curve, while the results of the three comparative methods are represented by blue, black, and green curves, respectively. The curves in the figure show the results of repeated experiments with five random number seeds, and the lighter shaded areas represent the fluctuation range of the five results. The overall trend of the results is that as the jammer's learning process deepens, the performance of all methods gradually decreases and then stabilizes. This indicates that the anti-jamming strategies learned by all methods still have some exploitability, meaning that some weaknesses remain in the strategies, which are then captured by the online learning jammer for targeted jamming attacks. It is easy to see that the performance decline of the proposed method is smaller, while the performance decline of the three comparative methods is more significant. This shows that the proposed method has lower exploitability and stronger robustness compared to the comparative methods, and can better resist jamming attacks, thus achieving higher anti-jamming communication throughput.

[0111] Figure 5 The comparison results of different methods are presented. The results of the proposed method are represented by red bar graphs. The results of the comparative methods ind-1 / 4 and ind-1 / 2 are represented by blue graphs filled with diagonal lines and orange graphs filled with grids, respectively. The results of the comparative method rdm are represented by green graphs filled with dots. The results show that the performance of the proposed method is more stable when facing different interference algorithms, while the performance of the comparative methods fluctuates more significantly. This indicates that the proposed method has stronger robustness and adaptability to different interference algorithms. Comparing the performance of the comparative methods against different interference algorithms reveals that these methods perform best against the SACD interference algorithm and worst against the PPO interference algorithm. This indicates that the jammer has a stronger interference capability when using the PPO algorithm, posing a greater threat to anti-jamming algorithms. Therefore, the PPO algorithm will be considered for the jammer in subsequent simulation experiments.

[0112] Figure 6The results of FANCT for four methods under different scenarios are presented. We calculated the average of the last 100 results after the ANCT curve convergence, denoted as FANCT (Final Average Normalized Communication Throughput). The results of the proposed method are represented by red bars, while the results of the three comparative methods are represented by blue, orange, and green bars, respectively. The overall trend conveyed by the simulation results is that the less interference signal in each time slot, the better the performance of each method; as the interference signal increases, the performance of each method decreases significantly. This phenomenon is consistent with expectations. Less interference signal in each time slot means more available channels, and the communication success rate of each method is naturally higher. As the interference signal increases, the interference capability continuously strengthens and the available channels gradually decrease, the communication success rate of each method gradually decreases, and the communication throughput naturally decreases as well.

[0113] Figure 7 The performance comparison results of the proposed method and the comparative method with each other are presented in three scenarios. The results of the proposed method are represented by red bar graphs, while the results of the comparative method are represented by blue graphs filled with diagonal lines. The results show that, except for the scenario where M=1, the performance of the parameter reset mechanism proposed in this invention is significantly better than the comparative method. Even in the scenario where M=1, the performance of the proposed method is very close to that of the comparative method. This indicates that the parameter reset mechanism, by enriching the virtual interference strategies, provides the communicating party with more powerful and diverse adversaries, improves the learning quality of the communicating party, and helps the communicating party learn more efficient and robust anti-interference strategies, thus achieving better anti-interference communication performance when facing different amounts of interference signals.

[0114] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A robust anti-interference spectrum access method for unknown intelligent interference, characterized in that, The method includes: Step 1: The intelligent decision-making device at the base station continuously senses the entire communication frequency band and decides the communication channel for each user at the beginning of the communication time slot. The decision result is fed back to each user through the downlink. The base station determines whether the communication is successful by receiving and demodulating the data. Step 2: The jammer continuously senses the entire frequency band and decides which channel to jam at the beginning of the time slot; the jammer releases the jamming signal on the corresponding channel and evaluates the jamming effect at the end of the time slot. Step 3: Model the intelligent spectrum adversarial problem as a partially observable adversarial team stochastic game model POATSG; Step 4: Using a virtual training and online adversarial learning framework, virtual interference is set to continuously learn and continuously find weaknesses in the communication side's strategy, thereby forcing the communication side to continuously improve its anti-interference strategy. Step 5: Divide the entire virtual training process into several update cycles, start virtual training, and have the communication party and virtual interference observe and perform their respective actions to obtain their respective benefits. Then, update the strategy at the end of the update cycle. Step 6: Design a parameter reset mechanism for virtual interference to break out of the dead loop of overfitting between the two strategies by restarting the learning process of virtual interference.

2. The method according to claim 1, characterized in that, Step 1 includes: make Let these represent the decision-making time, feedback time, reception time, and benefit calculation time, respectively. This represents the time spent receiving feedback and the time spent transmitting data, assuming that the time spent on decision-making, feedback, and calculating the benefit is much shorter than the time spent receiving data. And the time spent receiving feedback is much shorter than the time spent transmitting it. Therefore, it can be ignored; At time t, let and Representing users respectively The channel gain from the base station and from the jammer to the base station, let , and Let PSD represent the power spectral density functions of the user signal, interference signal, and noise, respectively; and let PSD represent the power spectral density function of the wireless signal received at the base station, considering the presence of all signals. Defined as: (1); in, and Let represent the center frequencies of the i-th user signal and the k-th interference signal, respectively; and These represent the antenna transmit gains of the user and the jammer, respectively. This indicates the antenna receiving gain of the base station; exist At any given time, the spectrum vector sensed by the base station is defined as: (2); in, This represents the number of samples for spectrum sensing, where B is the bandwidth of the entire frequency band. Frequency resolution for spectrum sensing; The x-th spectrum sample value is defined as: (3); in, , Indicates the starting frequency for spectrum sensing; user The signal-to-jamming plus noise ratio (SJNR) of the transmitted signal propagating to the base station is defined as: (4); in, Indicates the bandwidth of the user signal; let This represents the signal-to-interference-plus-noise ratio (SIR) threshold required for a base station to correctly demodulate the signal. Indicates user Channel gain from the base station and from the jammer to the base station, user The normalized communications throughput (NCT), i.e. the number of successful communications, is defined as: (5); in, For indicator functions, when When true ,otherwise .

3. The method according to claim 2, characterized in that, Step 2 includes: make These represent the decision-making time, disturbance time, and benefit calculation time, respectively. It is assumed that the decision-making time and benefit calculation time are much smaller than the disturbance time. Therefore, the time spent on decision-making and the time spent calculating benefits can be ignored; The jammer's intelligent decision-making equipment is deployed at the jammer; while the intelligent decision-making equipment continuously senses the spectrum, the jammer needs to release jamming signals, and the jammer needs to remove these jamming signals from the intelligent decision-making equipment's sensing results; by using self-interference cancellation technology to eliminate the influence of jamming signals, the intelligent decision-making equipment's sensing results only contain user signals and noise signals, without any jamming signals; At time t, the signal power spectral density function at the intelligent decision-making device is: (6); in, This indicates the receiving antenna gain of the interfering party's intelligent decision-making device. Indicates the user at time t Channel gain at the jammer; The spectrum vector sensed by the intelligent decision-making device is defined as: (7); in, This indicates the number of sampling points for spectrum sensing. The frequency resolution of spectrum sensing is defined as follows: The x-th spectrum sample value is defined as: (8) The interfering party makes intelligent interference decisions based on the perception results.

4. The method according to claim 3, characterized in that, Step 3 includes: Partially observable adversarial team stochastic game models are represented as follows: The definition of each element is as follows: : Set of environmental states; In the intelligent spectrum countermeasure problem, the environment refers to the electromagnetic spectrum environment, and the environmental state refers to the spectrum state; The observation set of the communicating party; observation refers to the result obtained by the communicating party through spectrum sensing of the spectrum state; the spectrum waterfall (SW) is defined as the observation of the communicating party; the spectrum waterfall is a sequence of spectrum vectors sensed at the current and historical moments, and its expression is: (9); in, This refers to the spectrum vector perceived by the defined communication party. This refers to the spectrum vectors perceived by the communicating party at the first n time points. It indicates the length of the sequence; the spectrum waterfall plot contains information in three dimensions: time, frequency, and power. The interference party's observation set; unlike the communication party's observation definition, the interference party's observation includes two heterogeneous data sets; interference signals are removed from the interference party's spectrum sensing results; the interference party's observation is defined as: (10); in, This refers to the sequence of historical spectrum vectors perceived by the interfering party, defined as: (11); The length of the corresponding sequence; These are additional action data, also presented in the form of historical sequences, and defined as follows: (12); in, This indicates that the action sequence is One-Hot encoded; That is, the action of the interfering party at time t; These are the observation functions of the communicating party and the interfering party, respectively; the observation functions of both parties are implicit; due to the different geographical locations of the two parties, under the influence of the wireless signal propagation law, the observations of the spectrum state by the two parties are naturally different; The set of actions of the communicating parties; the actions of the communicating parties are defined as the combination of channels used by the two users. Each user's selectable channel comes from the set of available channels: C represents the number of available channels; furthermore, the channels for two users within each time slot should be different to avoid internal interference. The set of actions taken by the interfering party; the actions of the interfering party are defined as the channel combination in which the interfering signal is located. The channel occupied by the interference signal is different in each time slot; The revenue function of the communicating party; the revenue of the communicating party is defined as the sum of the revenues of the two communicating users: (13); For each user, the benefit criterion is to reflect the communication quality and cost as much as possible; specifically, the benefit for each user is defined as: (14); in, That is, the defined user The normalized throughput obtained For indicator functions; and These represent the weighting factors corresponding to successful and failed communication, respectively. : The profit function of the interfering party; Since the opposing parties have conflicting interests and completely opposite goals, consider the scenario of a zero-sum game, that is, the profit of the interfering party is always the opposite of the profit of the communicating party, and the sum of the profits of both parties is 0. P: State transition function; Discount factor; Used to depict combat scenarios with infinite duration; make Indicates user The strategy, From the observation space To the action space The mapping; given observations, through Calculate a value in the action space The probability distribution on; let This indicates the joint strategy of the communicating parties, making The utility function, representing the expected long-term cumulative discount benefit of the communicating party, is expressed as: (15); in, This represents the interfering party's strategy; the communication party's optimization objective is to optimize its strategy. The utility function, expressed as: (16); Since the interfering party's intention is to disrupt the communicating party's communication as much as possible, which is completely in conflict with the communicating party's interests, its optimization objective is exactly the opposite of the communicating party's; that is, the interfering party attempts to optimize its strategy. To minimize the utility function of the communicating parties; considering the zero-sum game characteristics of both sides, the intelligent spectrum adversarial problem is expressed as: (17); Solution to the intelligent spectrum adversarial problem This corresponds to the equilibrium strategy in the game theory model (POATSG); the intelligent spectrum adversarial problem involves two decision-making entities, namely the two opposing sides.

5. The method according to claim 4, characterized in that, Step 4 includes: The virtual interference is configured as follows: Observation: The observation of virtual interference is exactly the same as that of the communication party; ensure that the virtual interference party has enough information, at least no less than the information held by the communication party; Action: The virtual jamming selects two channels in each time slot to release virtual jamming signals. The action is represented as follows. The channels available for virtual interference are all derived from the set of available channels. C represents the number of available channels; furthermore, the two channels used for interference are different in each time slot; the virtual interference signal is directly added to the sensing results of the communicating party without going through the wireless channel; Benefits: Virtual interference serves as a hypothetical enemy introduced by the communicating party. The objective is completely opposite to that of the communicating party, that is, the benefit of virtual interference is the opposite of the benefit of the communicating party, and the sum of the benefits of both parties is 0; Intelligent decision-making algorithm: The intelligent decision-making algorithm and neural network architecture of virtual interference are exactly the same as those of the communication side.

6. The method according to claim 5, characterized in that, Step 5 includes: Initiating the virtual training process, the communicating party and the virtual interference respectively base their observations on... Make their own actions The virtual environment generates new spectrum states based on the actions of both parties, allowing the communicating parties to obtain new observations. The communication provider calculates its revenue based on the signal demodulation at the base station. Benefits of virtual interference This is the inverse of the communication party's benefit; at the end of each time slot, empirical data is obtained: And store the experience data in the experience replay pool; When the communicating party and the virtual jammer update their policies, the original data needs to be split into data formats that both parties can access. Specifically, the data format used by the communicating party to update its policies is as follows: The data format used for the virtual interference update strategy is: After the split is completed, both parties will update their strategies using their respective experience data. The entire virtual training process is divided into several update cycles. Both sides maintain their strategies unchanged throughout each update cycle, with updates performed only at the end of the cycle.

7. The method according to claim 6, characterized in that, Step 6 includes: To add a parameter reset period to the virtual interference. At the end of each parameter reset cycle, with probability Update neural network parameters with probability Reset the neural network parameters; Reset refers to reinitializing the neural network parameters and restarting the learning process; the parameter reset period is... With update cycle The relationships between them are integer multiples, that is: (18); The execution process is as follows: Step 61, Initialization Process: Initialize the parameters of the communication decision network 1. Evaluate the parameters of the network Experience Data Pool Parameters of the decision network for virtual interference 1. Evaluate the parameters of the network Experience Data Pool ; Set the update interval for the communicating party. Reset cycle of virtual interference Reset probability Learning rates of decision-making networks and evaluation networks Learning parameters Weighting factors ; Step 62, enter the loop. ; The communicating party obtains current spectrum observations ; According to the policy Select Action Virtual interference selects actions based on the strategy. Both parties perform their respective actions; The communicating party obtains the spectrum observation at the next moment. And calculate the revenue. Gaining benefits through virtual interference ; The communication party and the virtual interference will respectively use empirical data and Stored in their respective experience data pools , ; The process begins; if it is the end of an update cycle, the neural network parameters need to be updated. Virtual interference with probability Reset network parameters with probability Update network parameters; The communicating parties synchronously update the neural network parameters; Clear the experience data pool of communication parties and virtual interference. , .