A multi-agent cooperative anti-interference method based on spectrum data aggregation

By employing spectrum data aggregation and multi-agent collaborative anti-interference methods, the problem of algorithm performance degradation caused by incomplete sensing data was solved, achieving efficient communication in dynamic interference environments and improving the throughput of wireless networks.

CN119300166BActive Publication Date: 2025-11-11ARMY ENG UNIV OF PLA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411250013.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2025-11-11
Estimated Expiration
2044-09-06

AI Technical Summary

Technical Problem

Existing anti-interference methods based on deep reinforcement learning suffer from prolonged convergence time and degraded performance in real-world communication environments due to incomplete or erroneous sensing data, making it impossible to guarantee reliable performance in real-world environments.

Method used

A multi-agent collaborative anti-interference method based on spectrum data aggregation is adopted. By aggregating user perception data through spectrum center, a partially observable Markov decision process is established. The mean field approximation and mellowmax operator are used to train the neural network, and the Q function is optimized to achieve joint decision-making.

Benefits of technology

It effectively reduces algorithm complexity, improves convergence speed and stability, and enables a strategy to achieve long-term cumulative discount communication rate under dynamic interference environments, significantly improving the network throughput of large-scale wireless communication networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119300166B_ABST
    Figure CN119300166B_ABST
Patent Text Reader

Abstract

This invention discloses a multi-agent cooperative anti-interference method based on spectrum data aggregation. First, a multi-agent cooperative anti-interference model based on spectrum data aggregation is established. Second, a joint action is selected based on a greedy strategy. Third, the joint action is executed to obtain a reward, the current spectrum is perceived, the state transitions to the next state, and the experience is stored in the agent's experience pool. Then, random batch sampling is performed from the experience pool, the target Q-value is calculated, the gradient of the loss function is calculated, and the weight values ​​are trained and updated. The algorithm terminates when the maximum number of iterations is reached. This invention features a complete model with clear physical meaning, avoids malicious interference and mutual interference between users, significantly improves the anti-interference performance of ultra-dense networks with incomplete perceptual information, and provides a fairer frequency usage strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication anti-interference technology, and in particular, it is a multi-agent collaborative anti-interference method based on spectrum data aggregation. Background Technology

[0002] Thanks to the rapid development of wireless communication technology, the Internet of Things (IoT), vehicle-to-everything (V2X) networks, and Ad Hoc networks have also developed rapidly. The explosive growth in the number of wireless devices has led to a continuous expansion of network scale and increasingly crowded device distribution, making the already scarce and limited spectrum resources even more difficult to allocate rationally. Due to its inherent openness, the electromagnetic spectrum environment is extremely vulnerable to interference and malicious attacks. Therefore, it is crucial to study dynamic spectrum access methods for large-scale, densely distributed wireless networks under malicious interference attacks.

[0003] Deep reinforcement learning-based dynamic spectrum access methods for anti-jamming based on cognitive radio technology and machine learning have been widely studied in academia (Oroojlooyjadid A, Hajinezhad D. "A Review of Cooperative Multi-Agent Deep Reinforcement Learning", 2019). However, the dynamic spectrum access decisions of these anti-jamming methods are all based on the premise that the sampled data obtained from spectrum sensing is complete and error-free. If the data obtained from spectrum sensing is incomplete or contains errors, the convergence time of the algorithm will be significantly increased and the performance will be severely degraded. In actual communication environments, due to the performance limitations of hardware devices and data transmission distortion, the data obtained by communication devices may have varying degrees of missing information and errors. Therefore, the aforementioned DRL-based anti-jamming methods cannot guarantee the same reliable performance in actual communication environments as in simulation environments. Some literature focuses on the sensing problem. For example, the literature (Abbas W, Koutsoukos X. "Efficient Complete Coverage Through Heterogeneous SensingNodes," IEEE Wireless Communications Letters, 2015, 4(1): 14-17.) studies the location deployment problem of communication nodes with heterogeneous spectrum sensing capabilities, reducing the cost required for full sensing coverage. The literature (Hu M, Zhu Q. "Secondary user utility optimization algorithm on cooperative spectrum sensing," 2019 IEEE 5th International Conference on Computer and Communications (ICCC). IEEE, 2019: 463-467) comprehensively considers factors such as sensing energy consumption, detection probability, and transmission distance of cognitive radio devices, and obtains the optimal sensing time for each CR device based on game theory, thereby reducing sensing energy consumption.

[0004] To achieve collaborative anti-interference dynamic spectrum access under conditions of incomplete sensing information, a feasible solution is to draw on the idea of ​​"cooperation and win-win". All users upload the missing sensing data to a central controller with powerful computing capabilities. The central controller then aggregates all the missing spectrum data to obtain more complete spectrum data. Based on the data uploaded by users, the central controller trains a joint decision deep neural network to guide users' anti-interference dynamic spectrum access decisions. Summary of the Invention

[0005] The purpose of this invention is to provide a multi-agent collaborative anti-interference method based on spectrum data aggregation, which effectively characterizes the anti-interference scenario of ultra-dense networks with incomplete perception information.

[0006] The technical solution to achieve the purpose of this invention is: a multi-agent cooperative anti-interference method based on spectrum data aggregation, comprising the following steps:

[0007] Step 1: Establish a multi-agent collaborative anti-interference model based on spectrum data aggregation, which includes a group of interference devices, N users, and a frequency management center;

[0008] Step 2: Using the spectrum waterfall plot as the environmental state, the spectrum sampling data aggregated at the spectrum center as the observation state, and the access channel selected by the user as the action, define the reward function based on the environmental state, the observation state, and the access states of other users. Model the multi-agent cooperative anti-interference process of spectrum data aggregation as a partially observable Markov decision process, where the environmental state of time slot t is represented by S. t The observation state is represented as o t The reward function for user n is represented as follows

[0009] Step 3, define the state value function of user n at time t. Loop through N users and select actions based on a greedy strategy

[0010] Step 4: For user n, sense the current spectrum and execute joint actions. Receive rewards And the next state O t+1 =[o t+1 ,o t ,...,o t-Φ+2 ], and will experience The experience pool E of user n is stored.

[0011] Step 5: Randomly sample in batches from user n's experience pool E. Using state and reward as input and user action as output, a neural network is trained using gradient descent optimization. During the training process, the mellowmax operator is introduced to calculate the target state value TargetQ, and a loss function is defined accordingly.

[0012] Step 6: Repeat steps 4 and 5. When the number of iterations reaches the total number of users, jump back to step 3 until the simulation time reaches the maximum number of iterations.

[0013] Further, in step 1, a multi-agent cooperative anti-interference model based on spectrum data aggregation is established. The specific method is as follows:

[0014] The multi-agent cooperative anti-interference model based on spectrum data aggregation is a super-dense wireless network scenario, including a group of interfering devices, N users, and a frequency management center, where users are represented as follows: The entire ultra-dense network shares a frequency band, which is equally divided into M non-overlapping available channels with bandwidth b, denoted as... The frequency range of channel k is [f k -b / 2,f k [+b / 2], the SINR of user n when transmitting data in channel k is:

[0015]

[0016] Among them, f k Let k be the center frequency of channel k, and b be the channel bandwidth. U represents the data transmission power of user n on a channel. n (f) represents the power spectral density equation when user n transmits data; Let n be the channel coefficient of user n in channel k;

[0017] I n,k Let J be the power of interference experienced by user n from other users on the same channel in channel k. n,k Let the power of malicious interference suffered by user n in channel k be expressed as:

[0018]

[0019] Among them, U m (f) is the power spectral density equation for user m. Let N be the channel coefficient of user m in channel k. n f is the set of transmitters other than user n. m The center frequency of the transmission channel selected by user m; Let J(f) be the channel coefficient of the interference signal j on channel k, and J(f) be the power spectral density equation of the jammer.

[0020] δ is the power of the additive white Gaussian noise, expressed as:

[0021]

[0022] Where n(f) is the power spectral density equation of additive white Gaussian noise;

[0023] Considering that there is a minimum transmission quality requirement β in communication transmission. th Data transmission quality needs to be greater than β th For effective transmission, that is, the communication transmission rate C of user n transmitting through channel k. n,k Represented as:

[0024]

[0025] If a user can fully perceive the spectrum data of the entire available communication frequency band, then user n performs a spectrum perception result s. n (f) is:

[0026]

[0027] Define the discrete spectrum sampling value as:

[0028]

[0029] Where Δf is the resolution of the spectrum sampling, and the result obtained from one spectrum sensing is o=[o1,o2,…,o X ] T X is the number of sampling points;

[0030] Each communication user aims to access a channel that achieves a higher cumulative transmission rate. This requires users to avoid malicious interference while also preventing frequency conflicts with other users. In other words, the optimization objective for communication user n is to find the channel that maximizes the long-term cumulative discounted transmission rate, expressed as:

[0031]

[0032] Where γ (0 < γ < 1) is the discount factor, and π n For user n's strategy, The communication rate of user n in time slot t, if user n chooses channel k for communication in time slot t, then

[0033] Furthermore, in step 2, using the spectrum waterfall plot as the environmental state, the spectrum sampling data aggregated by the spectrum center as the observation state, and the access channel selected by the communication user as the action, a reward function is defined based on the environmental state, the observation state, and the access states of other users. The multi-agent cooperative anti-interference process of spectrum data aggregation is modeled as a partially observable Markov decision process. The specific method is as follows:

[0034] (1) Environmental conditions

[0035] The environmental state is defined as a complete and accurate electromagnetic spectrum state, and the environmental state in time slot t is represented by S. t ;

[0036] (2) Observation status

[0037] The observation state O is a matrix sequence of spectrum sampling data aggregated at the spectrum center over a historical period. The observation state of time slot t is represented as O. t :

[0038] O t =[o t ,o t-1 ,…,o t-Φ+1 ] T (8)

[0039] Where Φ represents the duration of historical backtracking, o t The observation data after aggregating all user-perceived results at the t-slot frequency tube center;

[0040] (3) Actions

[0041] For the access channel of user in time slot t,

[0042] (4) Reward value function

[0043] The t-slot joint reward set is The reward value for user n is related to the environmental state as S. t Observation state O t The access status of other users is related, and is represented as follows:

[0044]

[0045] Where, β n Let be the SINR of user n. Within the transmission time slot, the SINR of user n is always greater than the minimum transmission quality requirement β. th The reward value for a successful transmission is 5; if user n's SINR is lower than the minimum transmission quality requirement β at any time... th If the transmission fails, the reward value is -0.1.

[0046] (5) Partially observable Markov decision processes

[0047] A partially observable Markov decision process can be described as a six-tuple. in, Similar to the quadruple of MDP, Ω represents the state, the joint action set, the environment state transition probability, and the joint reward value function, respectively, where Ω is the observation space and O is the observed state.

[0048] Further, in step 3, define the state value function of user n at time t. Loop through N users and select actions based on a greedy strategy Specifically as follows:

[0049] The state-value function of user n at time t is approximated by reducing dimensionality using the mean field method:

[0050]

[0051] Among them, a n Use length |A n The one-hot code representation of | is that the element corresponding to the access channel position selected by user n is 1, and the rest are 0. This represents the average action impact formed by the users other than user n, i.e.

[0052]

[0053] Further, in step 4, random batch sampling is performed from user n's experience pool E. Using state and reward as input and user action as output, a neural network is trained using gradient descent optimization, where:

[0054] Using the mellowmax operator mm w The target state value TargetQ is calculated using (·), and is expressed as:

[0055]

[0056] in, w (w>0) is a temperature parameter. The reward value obtained by user n in time slot t;

[0057] The Q-function update formula after introducing the mellowmax operator and mean-field approximation is:

[0058]

[0059] in, It is the action at time t. It is the average field motion at time t;

[0060] The nonlinear fitting Q-function of a neural network is defined as follows:

[0061]

[0062] When training a neural network using the gradient descent algorithm, the gradient of the loss function is:

[0063]

[0064] A multi-agent cooperative anti-interference system based on spectrum data aggregation is characterized by implementing the multi-agent cooperative anti-interference method based on spectrum data aggregation to achieve multi-agent cooperative anti-interference based on spectrum data aggregation.

[0065] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the multi-agent cooperative anti-interference method based on spectrum data aggregation to achieve multi-agent cooperative anti-interference based on spectrum data aggregation.

[0066] A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the aforementioned multi-agent cooperative anti-interference method based on spectrum data aggregation is implemented to achieve multi-agent cooperative anti-interference based on spectrum data aggregation.

[0067] Compared with existing technologies, the significant advantages of this invention are: (1) To address the problem of a large number of users and the large dimension of the Q function based on deep reinforcement learning, the mean field approximation method is used to reduce the dimension of the Q function, thereby reducing the algorithm complexity and improving the convergence speed. To address the dependence of the classic DQN algorithm on the target network, the mellwmax operator update algorithm is used to improve the stability of the DRL algorithm update; (2) The model is complete and the physical meaning is clear. The proposed multi-agent collaborative anti-interference algorithm based on spectrum data aggregation can effectively solve the proposed model and find the strategy for long-term cumulative discount communication rate; (3) It can effectively cope with dynamic interference and well characterize the multi-agent collaborative anti-interference scenario based on spectrum data aggregation. Attached Figure Description

[0068] Figure 1 This is a block diagram of the collaborative anti-interference algorithm under incomplete perception in this invention.

[0069] Figure 2 This is the convergence analysis in Embodiment 1 of the present invention.

[0070] Figure 3 This refers to the change in normalized network throughput as a function of the user's full-band perception ratio in Embodiment 1 of the present invention.

[0071] Figure 4 This relates to the impact of the number of users with sensing capabilities on network throughput in Embodiment 2 of the present invention. Detailed Implementation

[0072] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0073] like Figure 1 As shown, this invention is a multi-agent cooperative anti-interference method based on spectrum data aggregation, and the specific steps are as follows:

[0074] Step 1: Establish a multi-agent collaborative anti-interference model based on spectrum data aggregation;

[0075] The multi-agent cooperative anti-interference model based on spectrum data aggregation is a scenario of an ultra-dense wireless network. The network includes a group of jamming devices, N communication users, and a frequency management center. The communication users are represented as... The entire ultra-dense network shares a frequency band. Considering the network has multiple communication channels, this frequency band is divided into M non-overlapping available channels with bandwidth b. The channel set is represented as...

[0076] Furthermore, the channel The frequency range is [f k -b / 2,f k +b / 2], where the SINR of user n when transmitting data in channel k is:

[0077]

[0078] Among them, f k Let k be the center frequency of channel k, and b be the channel bandwidth. U represents the data transmission power of user n on a channel. n (f) represents the power spectral density equation when user n transmits data; Let n be the channel coefficient of user n in channel k;

[0079] I n,k Let J be the power of interference experienced by user n from other users on the same channel in channel k. n,k Let the power of malicious interference suffered by user n in channel k be expressed as:

[0080]

[0081] Among them, U m (f) is the power spectral density equation for user m. Let N be the channel coefficient of user m in channel k. n f is the set of transmitters other than user n. m The center frequency of the transmission channel selected by user m; Let J(f) be the channel coefficient of the interference signal j on channel k, and J(f) be the power spectral density equation of the jammer.

[0082] δ is the power of the additive white Gaussian noise, expressed as:

[0083]

[0084] Where n(f) is the power spectral density equation of additive white Gaussian noise;

[0085] Furthermore, considering the minimum transmission quality requirement β in communication transmission... th Data transmission quality greater than βth Effective transmission refers to the communication transmission rate C of user n transmitting through channel k. n,k Represented as:

[0086]

[0087] If a user can fully perceive the spectrum data of the entire available communication frequency band, then user n performs a spectrum perception result s. n (f) is:

[0088]

[0089] Furthermore, the discrete spectrum sample value is defined as:

[0090]

[0091] Where Δf is the resolution of the spectrum sampling, and the result obtained from one spectrum sensing is o=[o1,o2,…,o X ] T X represents the number of sampling points.

[0092] Step 2: Each communication user's goal is to access a channel that achieves a higher cumulative transmission rate. This requires users to avoid malicious interference while also preventing frequency conflicts with other users. In other words, the optimization objective for user n is to find the channel that maximizes the long-term cumulative discount communication rate:

[0093]

[0094] Where γ (0 < γ < 1) is the discount factor, and π n For user n's strategy, The communication rate of user n in time slot t, if user n chooses channel k for communication in time slot t, then

[0095] Step 3: Using the spectrum waterfall plot as the environmental state, the spectrum sampling data aggregated by the spectrum center as the observation state, and the access channel selected by the user as the action, define the reward function based on the environmental state, the observation state, and the access state of other users, and model the multi-agent cooperative anti-interference process of spectrum data aggregation as a partially observable Markov decision process.

[0096] (1) Environmental conditions

[0097] The environmental state is defined as a complete and accurate electromagnetic spectrum state, and the environmental state in time slot t is represented by S. t .

[0098] (2) Observation status

[0099] The observation state O is a matrix sequence of spectrum sampling data aggregated at the spectrum center over a historical period. The observation state of time slot t is represented as O. t :

[0100] O t =[o t ,o t-1 ,…,o t-Φ+1 ] T (8)

[0101] Where Φ represents the duration of historical backtracking, o t The observation data is the result of aggregating all user-perceived results at the t-slot frequency tube center.

[0102] (3) Actions

[0103] For each user's access channel, a n ∈{1,2,…,M}.

[0104] (4) Reward value function

[0105] Joint Awards The reward value for user n is related to the environmental state as S. t Observation state O t The reward value for user n is related to the access status of other users, etc.

[0106]

[0107] Where, β n Let be the SINR of user n. Throughout the transmission time slot, the SINR of user n is consistently greater than the minimum transmission quality requirement β. th The reward value for a successful transmission is 5; if user n's SINR is lower than the minimum transmission quality requirement β at any time... th If the transmission fails, the reward value is -0.1.

[0108] (5) Partially observable Markov decision processes

[0109] A partially observable Markov decision process can be described as a six-tuple. in, Similar to the quadruple of MDP, Ω represents the state, the joint action set, the environment state transition probability, and the joint reward value function, respectively, where Ω is the observation space and O is the observed state.

[0110] Step 4, define the state value function of user n at time t. Loop through N users and select actions based on a greedy strategy Specifically as follows:

[0111] The state-value function of user n at time t can be approximated with reduced dimensionality using the mean field method:

[0112]

[0113] Among them, a n Use length |A n The one-hot code representation of | is that the element corresponding to the access channel position selected by user n is 1, and the rest are 0. This represents the average action impact formed by the users other than user n, i.e.

[0114]

[0115] Step 5: Sensing the current spectrum and executing joint actions. Receive rewards The state transitions to the next state O. t+1 =[o t+1 ,o t ,...,o t-Φ+2 ], to experience Stored into the experience pool E of agent n;

[0116] Step 6: Iterate through N users and randomly sample from the experience pool E in batches. The mellowmax operator is introduced to calculate the target state value TargetQ, and the neural network is trained using the gradient descent optimization method.

[0117] Using the mellowmax operator mm w The target state value TargetQ is calculated using (·), and is expressed as:

[0118]

[0119] in, w (w>0) is a temperature parameter. This represents the reward value obtained by user n in time slot t.

[0120] Furthermore, the Q-function update formula after introducing the mellowmax operator and the mean-field approximation is:

[0121]

[0122] in, It is the action at time t. It is the average field motion at time t;

[0123] The nonlinear fitting Q-function of a neural network is defined as follows:

[0124]

[0125] When training a neural network using the gradient descent algorithm, the gradient of the loss function is:

[0126]

[0127] Step 7: Repeat steps 5 and 6. When the number of iterations reaches the total number of users, jump back to step 4. Continue until the simulation time reaches the maximum number of iterations, at which point the algorithm ends.

[0128] In summary, each user samples the spectrum through a spectrum sensing device. Due to hardware limitations, the user's sensing capability is limited, and the obtained spectrum data may contain omissions or errors. The user then uploads the obtained spectrum sampling data and the communication throughput of the previous time slot to the frequency management center via the control link. The frequency management center centrally aggregates all the spectrum data uploaded by users to obtain more complete spectrum data. This aggregated spectrum data is then used to input a trained joint decision neural network for joint channel access decision-making. Finally, the joint access decision is distributed to each communication user via the control channel. This invention effectively addresses dynamic interference and accurately characterizes multi-agent collaborative anti-interference scenarios based on spectrum data aggregation.

[0129] Example 1

[0130] The first embodiment of the present invention is described in detail below. The system simulation is performed using Python, based on a TensorFlow neural network, and the parameter settings do not affect the generality. This embodiment verifies the effectiveness of the proposed model and method. Figure 2 The effectiveness of the proposed algorithm was verified using three comparative algorithms. The parameters were set as follows: the entire communication network shared a 20MHz spectrum, which was divided into M non-overlapping channels; the number of users was N=20; users sampled the spectrum at a resolution of Δf=100kHz, generating X=200 sampling points; and the frequency management center stored 100ms of spectrum data to construct a spectrum waterfall plot. The user communication signal used a raised cosine wave with a rise / fall factor of 0.5, the user power was 0dBm, and the minimum transmission quality requirement β was set. th =10dB. Strong discount factor λ=0.9, learning rate α=0.01, temperature parameter w=0.01, algorithm iterations 21000 times. The jammer transmits a frequency sweep jamming mode, the jamming signal is a raised cosine wave with a rise / fall factor of 0.5, power of 40dBm, frequency sweep speed of 1GHz / s, the jamming signal completes full-band jamming once every 20ms, and the background noise power is -90dBm / Hz.

[0131] First, a neural network Q is generated, with weights randomly assigned, and the initial observation state O0 is set. Second, a joint action is selected based on a greedy policy. Third, the joint action is executed to obtain a reward, the current spectrum is perceived, the state is transitioned to the next state, and the experience is stored in the agent's experience pool. Then, random batch sampling is performed from the experience pool, the target Q value is calculated, the gradient of the loss function is calculated, and the weight values ​​are trained and updated. The algorithm ends when the maximum number of iterations is reached. Figure 3 This refers to the change in network throughput with the number of users. Compared to the centralized DQN algorithm, which uses a traditional target random network mechanism, the proposed algorithm does not rely on a target neural network, resulting in more stable training and updates, and a slower rate of network throughput decline compared to the centralized DQN algorithm.

[0132] Example 2

[0133] The second embodiment of the present invention is described in detail below. The system simulation is performed using Python, based on a TensorFlow neural network, and the parameter settings do not affect the generality. This embodiment verifies the effectiveness of the proposed model and method. Figure 4 The effectiveness of the proposed algorithm was verified using three comparative algorithms. The parameters were set as follows: the entire communication network shared a 20MHz spectrum, which was divided into M non-overlapping channels; the number of users was N=20; users sampled the spectrum at a resolution of Δf=100kHz, generating X=200 sampling points; and the frequency management center stored 100ms of spectrum data to construct a spectrum waterfall plot. The user communication signal used a raised cosine wave with a rise / fall factor of 0.5, the user power was 0dBm, and the minimum transmission quality requirement β was set. th =10dB. Strong discount factor λ=0.9, learning rate α=0.01, temperature parameter w=0.01, algorithm iterations 21000 times. The jammer transmits a frequency sweep jamming mode, the jamming signal is a raised cosine wave with a rise / fall factor of 0.5, power of 40dBm, frequency sweep speed of 1GHz / s, the jamming signal completes full-band jamming once every 20ms, and the background noise power is -90dBm / Hz.

[0134] With a fixed user base and user perception capabilities, the more devices that have spectrum sensing capabilities, the more complete the spectrum the frequency management center can obtain when performing spectrum data aggregation and processing.

[0135] Figure 4 This relates to the impact of the number of users with sensing capabilities on network throughput. As the number of iterations increases, the normalized network throughput continuously improves and converges; the more users with sensing capabilities, the higher the normalized network throughput during the convergence phase.

[0136] In summary, the multi-agent cooperative anti-interference model based on spectrum data aggregation proposed in this invention primarily considers the attacks on communication networks by malicious interference devices and the limited spectrum sensing capabilities of communication equipment. Addressing the issue of limited user sensing capabilities preventing the perception of the complete spectrum, it proposes a spectrum data aggregation approach, centrally processing all user-perceived spectrum data to obtain a more complete spectrum sensing result, which is more practically significant than traditional models. The proposed multi-agent cooperative anti-interference algorithm based on spectrum data aggregation can effectively solve the proposed model. To address the problem of a large number of users and the high dimensionality of the Q-function based on deep reinforcement learning, a mean-field approximation method is used to reduce the dimensionality of the Q-function, thereby reducing algorithm complexity and improving convergence speed. To address the dependence of the classic DQN algorithm on the target network, a mellwmax operator update algorithm is adopted to improve the stability of the DRL algorithm update. Under malicious interference environments, the proposed cooperative anti-interference algorithm can reduce the impact of incomplete spectrum sensing on the performance and convergence of the DRL algorithm, significantly improving the network throughput of large-scale wireless communication networks.

Claims

1. A multi-agent cooperative anti-interference method based on spectrum data aggregation, characterized in that, Includes the following steps: Step 1: Establish a multi-agent collaborative anti-interference model based on spectrum data aggregation, which includes a group of interference devices, N users, and a frequency management center; Step 2: Using the spectrum waterfall plot as the environmental state, the spectrum sampling data aggregated at the spectrum center as the observation state, and the access channel selected by the user as the action, define the reward function based on the environmental state, the observation state, and the access states of other users. Model the multi-agent cooperative anti-interference process of spectrum data aggregation as a partially observable Markov decision process, where the environmental state of time slot t is represented by S. t The observed state is represented as o t The reward function for user n represents Step 3, define the state value function of user n at time t. Loop through N users and select actions based on a greedy strategy Step 4: For user n, sense the current spectrum and execute joint actions. Receive rewards And the next state O t+1 =[o t+1 ,o t ,...,o t-Φ+2 ], and will experience The experience pool E of user n is stored. Step 5: Randomly sample in batches from user n's experience pool E. Using state and reward as input and user action as output, a neural network is trained using gradient descent optimization. During the training process, the mellowmax operator is introduced to calculate the target state value TargetQ, and a loss function is defined accordingly. Step 6: Repeat steps 4 and 5. When the number of iterations reaches the total number of users, jump back to step 3 until the simulation time reaches the maximum number of iterations.

2. The multi-agent cooperative anti-interference method based on spectrum data aggregation according to claim 1, characterized in that, Step 1: Establish a multi-agent cooperative anti-interference model based on spectrum data aggregation. The specific method is as follows: The multi-agent cooperative anti-interference model based on spectrum data aggregation is a super-dense wireless network scenario, including a group of interfering devices, N users, and a frequency management center, where users are represented as follows: The entire ultra-dense network shares a frequency band, which is equally divided into M non-overlapping available channels with bandwidth b, denoted as... The frequency range of channel k is [f k -b / 2,f k [+b / 2], the SINR of user n when transmitting data in channel k is: Among them, f k Let k be the center frequency of channel k, and b be the channel bandwidth. U represents the data transmission power of user n on a channel. n (f) represents the power spectral density equation when user n transmits data; Let n be the channel coefficient of user n in channel k; I n,k Let J be the power of interference experienced by user n from other users on the same channel in channel k. n,k Let the power of malicious interference suffered by user n in channel k be expressed as follows: Among them, U m (f) is the power spectral density equation for user m. Let N be the channel coefficient of user m in channel k. n f is the set of transmitters other than user n. m The center frequency of the transmission channel selected by user m; Let J(f) be the channel coefficient of the interference signal j on channel k, and J(f) be the power spectral density equation of the jammer. δ is the power of the additive white Gaussian noise, expressed as: Where n(f) is the power spectral density equation of additive white Gaussian noise; Considering that there is a minimum transmission quality requirement β in communication transmission. th Data transmission quality needs to be greater than β th For effective transmission, that is, the communication transmission rate C of user n transmitting through channel k. n,k Represented as: If a user can fully perceive the spectrum data of the entire available communication frequency band, then user n performs a spectrum perception result s. n (f) is: Define the discrete spectrum sampling value as: Where Δf is the resolution of the spectrum sampling, and the result obtained from one spectrum sensing is o=[o1,o2,…,o X ] T X is the number of sampling points; Each communication user aims to access a channel that achieves a higher cumulative transmission rate. This requires users to avoid malicious interference while also preventing frequency conflicts with other users. In other words, the optimization objective for communication user n is to find the channel that maximizes the long-term cumulative discounted transmission rate, expressed as: Where γ is the discount factor, π n For user n's strategy, The communication rate of user n in time slot t, if user n chooses channel k for communication in time slot t, then 3. The multi-agent cooperative anti-interference method based on spectrum data aggregation according to claim 1, characterized in that, Step 2: Using the spectrum waterfall plot as the environmental state, the spectrum sampling data aggregated at the spectrum center as the observation state, and the access channel selected by the communication user as the action, define the reward function based on the environmental state, the observation state, and the access states of other users. Model the multi-agent cooperative anti-interference process of spectrum data aggregation as a partially observable Markov decision process. The specific method is as follows: (1) Environmental conditions The environmental state is defined as a complete and accurate electromagnetic spectrum state, and the environmental state in time slot t is represented by S. t ; (2) Observation status The observation state O is a matrix sequence of spectrum sampling data aggregated at the spectrum center over a historical period. The observation state of time slot t is represented as O. t : The t =[o t ,o t-1 ,…,o t-Φ+1 ] T (8) Where Φ represents the duration of historical backtracking, o t The observation data after aggregating all user-perceived results at the t-slot frequency tube center; (3) Actions For the access channel of user in time slot t, (4) Reward value function The t-slot joint reward set is The reward value for user n is related to the environmental state as S. t Observation state O t The access status of other users is related, and is represented as follows: Where, β n Let be the SINR of user n. Within the transmission time slot, the SINR of user n is always greater than the minimum transmission quality requirement β. th The reward value for a successful transmission is 5; if user n's SINR is lower than the minimum transmission quality requirement β at any time... th If the transmission fails, the reward value is -0.

1. (5) Partially observable Markov decision processes A partially observable Markov decision process can be described as a six-tuple. Among them, S, P, Similar to the quadruple of MDP, Ω represents the state, the joint action set, the environment state transition probability, and the joint reward value function, respectively, where Ω is the observation space and O is the observed state.

4. The multi-agent cooperative anti-interference method based on spectrum data aggregation according to claim 1, characterized in that, Step 3, define the state value function of user n at time t. Loop through N users and select actions based on a greedy strategy Specifically as follows: The state-value function of user n at time t is approximated by reducing dimensionality using the mean field method: Among them, a n Use length |A n The one-hot code representation of | is that the element corresponding to the access channel position selected by user n is 1, and the rest are 0. This represents the average action impact formed by the users other than user n, i.e.

5. The multi-agent cooperative anti-interference method based on spectrum data aggregation according to claim 1, characterized in that, Step 4: Randomly sample from user n's experience pool E in batches. Using state and reward as input and user action as output, a neural network is trained using gradient descent optimization, where: Using the mellowmax operator mm w The target state value TargetQ is calculated using (·), and is expressed as: in, For temperature parameters, r t n The reward value obtained by user n in time slot t; The Q-function update formula after introducing the mellowmax operator and the mean-field approximation is: Where α is the learning rate. It is the action at time t. It is the average field motion at time t; The nonlinear fitting Q-function of a neural network is defined as follows: When training a neural network using the gradient descent algorithm, the gradient of the loss function is:

6. A multi-agent cooperative anti-interference system based on spectrum data aggregation, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it is characterized in that, Implement the multi-agent cooperative anti-interference method based on spectrum data aggregation as described in any one of claims 1-5.

7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the multi-agent cooperative anti-interference method based on spectrum data aggregation as described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, wherein when executed by a processor, the computer program implements the multi-agent cooperative anti-interference method based on spectrum data aggregation as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Master user-friendly anti-interference dynamic spectrum access method

    CN113938897A

  • Dynamic spectrum multi-domain anti-interference method and system based on cognitive anti-interference model

    CN115276858A