Reactive interference-oriented secure cognitive radio network power distribution method

By constructing a random reactive interference model and a deep reinforcement learning framework, the power allocation of secondary users is optimized, the problem of limited communication rate of cognitive radio networks under reactive interference is solved, and efficient and secure communication is achieved.

CN120614680APending Publication Date: 2025-09-09BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510622509.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing cognitive radio networks based on deep reinforcement learning are unable to effectively deal with the randomness and concealment of interferers in reactive interference scenarios, resulting in limited communication rates for secondary users and susceptibility to interference detection.

Method used

A reactive interference model including time-varying channel gain and randomness is constructed. The dynamic optimal detection threshold of the interferer is derived by combining energy detection theory. A Markov decision process is constructed based on the deep reinforcement learning framework. The CR network power allocation scheme is solved through a dual-deep Q network, and the transmission power allocation of secondary users is optimized to maximize the communication rate and avoid interference.

Benefits of technology

The anti-interference capability of the cognitive radio network is enhanced, achieving high communication rate and secure transmission for downstream users under reactive interference scenarios, and meeting the secure transmission and rate requirements of the CR network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120614680A_ABST
    Figure CN120614680A_ABST
Patent Text Reader

Abstract

The invention discloses a reactive interference-oriented secure cognitive radio network power allocation method, and belongs to the field of communication resource allocation and anti-interference secure communication of a cognitive radio network. The implementation method comprises the following steps: analyzing secondary user communication requirements and disturber behavior characteristics, constructing a CR network model containing time-varying channel gain and a random reactive interference model, and deducing an optimal dynamic energy detection threshold of a disturber; constructing a CR network power distribution optimization problem by taking maximization of a secondary user communication rate and avoiding of reactive interference attacks as targets; aiming at the dynamic uncertainty of channel gain and interference behaviors, constructing a Markov decision process based on a deep reinforcement learning framework, and giving definitions of basic elements such as states, actions, rewards and the like; and constructing a double-depth Q network, and solving to obtain a CR network power distribution scheme, thereby realizing the power distribution of the CR network in the reactive interference scene. The method has the advantages of high communication security, high robustness in a complex environment and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communication resource allocation and anti-interference safety communication of cognitive radio networks, and relates to a reactive interference-oriented safety cognitive radio network power allocation method. Background Art

[0002] In recent years, with the rapid development of emerging communications industries such as the Industrial Internet of Things, smart healthcare, and smart cities, spectrum resource shortages have become increasingly prominent. Cognitive Radio (CR) technology has emerged to address this problem. This technology dynamically senses idle spectrum resources in the surrounding wireless environment and allows primary users (PUs) and secondary users (SUs) to share spectrum without causing interference. This significantly improves spectrum utilization efficiency and provides important technical support for the sustainable development of future wireless communication networks. However, the openness and dynamic nature of CR networks also present unprecedented security challenges, the most threatening of which is reactive jamming. Unlike traditional jamming methods, reactive jammers employ an intelligent "listen and jam" strategy, transmitting jamming signals only when they detect legitimate user communications. This "on-demand jamming" approach not only significantly reduces the attacker's energy consumption but also significantly improves the stealth and targeted nature of the jamming. Therefore, effectively countering reactive jamming attacks has become a key issue that needs to be addressed in the field of cognitive radio networks.

[0003] In recent years, with the rapid development of deep learning technology, especially the successful application of deep reinforcement learning (DRL) technology to complex decision-making problems, DRL-based anti-interference communication has gradually become a research hotspot. In CR networks, DRL does not rely on pre-defined interference models or prior knowledge. Instead, it can autonomously learn optimal strategies through interaction with the dynamic environment, significantly improving the anti-interference and communication performance of CR networks. However, existing DRL-based research mostly assumes that the interferer adopts a fixed or predictable interference pattern (such as single-tone interference, multi-tone interference, and frequency sweeping interference), and rarely considers reactive interference scenarios. In existing research on reactive interference, the reactive interference model often assumes that the interferer adopts a preset "detect-and-interfere" mode (such as selecting the interference channel based on maximum power), which is easily learned and adaptively avoided by the DRL agent. Therefore, introducing a randomized reactive interference strategy into the DRL-driven secure CR network architecture can provide an innovative adaptive security protection strategy for CR networks in complex electromagnetic environments. It has important theoretical value and practical application prospects in highly dynamic adversarial scenarios such as the Industrial Internet of Things, smart healthcare, and emergency communications. Summary of the Invention

[0004] To address the issues of limited communication rates and susceptibility to detection by interferers in reactive jamming scenarios for secondary users in CR networks, the present invention aims to provide a secure CR network power allocation method for reactive jamming. By analyzing the communication requirements of secondary users and the dynamic behavior characteristics of interferers, a CR network model is constructed that includes a time-varying channel gain and a stochastic reactive jamming model, and the optimal dynamic energy detection threshold for the interferer is derived. With the goal of maximizing the communication rate of secondary users while avoiding reactive jamming attacks, an optimization problem for CR network power allocation is constructed. A Markov decision process is constructed based on a deep reinforcement learning framework to address the dynamic uncertainty of channel gain and jamming behavior, defining basic elements such as state, action, and reward. A dual-deep Q network is constructed to solve the CR network power allocation solution, thereby achieving power allocation for CR networks in reactive jamming scenarios.

[0005] The purpose of the present invention is achieved through the following technical solutions.

[0006] The present invention discloses a reactive interference-oriented secure cognitive radio network power allocation method, comprising the following steps:

[0007] Step 1: Combining the cognitive radio network working mode with the loss characteristics of the wireless channel, a secure CR network scenario model is constructed. The channel model of the secondary user-base station link and the channel model of the secondary user-interferer link are constructed, and the achievable communication rate index is characterized.

[0008] The secure CR network scenario includes N primary users, one secondary user, one interferer, and one communication base station. The distance between the secondary user and the communication base station is d, and the distance between the secondary user and the interferer is d. w The spectrum is divided into N frequency bands with a bandwidth of B0. The primary user selects different frequency bands to communicate with the communication base station to avoid mutual interference. The secondary user performs spectrum sensing and spectrum access in each time slot. The number of time slots is T. In the tth time slot, the secondary user obtains the number of idle frequency bands in the current time slot through spectrum sensing, which is recorded as N t , and then randomly select an idle frequency band with power P t Communicate with the base station, where P t From the predefined power set {P k |k=1,2,...,K}, K is the number of predefined powers, and the carrier frequency of the transmitted signal is recorded as f c At the same time, the jammer randomly selects an idle frequency band for energy detection, determines whether the secondary user is communicating in the frequency band, and chooses whether to transmit a signal to interfere with the secondary user.

[0009] The channel model of the secondary user-base station link is a Rice channel, which is expressed as:

[0010]

[0011] Where κ is the Rice factor, which represents the power ratio of the direct path to the scattered path of the channel. Indicates the random phase of the direct path arrival signal, which obeys the uniform distribution. represents a random variable with complex circular Gaussian distribution. 2 =10 -PL / 10 represents the free space path loss of the secondary user-base station link, PL is expressed as:

[0012]

[0013] The channel model of the secondary user-interferer link is:

[0014]

[0015] in represents the free space path loss of the secondary user-interferer link, PL w Expressed as:

[0016]

[0017] In the tth time slot, the achievable communication rate of the secondary user is expressed as:

[0018]

[0019] in is the noise power received by the base station.

[0020] Step 2: Construct a random reactive interference model and combine it with energy detection theory to derive the dynamic optimal energy detection threshold of the interferer. Construct a cognitive radio network power allocation optimization problem for reactive interference.

[0021] In the CR network of step 1, the interferer has p(1)=1 / N t The probability of selecting the same frequency band as the secondary user is p(0)=(N t -1) / N t The probability of selecting other frequency bands is . Perform energy detection on the selected frequency band:

[0022]

[0023] Where D represents the detection statistic, y[i] represents the i-th signal sample received by the interferer, M represents the total number of samples of the interferer's sampled signal in a time slot, ε opt Indicates the optimal energy detection threshold of the interferer. It means that the interferer only receives noise in the selected frequency band and does not transmit interference. It means that the interferer receives noise and secondary user signals in the selected frequency band and transmits interference signals. It is expressed as follows:

[0024]

[0025] Where x[i] represents the i-th signal sample from the secondary user received by the interferer, and n[i] represents the variance. Additive Gaussian white noise. Combining equation (6) and equation (7), and Re-expressed as follows:

[0026]

[0027] Where D0 and D1 represent and The test statistics in the case all obey Gaussian distribution. max represents the maximum transmit power of the secondary user, represents the variance of the test statistic.

[0028] The total probability of detecting an interferer error is expressed as a function of the energy detection threshold ε:

[0029] p e (ε)=p(1)p MD +p(0)p FA (9)

[0030] where p MD =p(D1<ε) represents the probability of missed detection, p FA =p(D0≥ε) represents the false alarm probability. According to the properties of Gaussian distribution, equation (9) can be further expressed as:

[0031]

[0032] in Denotes the Gaussian Q function. Derivative of Equation (10) with respect to ε, let get:

[0033] p(1)f1(ε opt )=p(0)f0(ε opt ) (11)

[0034] Where f0(x) and f1(x) represent the probability density functions of Gaussian variables D0 and D1, respectively, and the optimal energy detection threshold ε of the interferer is opt Expressed as:

[0035]

[0036] The optimization goal is to maximize the total communication rate of the secondary user while avoiding reactive jamming attacks. The optimization variable is the transmit power of the secondary user in each time slot, which is expressed as the set The optimization problem is then expressed as:

[0037]

[0038] where R P is a constant penalty term used to represent the impact of interference attacks on the communication rate. Binary auxiliary variable φ t =1 means that the transmission activity of the secondary user in the tth time slot is detected by the jammer and is attacked by reactive jamming. At this time, the jammer selects the same frequency band as the secondary user for monitoring and meets the received signal power φ t =0 means that the secondary user is not subject to reactive interference.

[0039] Step 3: Define the basic elements, including state, action, and reward. Construct a dual-depth Q network to solve the CR network power allocation scheme. This scheme is used to implement power allocation for the CR network in a reactive interference scenario.

[0040] The power allocation optimization problem in step 2 is modeled as a Markov decision process (MDP). In MDP, the secondary user acts as an intelligent agent and observes the current state s at each time slot t. t , and select action a according to the strategy t , and obtain the reward r for evaluating the long-term effectiveness of the action t Through continuous interaction with the environment, the agent gradually optimizes its action selection strategy to obtain the maximum long-term reward, that is, the total communication rate of secondary users. The basic elements of MDP are defined as follows.

[0041] State: The state represents the agent's observation of the environment at a given moment. In a CR network scenario, secondary users observe the occupancy of each frequency band through spectrum sensing. Therefore, the state of the tth time slot can be represented by the number of idle frequency bands, expressed as:

[0042] s t =N t (14)

[0043] Action: An action represents the decision made by an agent at a certain moment based on its current state. In the CR network scenario, the action at the tth time slot represents the secondary user’s choice of transmit power. It can be expressed as:

[0044] at =P t ,P t ∈{P k |k=1,2,...,K} (15)

[0045] Reward: Reward is used to evaluate the effectiveness of the agent's actions at a certain moment and is a direct reflection of the optimization goal. In the CR network scenario, the reward for the tth time slot is expressed as:

[0046] r t =R t -φ t R P (16)

[0047] Construct a Double Deep Q Network (DDQN), which evaluates the neural network Q(s t ,a t ;θ) and the target neural network Q′(s t ,a t ; θ′), where θ and θ′ represent the internal parameters of the two neural networks. The two neural networks have the same structure, and the input is the state s of the tth time slot t , the output is a set of action value vectors, representing the state s t The rewards for executing each action are as follows.

[0048] In the training process of the dual deep Q network, the secondary user acts as an agent and first collects training data based on a greedy strategy. The agent selects a random action with probability α and selects an action with probability 1-α. Where α∈[0,1] is the exploration rate. The training data (s t ,a t ,r t ,s t+1 ) are stored in the experience pool in the form of tuples, and the agent randomly samples small batches of data from the experience pool to train and evaluate the neural network. The loss function of the evaluation neural network is defined as:

[0049]

[0050] in is the target network Q value for the tth time slot, expressed as:

[0051]

[0052] Where γ∈[0,1] is the discount factor, is the action set. The evaluation neural network continuously backpropagates through Equation (17) and updates the network parameters θ. After each F-step training, the parameters θ of the evaluation neural network are copied to the parameters θ' of the target neural network. After continuous iterative updates of the dual-depth Q network, the optimal evaluation neural network is obtained. In the t-th time slot, the optimal allocation scheme for secondary users is expressed as:

[0053]

[0054] where θ opt represents the network parameters of the optimal evaluation neural network. Secondary users execute power allocation schemes based on the spectrum sensing results of each time slot, thus achieving secure power allocation in reactive interference-resistant cognitive radio networks.

[0055] Beneficial effects:

[0056] 1. Unlike existing DRL-based anti-interference cognitive radio networks, the present invention discloses a power allocation method for secure cognitive radio networks oriented to reactive interference, which considers the communication security issues in reactive interference scenarios. By constructing a random reactive interference model and combining it with energy detection theory to derive the dynamic optimal detection threshold of the interferer, the anti-interference capability of the CR network is enhanced.

[0057] 2. The present invention discloses a secure cognitive radio network power allocation method for reactive interference, which constructs a CR network power allocation optimization problem for reactive interference, and achieves a higher achievable communication rate while ensuring that the transmission signal of the secondary user is not detected by the interferer.

[0058] 3. The present invention discloses a secure cognitive radio network power allocation method for reactive interference. It constructs a Markov decision process based on the DRL framework, defines basic elements such as state, action, and reward, and constructs a dual-depth Q network to solve the CR network power allocation scheme. According to the CR network power allocation scheme, the power allocation of the CR network in the reactive interference scenario is realized to meet the secure transmission and rate requirements of the CR network. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 The present invention is a flowchart of a reactive interference-oriented secure cognitive radio network power allocation method.

[0060] Figure 2 This is a scenario diagram of a reactive interference-oriented secure cognitive radio network power allocation method according to the present invention.

[0061] Figure 3 The diagram is a relationship diagram between the secondary user transmission power and the number of idle frequency bands in a reactive interference-oriented secure cognitive radio network power allocation method according to the present invention.

[0062] Figure 4 The invention discloses a received signal power distribution histogram of a secure cognitive radio network power allocation method oriented to reactive interference.

[0063] Figure 5 This is a training effect diagram of a reactive interference-oriented secure cognitive radio network power allocation method according to the present invention.

[0064] Figure 6 This is a communication performance diagram corresponding to different Ricean factors and subchannel bandwidths of a reactive interference-oriented secure cognitive radio network power allocation method of the present invention. DETAILED DESCRIPTION

[0065] The specific implementation of the present invention is further described below in conjunction with the technical solution.

[0066] like Figure 1 As shown, this embodiment discloses a method for power allocation in a secure cognitive radio network oriented to reactive interference, and the specific implementation steps are as follows:

[0067] Step 1: Combining the cognitive radio network operating mode with the loss characteristics of the wireless channel, a secure CR network scenario model is constructed. Channel models for the secondary user-base station link and the secondary user-interferer link are constructed, and the achievable communication rate indicator is characterized.

[0068] A smart medical communication system consists of N key medical devices (primary users), an emergency data transmission terminal (secondary user), an interferer, and a communication base station. The distance between the secondary user and the base station is d, and the distance between the secondary user and the interferer is d. w , the relative distance between each node remains unchanged; the spectrum is divided into N frequency bands with bandwidth B0, and the primary user selects different frequency bands to transmit medical data to the base station to avoid mutual interference; the secondary user performs spectrum sensing and spectrum access in each time slot, and the number of time slots is T; in the tth time slot, the secondary user obtains the number of idle frequency bands in the current time slot through spectrum sensing, which is recorded as N t , and then randomly select an idle frequency band with power P t Communicate with the base station, where P t From the predefined power set {P k |k=1,2,...,K}, K is the number of predefined powers, and the carrier frequency of the transmitted signal is recorded as f c At the same time, the jammer randomly selects an idle frequency band for energy detection to determine whether the secondary user is communicating in the frequency band, and then chooses whether to transmit a signal to interfere with the secondary user. The attack may cause the loss of critical medical data.

[0069] The channel model of the secondary user-base station link is a Rice channel, which is expressed as:

[0070]

[0071] Where κ is the Rice factor, which represents the power ratio of the channel direct path to the scattered path; Indicates the random phase of the direct path arrival signal, which obeys the uniform distribution; represents a random variable with complex circular Gaussian distribution; where ξ 2 =10 -PL / 10 represents the free space path loss of the secondary user-base station link, PL is expressed as:

[0072]

[0073] The channel model of the secondary user-interferer link is:

[0074]

[0075] in represents the free space path loss of the secondary user-interferer link, PL w Expressed as:

[0076]

[0077] In the tth time slot, the achievable communication rate of the secondary user is expressed as:

[0078]

[0079] in is the noise power received by the base station.

[0080] Specifically in this embodiment, N=10, f c =5GHz, B0=100MHz, d=30m, d w =200m,κ=30, The number of idle frequency bands in each time slot N t Obeying the 2-10 discrete uniform distribution, K = 4, the set of secondary user transmit powers is {10mW, 20mW, 30mW, 40mW};

[0081] Step 2: Construct a random reactive interference model and derive the dynamic optimal energy detection threshold of the interferer by combining it with energy detection theory; construct a cognitive radio network power allocation optimization problem for reactive interference;

[0082] In the CR network of step 1, the interferer has p(1)=1 / N t The probability of selecting the same frequency band as the secondary user is p(0)=(Nt -1) / N t The probability of selecting other frequency bands; energy detection is performed on the selected frequency band:

[0083]

[0084] Where D represents the detection statistic, y[i] represents the i-th signal sample received by the interferer, M represents the total number of samples of the interferer's sampled signal in a time slot, ε opt represents the optimal energy detection threshold of the interferer; It means that the interferer only receives noise in the selected frequency band and does not transmit interference; It indicates that the interferer receives noise and secondary user signals in the selected frequency band and transmits an interference signal; It is expressed as follows:

[0085]

[0086] Where x[i] represents the i-th signal sample from the secondary user received by the interferer, and n[i] represents the variance. Additive Gaussian white noise; Combining Equation (25) and Equation (26), the detection statistic obeys the chi-square distribution and is approximately Gaussian when the number of sampling points is large enough. and Re-expressed as follows:

[0087]

[0088] Where D0 and D1 represent and The test statistics in the case all obey Gaussian distribution; P max represents the maximum transmit power of the secondary user, represents the variance of the test statistic;

[0089] The total probability of detecting an interferer error is expressed as a function of the energy detection threshold ε:

[0090] p e (ε)=p(1)p MD +p(0)p FA (28)

[0091] where p MD =p(D1<ε) represents the probability of missed detection, p FA =p(D0≥ε) represents the false alarm probability. According to the properties of Gaussian distribution, equation (28) can be further expressed as:

[0092]

[0093] in Denotes the Gaussian Q function; Derivative of Equation (29) with respect to ε, let get:

[0094] p(1)f1(ε opt )=p(0)f0(ε opt ) (30)

[0095] Where f0(x) and f1(x) represent the probability density functions of Gaussian variables D0 and D1, respectively, and the optimal energy detection threshold ε of the interferer is opt Expressed as:

[0096]

[0097] The optimization goal is to maximize the total communication rate of the secondary user while avoiding reactive jamming attacks; the optimization variable is the secondary user's transmit power in each time slot, which is expressed as the set The optimization problem is then expressed as:

[0098]

[0099] where R P is a constant penalty term used to represent the impact of interference attacks on the communication rate; the binary auxiliary variable φ t =1 means that the transmission activity of the secondary user in the tth time slot is detected by the jammer and is attacked by reactive jamming. At this time, the jammer selects the same frequency band as the secondary user for monitoring and meets the received signal power φ t =0 means that the secondary user is not subject to reactive interference;

[0100] Specifically in this embodiment, T=200, P max =50mW, R P =20Mbps;

[0101] Step 3: Based on the deep reinforcement learning framework, a Markov decision process is constructed to define basic elements such as state, action, and reward. A dual-depth Q network is constructed to solve the CR network power allocation solution.

[0102] The power allocation optimization problem in step 2 is modeled as an MDP. In the MDP, the secondary user acts as an intelligent agent and observes the current state s in each time slot t. t , and select action a according to the strategy t , and obtain the reward r for evaluating the long-term effectiveness of the action t Through continuous interaction with the environment, the agent gradually optimizes its action selection strategy to obtain the maximum long-term reward, that is, the total communication rate of secondary users. The basic elements of MDP are defined as follows:

[0103] State: The state represents the agent's observation of the environment at a certain moment. In the CR network scenario, secondary users observe the occupancy of each frequency band through spectrum sensing. Therefore, the state of the t-th time slot can be represented by the number of idle frequency bands, expressed as:

[0104] s t =N t (33)

[0105] Action: An action represents the decision made by the agent at a certain moment based on the current state. In the CR network scenario, the action at the tth time slot represents the secondary user's choice of transmit power. It is expressed as:

[0106] a t =P t ,P t ∈{P k |k=1,2,...,K} (34)

[0107] Reward: Reward is used to evaluate the effectiveness of the agent's actions at a certain moment and is a direct reflection of the optimization goal. In the CR network scenario, the reward for the tth time slot is expressed as:

[0108] r t =R t -φ t R P (35)

[0109] Construct a dual-depth Q network, that is, evaluate the neural network Q(s t ,a t ;θ) and the target neural network Q′(s t ,a t ; θ′), where θ and θ′ represent the internal parameters of the two neural networks; the two neural networks have the same structure, and the input is the state s of the tth time slot t , the output is a set of action value vectors, representing the state s t The reward for executing each action;

[0110] In the training process of the dual-depth Q network, the secondary user acts as an intelligent agent and first collects training data based on a greedy strategy; the intelligent agent selects a random action with probability α and selects an action with probability 1-α. Where α∈[0,1] is the exploration rate; the training data collected by the agent (s t ,a t ,r t ,s t+1 ) are stored in the experience pool in the form of tuples, and the agent randomly samples small batches of data from the experience pool to train and evaluate the neural network; the loss function of the evaluation neural network is defined as:

[0111]

[0112] in is the target network Q value for the tth time slot, expressed as:

[0113]

[0114] Where γ∈[0,1] is the discount factor, is the action set; the evaluation neural network continuously backpropagates through formula (36) and updates the network parameters θ; after each F-step training, the parameters θ of the evaluation neural network are copied to the parameters θ' of the target neural network; after continuous iterative updates of the dual-depth Q network, the optimal evaluation neural network is obtained; in the t-th time slot, the optimal allocation scheme for secondary users is expressed as:

[0115]

[0116] where θ opt represents the network parameters of the optimal evaluation neural network; the secondary user executes the power allocation scheme according to the spectrum sensing results of each time slot, thereby realizing the power allocation of the secure cognitive radio network oriented to reactive interference;

[0117] Specifically in this embodiment, a fully connected layer is used as a hidden layer to build an evaluation neural network and a target neural network. The number of fully connected layers is 2, the number of neurons is 64, and the RELU function is used as the hidden layer activation function. The exploration rate α decreases linearly from the initial value of 0.95 to 0.001 during the training process. γ = 0.95, the learning rate is 0.001, and the optimizer is Adam. The amount of training data per round is 512, the amount of data that the experience pool can accommodate is 800, and F = 200.

[0118] According to the above parameters, the simulation results are sorted and plotted. Figure 3 As shown in , as the number of idle frequency bands in the system increases, the detection threshold of the interferer increases accordingly, so the secondary user can choose a larger transmission power. The impact of dynamic channel gain on the received signal power is shown in Figure 4 As shown in the figure, the histogram of 2000 data sets shows that for a given transmit power, the distribution of received signal power is somewhat random. For the same transmit power level, a transmit power level is considered safe only when the majority of received signal powers exceed the detection threshold. Figure 5The training results of DDQN are presented, along with a conservative scheme for comparison. In this conservative scheme, the interferer is assumed to always transmit in the same frequency band as the secondary user for detection. Therefore, the secondary user will prefer to use a lower transmit power to avoid interference attacks. As can be seen from the figure, at the end of training, both schemes can effectively avoid the interferer's reactive interference. Therefore, the proposed scheme can achieve higher communication performance while ensuring communication security. Figure 6 The communication rates corresponding to different Ricean factors and subchannel bandwidths are shown. It can be seen that the communication system grows linearly with the subchannel bandwidth, while an increase in the Ricean factor leads to a decrease in the communication rate, reflecting the impact of Ricean channel uncertainty on the power decisions of secondary users. In summary, the power allocation method for secure cognitive radio networks oriented to reactive interference can meet the communication security and rate requirements of cognitive radio networks.

[0119] The above specific description further illustrates the purpose, technical solutions and beneficial effects of the invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A secure cognitive radio network power allocation method for reactive interference, characterized by: The steps include: Step 1: Combine the cognitive radio network operating mode with the loss characteristics of the wireless channel to build a secure CR network scenario model, construct the channel model of the secondary user-base station link and the channel model of the secondary user-interferer link, and characterize the achievable communication rate indicator; Step 2: Construct a random reactive interference model and derive the dynamic optimal energy detection threshold of the interferer by combining it with energy detection theory; construct a cognitive radio network power allocation optimization problem for reactive interference; Step 3: Based on the deep reinforcement learning framework, a Markov decision process is constructed to define the basic elements such as state, action, and reward. The basic elements include state, action, and reward. A dual-deep Q network is constructed to solve the CR network power allocation scheme. According to the CR network power allocation scheme, the power allocation of the CR network in the reactive interference scenario is realized.

2. The method according to claim 1, wherein: The secure CR network scenario includes N primary users, one secondary user, one interferer, and one communication base station. The distance between the secondary user and the communication base station is d, and the distance between the secondary user and the interferer is d. w ; Divide the spectrum into N frequency bands with bandwidth B0, and the primary user selects different frequency bands to communicate with the communication base station to avoid mutual interference; The secondary user performs spectrum sensing and spectrum access in each time slot, and the number of time slots is T. In the tth time slot, the secondary user obtains the number of idle frequency bands in the current time slot through spectrum sensing, which is recorded as N t , and then randomly selects an idle frequency band to communicate with the base station with a power of Pt, where Pt is selected from a predefined power set {P k |k=1,2,...,K}, K is the number of predefined powers, and the carrier frequency of the transmitted signal is recorded as f c At the same time, the interferer randomly selects an idle frequency band for energy detection, determines whether the secondary user is communicating in the frequency band, and chooses whether to transmit a signal to interfere with the secondary user.

3. The method according to claim 1, wherein: The channel model of the secondary user-base station link is a Rice channel, which is expressed as: Where κ is the Rice factor, which represents the power ratio of the channel direct path to the scattered path; Indicates the random phase of the direct path arrival signal, which obeys the uniform distribution; represents a random variable with complex circular Gaussian distribution; where ξ 2 =10 -PL / 10 represents the free space path loss of the secondary user-base station link, PL is expressed as:

4. The method according to claim 1, wherein: The channel model of the secondary user-interferer link is: in represents the free space path loss of the secondary user-interferer link, PL w Expressed as:

5. The method according to claim 1, wherein: In the tth time slot, the achievable communication rate of the secondary user is expressed as: in is the noise power received by the base station.

6. The method according to claim 1, wherein: The specific implementation method of step 2 is: In the CR network of step 1, the interferer has p(1)=1 / N t The probability of selecting the same frequency band as the secondary user is p(0)=(N t -1) / N t The probability of selecting other frequency bands; energy detection is performed on the selected frequency band: Where D represents the detection statistic, y[i] represents the i-th signal sample received by the interferer, M represents the total number of samples of the interferer's sampled signal in a time slot, ε opt represents the optimal energy detection threshold of the interferer; It means that the interferer only receives noise in the selected frequency band and does not transmit interference; It indicates that the interferer receives noise and secondary user signals in the selected frequency band and transmits an interference signal; It is expressed as follows: Where x[i] represents the i-th signal sample from the secondary user received by the interferer, and n[i] represents the variance. Additive Gaussian white noise; Combining formula (6) and formula (7), and Re-expressed as follows: Where D0 and D1 represent and The test statistics in the case all obey Gaussian distribution; P max represents the maximum transmit power of the secondary user, represents the variance of the test statistic; The total probability of detecting an interferer error is expressed as a function of the energy detection threshold ε: p e (ε)=p(1)p MD +p(0)p FA (9) where p MD =p(D1<ε) represents the probability of missed detection, p FA =p(D0≥ε) represents the false alarm probability. According to the properties of Gaussian distribution, equation (9) can be further expressed as: in Denotes the Gaussian Q function; Derivative of Equation (10) with respect to ε, let get: p(1)f1(e opt )=p(0)f0(ε opt ) (11) Where f0(x) and f1(x) represent the probability density functions of Gaussian variables D0 and D1, respectively, and the optimal energy detection threshold ε of the interferer is opt Expressed as: The optimization goal is to maximize the total communication rate of the secondary user while avoiding reactive jamming attacks; the optimization variable is the secondary user's transmit power in each time slot, which is expressed as the set The optimization problem is then expressed as: where R P is a constant penalty term, which is used to represent the impact of interference attacks on the communication rate; Binary auxiliary variable φ t =1 means that the transmission activity of the secondary user in the tth time slot is detected by the jammer and is attacked by reactive jamming. At this time, the jammer selects the same frequency band as the secondary user for monitoring and meets the received signal power φ t =0 means that the secondary user is not subject to reactive interference.

7. The method according to claim 1, wherein: Step 3: Based on the deep reinforcement learning framework, we construct a Markov decision process and define the state, action, and reward as follows: In the Markov decision process, the secondary user acts as an intelligent agent and observes the current state s at each time slot t. t , and select action a according to the strategy t , and obtain the reward r for evaluating the long-term effectiveness of the action t Through continuous interaction with the environment, the agent gradually optimizes its action selection strategy to obtain the maximum long-term reward, that is, the total communication rate of secondary users. The basic elements of MDP are defined as follows: State: The state represents the agent's observation of the environment at a certain moment. In the CR network scenario, secondary users observe the occupancy of each frequency band through spectrum sensing. Therefore, the state of the tth time slot is represented by the number of idle frequency bands, expressed as: s t =N t (14) Action: An action represents the decision made by the agent at a certain moment based on the current state. In the CR network scenario, the action at the tth time slot represents the secondary user's choice of transmit power. It is expressed as: a t =P t ,P t ∈{P k |k=1,2,...,K} (15) Reward: Reward is used to evaluate the effectiveness of the agent's actions at a certain moment and is a direct reflection of the optimization goal. In the CR network scenario, the reward for the tth time slot is expressed as: r t =R t -φ t R P (16) 8. The method according to claim 1, wherein: Step 3: Construct a dual-depth Q network to solve the CR network power allocation solution as follows: Construct a dual-depth Q network, that is, evaluate the neural network Q(s t ,a t ;θ) and the target neural network Q′(s t ,a t ; θ′), where θ and θ′ represent the internal parameters of the two neural networks; the two neural networks have the same structure, and the input is the state s of the tth time slot t , the output is a set of action value vectors, representing the state s t The reward for executing each action; In the training process of the dual-depth Q network, the secondary user acts as an intelligent agent and first collects training data based on a greedy strategy; the intelligent agent selects a random action with probability α and selects an action with probability 1-α. Where α∈[0,1] is the exploration rate; the training data collected by the agent (s t ,a t ,r t ,s t+1 ) are stored in the experience pool in the form of tuples, and the agent randomly samples small batches of data from the experience pool to train and evaluate the neural network; the loss function of the evaluation neural network is defined as: in is the target network Q value for the tth time slot, expressed as: Where γ∈[0,1] is the discount factor; the evaluation neural network continuously backpropagates through Equation (17) and updates the network parameters θ; after each F-step training, the parameters θ of the evaluation neural network are copied to the parameters θ' of the target neural network; After continuous iterative updates of the dual-depth Q network, the optimal evaluation neural network is obtained; in the tth time slot, the optimal allocation scheme for secondary users is expressed as: where θ opt represents the network parameters of the optimal evaluation neural network; the secondary user executes the power allocation scheme according to the spectrum sensing results of each time slot, thereby realizing the secure cognitive radio network power allocation oriented to reactive interference.