Fast self-adaptive anti-interference channel access method based on coarse-grained spectrum prediction

By adopting a fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction in wireless communication systems, the problem of slow convergence speed of anti-interference solutions in the prior art under intelligent interference attacks is solved, and faster convergence speed and superior anti-interference performance are achieved, and there is no need to understand the action space and action information of the jammer.

CN119997248APending Publication Date: 2025-05-13NANJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510162689.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing anti-interference solution converges slowly under intelligent interference attacks, and requires legal users to understand the action space and action information of the jammer, making it difficult to directly apply to actual anti-interference communication.

Method used

Using a fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction, by constructing an intelligent interference decision model and an intelligent anti-interference channel access decision model, and using the Markov game framework and deep Q network, it is possible to quickly adaptive anti-interference channel access for legitimate users without understanding the action space of the jammer.

Benefits of technology

It significantly improves the ability to represent the status feature of legal users in intelligent interference scenarios, learns better Q functions, improves anti-interference performance and convergence speed, and does not require legal users to understand the action space and action information of the jammer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119997248A_ABST
    Figure CN119997248A_ABST
Patent Text Reader

Abstract

The invention discloses a fast self-adaptive anti-interference channel access method based on coarse-grained spectrum prediction, which belongs to the technical field of communication anti-interference, and comprises the following steps: respectively constructing an intelligent interference decision model and an intelligent anti-interference channel access decision model, modeling an interaction process between the legal user and the intelligent jammer as a Markov game; outputting an interference decision of the current hop by using the intelligent interference decision model; online training and updating the intelligent anti-interference channel access decision model; and predicting a coarse-grained spectrum and a Q function value by using the intelligent anti-interference channel access decision model after training and parameter updating so as to output a channel access decision of the legal user at the current hop, thereby realizing real-time access of the anti-interference channel of the legal user. Under the condition that a legal user does not need to know any related jammer action space and jammer action information, the method has higher convergence speed and excellent anti-interference performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of communication anti-interference, and in particular relates to a fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction. Background Art

[0002] Wireless communication technology is an indispensable part of modern society, and it is the main or even the only means of command and control in special environments such as emergency situations. However, due to its shared and open nature, wireless communication technology is extremely vulnerable to various malicious interference attacks. With the continuous advancement of signal processing technology and chip technology, the interference capabilities of malicious interference equipment are increasing. Jammers can easily generate intelligent jamming attacks using intelligent agents based on reinforcement learning.

[0003] Although the anti-interference scheme based on adversary modeling can obtain additional performance gains under intelligent jamming attacks and achieve anti-interference performance beyond Nash equilibrium, the scheme assumes that the legitimate user accurately understands the action space of the jammer and its jamming actions at each hop, which is unreasonable in practice and is therefore difficult to directly apply to actual anti-interference communications. On the other hand, the ε-greedy-based exploration and exploitation method widely used in reinforcement learning-based anti-interference schemes and adversary modeling-based anti-interference schemes, as well as the inefficient method of testing only one action per iteration, lead to slow convergence. If the anti-interference algorithm cannot converge before the jamming strategy changes, its effectiveness will be greatly reduced and the service quality of the wireless communication system will be impaired.

[0004] In view of this, it is necessary to propose a fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction to solve the above problems. Summary of the invention

[0005] In view of the deficiencies in the prior art, the present invention provides a fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction, which can have faster convergence speed and superior anti-interference performance without requiring legitimate users to know any jammer action space and jammer action information.

[0006] The present invention provides the following technical solutions:

[0007] A fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction includes:

[0008] An intelligent jamming decision model and an intelligent anti-jamming channel access decision model are constructed respectively, and the interaction process between legitimate users and intelligent jammers is modeled as a Markov game.

[0009] Based on the spectrum waterfall observed by the intelligent jammer at the current hop, the intelligent jammer decision model is used to output the jamming decision for the current hop;

[0010] Online training and updating of intelligent anti-interference channel access decision model;

[0011] Based on the spectrum waterfall observed by the legitimate user in the current hop, the coarse-grained spectrum and Q function value are predicted using the trained and parameter-updated intelligent anti-interference channel access decision model to output the channel access decision of the legitimate user in the current hop, thereby realizing real-time access of the legitimate user to the anti-interference channel.

[0012] Optionally, the intelligent interference decision model adopts a standard deep reinforcement learning model DQN, and the intelligent interference decision model performs online training and parameter update after outputting an interference decision in each hop; the legitimate user includes a transmitter and a receiver, and the transmitter and the receiver communicate in a frequency hopping manner; during the frequency hopping communication of the legitimate user, the intelligent jammer attacks the corresponding channel according to the output of the intelligent interference decision model; the basic time slots of the legitimate user and the intelligent jammer are aligned.

[0013] Optionally, the intelligent interference decision model and the intelligent anti-interference channel access decision model are constructed respectively, and the interaction process between the legitimate user and the intelligent jammer is modeled as a Markov game. The specific process is:

[0014] Construct an intelligent interference decision model and an intelligent anti-interference channel access decision model respectively;

[0015] The interaction process between legitimate users and intelligent jammers is represented by a tuple Description, where represents the set of environmental states, and denote the action sets of legitimate users and intelligent jammers respectively, is the state transfer function, and denote the reward functions of the legitimate user and the intelligent jammer, respectively, γ∈(0,1] is the discount factor;

[0016] The goal of the intelligent anti-interference channel access decision model is to find the maximum cumulative reward The optimal strategy π * The goal of the intelligent interference decision model is to find the maximum cumulative reward The optimal strategy μ * ;

[0017] The goal of legitimate users is to find u (S1;π b ,μ)≥R j (S1;π,μ), The best response strategy π b .

[0018] Optionally, the intelligent anti-interference channel access decision model includes a feature extraction module, a coarse-grained spectrum prediction module, a Q function estimation module and a joint decision function module;

[0019] The feature extraction module includes a series of convolutional layers for extracting features from the spectrum waterfall currently observed by the legitimate user.

[0020] The coarse-grained spectrum prediction module includes a series of fully connected layers, wherein the first fully connected layer uses the features extracted by the feature extraction module As input, and from the features Extract features from

[0021] The Q function estimation module includes a series of fully connected layers, wherein the first fully connected layer uses the features extracted by the feature extraction module As input, and from the features Extract features from

[0022] The characteristics and Features Connect to construct hidden representation Hide the representation As the input of the second fully connected layer in the coarse-grained spectrum prediction module and the Q function estimation module, and processed separately by the corresponding subsequent fully connected layers, so that the coarse-grained spectrum prediction module outputs the predicted coarse-grained spectrum The Q function estimation module outputs the estimated Q function value Q(S n ,a u ),

[0023] The joint decision function module is used to construct a joint decision function based on the output of the coarse-grained spectrum prediction model and the Q function estimation model. The transmitter of the legitimate user accesses the corresponding channel for transmission according to the output of the joint decision function.

[0024] Optionally, the online training and updating of the intelligent anti-interference channel access decision model comprises the following specific processes:

[0025] Collect coarse-grained spectrum sample-label data from the interaction between legitimate users and intelligent jammers, build a coarse-grained spectrum prediction model, and calculate the coarse-grained spectrum prediction loss;

[0026] Collect state transfer experience data from the interaction between legitimate users and intelligent jammers, build a Q function estimation model, and calculate the Q function value estimation loss;

[0027] The aggregation loss is calculated based on the coarse-grained spectrum prediction loss and the Q-function estimation loss, and the aggregation loss is used to update the parameters of the coarse-grained spectrum prediction model and the parameters of the Q-function estimation model.

[0028] Optionally, the coarse-grained spectrum sample-label data is collected from the interaction process between the legitimate user and the intelligent jammer, a coarse-grained spectrum prediction model is constructed, and the coarse-grained spectrum prediction loss is calculated. The specific process is:

[0029] Construct a coarse-grained spectrum prediction model F(·; ψ), where ψ is the parameter of the coarse-grained spectrum prediction model, the model input is the spectrum waterfall observed by the legitimate user at the beginning of the current hop, and the model output is the predicted coarse-grained spectrum of the current hop;

[0030] At the beginning of the nth hop, the legitimate user calculates the coarse-grained spectrum during the (n-1)th hop Among them, the sample value in the coarse-grained spectrum is the power spectral density function estimated by the P-Welch algorithm using the time domain signal sampled at the (n-1)th hop, Δf ′ =B / M=b is the resolution of coarse-grained spectrum analysis, and B is the total bandwidth;

[0031] Building a dynamically refreshed cache cache The capacity is And follow the first-in-first-out principle when storing data;

[0032] The legitimate user adds the coarse-grained spectrum sample label pair (S n-1 ,c n-1 ) is stored in the cache In cache The latest sample-label pairs; among them, S n-1 is the spectrum waterfall observed before the (n-1)th jump, c n-1 The coarse-grained spectrum during the (n-1)th jump;

[0033] from Randomly select a batch of sample label pairs Calculate the coarse-grained spectrum prediction loss L C (ψ):

[0034]

[0035] in, represents the output of the coarse-grained spectrum prediction model, Represents sample S i The real label.

[0036] Optionally, the state transfer experience data is collected from the interaction process between the legitimate user and the intelligent jammer, a Q function estimation model is constructed, and the Q function value estimation loss is calculated. The specific process is:

[0037] Construct a Q function estimation model Q(·; θ), where θ is the parameter of the Q function estimation model, the model input is the spectrum waterfall observed by the legitimate user at the beginning of the current hop and the available anti-interference actions, and the model output is the estimated Q function value;

[0038] Construct the target Q network. The structure of the target Q network is the same as the Q function estimation model, and its parameter θ - The Q function is used to estimate the model with each N u The frequency of jump update is dynamically updated;

[0039] At the beginning of the nth hop, the legitimate user takes the anti-interference action taken during the (n-1)th hop. Calculate the immediate reward obtained at the (n-1)th hop

[0040]

[0041] in, is the signal-to-interference-noise ratio at the receiver of the legitimate user in the kth time slot, N h The number of basic time slots included in the duration of each hop;

[0042] Build a dynamically refreshed experience replay pool Experience Replay Pool The capacity is And follow the first-in-first-out principle when storing data;

[0043] Legitimate users collect transfer samples during their interactions with the environment And store it in the experience replay pool Experience Replay Pool The latest transfer samples; among them, S n-1 and S n are the spectrum waterfalls observed before the start of the (n-1)th and nth jumps, respectively;

[0044] from Randomly select a batch of sample label pairs Calculate the Q function estimation loss L Q (θ):

[0045]

[0046] Among them, η i It is an experience replay pool The target value of the element in , θ- represents the parameters of the target network, is the immediate reward obtained at the i-th hop.

[0047] Optionally, the aggregation loss is calculated based on the coarse-grained spectrum prediction loss and the Q function estimation loss, and the parameters of the coarse-grained spectrum prediction model and the parameters of the Q function estimation model are updated using the aggregation loss. The specific process is:

[0048] The aggregate loss L(θ,ψ) is calculated based on the coarse-grained spectrum prediction loss and the Q-function estimation loss:

[0049] L(θ,ψ)=L C (ψ)+λL Q (θ)

[0050] Among them, L C (ψ) is the coarse-grained spectral prediction loss, L Q (θ) is the Q function estimation loss, YesL Q Adaptive scaling factor of (θ);

[0051] Update the parameters ψ of the coarse-grained spectrum prediction model:

[0052]

[0053] Among them, α C is the learning rate for updating the parameters of the coarse-grained spectrum prediction model;

[0054] Update the Q function to estimate the model parameters θ:

[0055]

[0056] Among them, α Q is the learning rate at which the Q function estimates the model parameters.

[0057] Optionally, the output of the coarse-grained spectrum prediction model and the Q function estimation model is used to construct a joint decision function, and the transmitter of the legitimate user accesses the corresponding channel for transmission according to the output of the joint decision function. The specific process is as follows:

[0058] Based on the output of the coarse-grained spectrum prediction model and the Q function estimation model, a joint decision function is constructed.

[0059]

[0060] in, Represents the predicted coarse-grained spectrum Middle a u elements; Q(S n ,a u; θ) is the predicted Q function value, where θ is the parameter of the constructed Q function estimation model;

[0061] According to the output of the joint decision function, the anti-interference action of the legitimate user during the nth hop is determined;

[0062]

[0063] During the nth hop, the transmitter acts according to the anti-jamming Select the corresponding channel to access to transfer.

[0064] Compared with the prior art, the present invention has the following beneficial effects:

[0065] The present invention uses coarse-grained spectrum prediction as an auxiliary task to train an anti-interference channel access decision model based on a deep Q network. The present invention can significantly improve the representation ability of legitimate users of state features in intelligent interference scenarios, and learn a better Q function than the traditional deep reinforcement learning method. In addition, in terms of anti-interference performance and convergence speed, the present invention shows more significant superiority than the anti-interference channel access method based on deep reinforcement learning and the anti-interference channel access method based on opponent modeling, and in the training and decision-making process of the anti-interference channel access decision model, the legitimate user does not need to know any information about the jammer action space and the jammer action. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a structural schematic diagram of the intelligent jammer attacking the channel when the legitimate user of the present invention is frequency hopping communication;

[0067] Figure 2 It is a diagram of a fast adaptive anti-interference channel access solution based on coarse-grained spectrum prediction of the present invention;

[0068] Figure 3 It is a schematic diagram of the time slot structure of the fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction of the present invention;

[0069] Figure 4 It is a thermodynamic diagram of the coarse-grained spectrum of the fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction of the present invention;

[0070] Figure 5 It is a schematic diagram of the structure of a joint decision-making network based on a coarse-grained spectrum prediction model and a Q function estimation model of the present invention;

[0071] Figure 6 This is a diagram of the anti-interference performance of Example 2 of the present invention under intelligent interference attack. DETAILED DESCRIPTION

[0072] The present invention is further described below in conjunction with the accompanying drawings. The following examples are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the scope of protection of the present invention. It should be noted that the term "comprising" and any variation thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0073] Example 1

[0074] A fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction is provided, comprising the following steps:

[0075] Step S1: construct an intelligent interference decision model and an intelligent anti-interference channel access decision model respectively, and model the interaction process between the legitimate user and the intelligent jammer as a Markov game.

[0076] like Figure 1 As shown, the legitimate users include transmitters and receivers, which communicate in a frequency hopping manner; the intelligent jammer attacks the corresponding channel according to the output of the intelligent jamming decision model during the frequency hopping communication of the legitimate users; the basic time slots of the legitimate users and the intelligent jammer are aligned. The intelligent jammer is equipped with an intelligent agent with learning ability, which can adaptively adjust its jamming behavior to enhance the jamming performance.

[0077] More specifically, the total bandwidth is B = f U -f L The communication frequency band [f L ,f U ] is divided into M channels with bandwidth b = B / M and no overlap where f m is the center frequency of the mth channel; the legitimate user communicates in a frequency hopping manner, and the basic time slots of the legitimate user and the intelligent jammer are aligned, the duration of each basic time slot is Δt, and the duration of each hop is T h =N h Δt.

[0078] During the nth hop, the legitimate transmitter accesses the channel Transmitting, intelligent jammer attacks N I Continuous Channel in

[0079] In this application, the interaction process between the legitimate user and the intelligent jammer is modeled as a Markov game, and the specific process is:

[0080] Step a: Construct an intelligent interference decision model and an intelligent anti-interference channel access decision model respectively.

[0081] The intelligent interference decision model of the present application adopts the standard deep reinforcement learning model DQN. For details, reference may be made to the prior art. The intelligent interference decision model performs online training and parameter updates after outputting the interference decision at each hop. The online training of the intelligent interference decision model and the calculation of the reward for attack during each hop of communication of the legitimate user may all refer to the prior art. At the nth hop, the intelligent interference decision model performs N consecutive T Spectrum waterfall composed of spectrum vectors observed by the intra-hop intelligent jammer As input, output the interference decision of the current hop

[0082] Step b: The interaction process between the legitimate user and the intelligent jammer is represented by a tuple Description, where represents the set of environmental states, and denote the action sets of legitimate users and intelligent jammers respectively, is the state transfer function, and denote the reward functions of the legitimate user and the intelligent jammer respectively, and γ∈(0,1] is the discount factor.

[0083] Step c: The goal of the intelligent anti-interference channel access decision model is to find the maximum cumulative reward The optimal strategy π * The goal of the intelligent interference decision model is to find the maximum cumulative reward The optimal strategy μ * ; Among them, γ n is the discount factor for the nth hop, and are the rewards of the n-th hop intelligent anti-interference channel access decision model and the intelligent interference decision model, respectively, and S1 is the initial state.

[0084] Step d: The goal of the legitimate user is to find u (S1;π b ,μ)≥R j (S1;π,μ), The best response strategy π b .

[0085] It is worth noting that in this application, at the nth hop, the intelligent anti-interference channel access decision model is based on N consecutive T Spectral waterfall consisting of the spectrum vectors observed within a jump As input, output the channel access decision of the current hop

[0086] Specifically, at the beginning of the nth hop, the legitimate user calculates the spectrum vector observed within (n-1) hops: The sample values ​​in the spectrum vector is the power spectral density function estimated by the P-Welch algorithm using the time domain signal sampled at the kth time slot, Δf = B / N F is the resolution of spectrum analysis; B is the total bandwidth, [f L ,f U ] is the communication frequency, N h is the number of basic time slots contained in each hop duration, N F is the number of samples in the spectrum vector.

[0087] like Figure 5 As shown, the intelligent anti-interference channel access decision model includes a feature extraction module, a coarse-grained spectrum prediction module, a Q function estimation module and a joint decision function module.

[0088] Specifically:

[0089] 1. Feature Extraction Module

[0090] It consists of a series of convolutional layers to extract features from the spectrum waterfall currently observed by the legitimate user.

[0091] 2. Coarse-grained spectrum prediction module

[0092] It consists of a series of fully connected layers, where the first fully connected layer uses the features extracted by the feature extraction module As input, and from the features Extract features from

[0093] 3. Q function estimation module

[0094] It consists of a series of fully connected layers, where the first fully connected layer uses the features extracted by the feature extraction module As input, and from the features Extract features from

[0095] The characteristics and Features Connect to construct hidden representation Hide the representation As the input of the second fully connected layer in the coarse-grained spectrum prediction module and the Q function estimation module, and processed separately by the corresponding subsequent fully connected layers, so that the coarse-grained spectrum prediction module outputs the predicted coarse-grained spectrum The Q function estimation module outputs the estimated Q function value Q(Sn ,a u ),

[0096] 4. Joint Decision-Making Function Module

[0097] A joint decision function is constructed based on the output of the coarse-grained spectrum prediction model and the Q function estimation model. The transmitter of the legitimate user accesses the corresponding channel for transmission according to the output of the joint decision function.

[0098] More specifically, the joint decision function module includes the following sub-steps: constructing a joint decision function based on the outputs of the coarse-grained spectrum prediction model and the Q function estimation model

[0099]

[0100] in, Represents the predicted coarse-grained spectrum Middle a u elements; Q(S n ,a u ; θ) is the predicted Q function value, where θ is the parameter of the constructed Q function estimation model;

[0101] According to the output of the joint decision function, the anti-interference action of the legitimate user during the nth hop is determined;

[0102]

[0103] During the nth hop, the transmitter acts according to the anti-jamming Select the corresponding channel to access to transfer.

[0104] Step S2: Based on the spectrum waterfall observed by the intelligent jammer at the current hop, the intelligent interference decision model is used to output the interference decision of the current hop.

[0105] Step S3: Online training and updating of the intelligent anti-interference channel access decision model.

[0106] Specifically, step S3 includes the following sub-steps:

[0107] Step S31: Collect coarse-grained spectrum sample-label data from the interaction process between the legitimate user and the intelligent jammer, build a coarse-grained spectrum prediction model, and calculate the coarse-grained spectrum prediction loss.

[0108] like Figure 3 and Figure 4 As shown, step S31 specifically includes the following sub-steps:

[0109] Step S311: Based on the feature extraction module and the coarse-grained spectrum prediction module of the intelligent anti-interference channel access decision model, a coarse-grained spectrum prediction model F(·; ψ) is constructed, where ψ is a parameter of the coarse-grained spectrum prediction model, the model input is the spectrum waterfall observed by the legitimate user at the beginning of the current hop, and the model output is the predicted coarse-grained spectrum of the current hop;

[0110] Step S312: At the beginning of the nth hop, the legitimate user calculates the coarse-grained spectrum during the (n-1)th hop Among them, the sample value in the coarse-grained spectrum 1,2,…,M,S n-1 (f+f L ) is the power spectral density function estimated by the P-Welch algorithm using the time domain signal sampled at the (n-1)th hop, Δf ′ =B / M=b is the resolution of coarse-grained spectrum analysis, and B is the total bandwidth.

[0111] Step S313: Build a dynamically refreshed cache cache The capacity is And the first-in-first-out principle is followed when storing data.

[0112] Step S314: The legal user adds the coarse-grained spectrum sample label pair (S n-1 ,c n-1 ) is stored in the cache In cache The latest sample-label pairs; among them, S n-1 is the spectrum waterfall observed before the (n-1)th jump, c n-1 is the coarse-grained spectrum during the (n-1)th jump.

[0113] Step S315: From Randomly select a batch of sample label pairs Calculate the coarse-grained spectrum prediction loss L C (ψ):

[0114]

[0115] in, represents the output of the coarse-grained spectrum prediction model, Represents sample S i The real label.

[0116] Step S32: collecting state transfer experience data from the interaction process between the legitimate user and the intelligent jammer, constructing a Q function estimation model, and calculating the Q function value estimation loss.

[0117] Step S32 specifically includes the following sub-steps:

[0118] Step S321: Based on the feature extraction module and the coarse-grained spectrum prediction module of the intelligent anti-interference channel access decision model, a Q function estimation model Q(·; θ) is constructed, where θ is a parameter of the Q function estimation model, the model input is the spectrum waterfall observed by the legitimate user at the beginning of the current hop and the available anti-interference actions, and the model output is the estimated Q function value;

[0119] Step S322: construct a target Q network. The structure of the target Q network is the same as the Q function estimation model. Its parameters θ - The Q function is used to estimate the model for each N u The frequency of jump update is dynamically updated;

[0120] Step S323: At the beginning of the nth hop, the legitimate user takes the anti-interference action taken during the (n-1)th hop. Calculate the immediate reward obtained at the (n-1)th hop

[0121]

[0122] in, is the signal-to-interference-noise ratio at the receiver of the legitimate user in the kth time slot, N h The number of basic time slots included in the duration of each hop;

[0123]

[0124] Step S324: Construct a dynamically refreshed experience replay pool Experience Replay Pool The capacity is And follow the first-in-first-out principle when storing data;

[0125] Step S325: Legitimate users collect transfer samples during the interaction with the environment And store it in the experience replay pool Experience Replay Pool The latest transfer samples; among them, S n-1 and S n are the spectrum waterfalls observed before the start of the (n-1)th and nth jumps, respectively;

[0126] Step S326: From Randomly select a batch of sample label pairs Calculate the Q function estimation loss L Q (θ):

[0127]

[0128] Among them, η i It is an experience replay pool The target value of the element in , θ - Represents the parameters of the target network.

[0129] Step S33: Calculate the aggregation loss based on the coarse-grained spectrum prediction loss and the Q-function estimation loss, and use the aggregation loss to update the parameters of the coarse-grained spectrum prediction model and the parameters of the Q-function estimation model.

[0130] Specifically, the aggregation loss L(θ,ψ) is:

[0131] L(θ,ψ)=L C (ψ)+λL Q (θ)

[0132] Among them, L C (ψ) is the coarse-grained spectral prediction loss, L Q (θ) is the Q function estimation loss, YesL Q Adaptive scaling factor of (θ);

[0133] Update the parameters ψ of the coarse-grained spectrum prediction model:

[0134]

[0135] Among them, α C is the learning rate for updating the parameters of the coarse-grained spectrum prediction model;

[0136] Update the Q function to estimate the model parameters θ:

[0137]

[0138] Among them, α Q is the learning rate at which the Q function estimates the model parameters.

[0139] Step S4: Figure 2 As shown, based on the spectrum waterfall observed by the legitimate user in the current hop, the intelligent anti-interference channel access decision model after training and parameter update is used to predict the coarse-grained spectrum and Q function value to output the channel access decision of the legitimate user in the current hop, thereby realizing real-time access of the legitimate user to the anti-interference channel.

[0140] A Markov game framework is adopted to model the interaction between legitimate users and intelligent jammers, and the anti-interference channel access decision model based on deep Q-network is trained by taking the synchronously updated coarse-grained spectrum prediction as an auxiliary task to find the optimal response strategy of legitimate users under time-varying intelligent jamming attacks, thereby achieving real-time access of legitimate users to anti-interference channels.

[0141] Example 2

[0142] A specific example of a fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction is given. The system simulation adopts Python's Pytorch framework. The system model includes a group of legitimate users and an intelligent jammer.

[0143] In this implementation, both the legitimate user and the intelligent jammer work in the frequency band from 0 MHz to 20 MHz. This frequency band can be divided into M = 10 channels with a bandwidth of b = 20 MHz and no overlap. The duration of each hop is T h = 10ms, the intelligent jammer can attack N I = 3 consecutive channels, the legitimate users and intelligent jammers can perform full-band spectrum sensing with a frequency resolution of Δf = 100kHz every Δt = 1ms, and retain the most recent N T = The spectrum data within 20 hops is used to construct a spectrum waterfall. Based on this, the anti-interference performance diagram of the legitimate user in this example is given, as shown in Figure 6 As shown, the anti-interference performance of the fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction proposed by the present invention is much better than that of the traditional frequency hopping anti-interference channel access method and the random frequency hopping anti-interference channel access method, and the anti-interference performance of the anti-interference channel access method proposed by the present invention is improved by about 10% compared with the anti-interference channel access method based on deep reinforcement learning and the anti-interference channel access method based on Nash equilibrium. Compared with the anti-interference channel access method based on opponent modeling, the fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction proposed by the present invention has a faster convergence speed, and the anti-interference performance is slightly improved while reducing the number of training rounds by up to 70%, and does not require any information related to the interference action and the jammer action space.

[0144] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solutions in the embodiments of the present invention are essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention or certain parts of the embodiments.

[0145] The above are only preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should be regarded as the protection scope of the present invention.

Claims

1. A fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction, characterized in that: include: An intelligent jamming decision model and an intelligent anti-jamming channel access decision model are constructed respectively, and the interaction process between legitimate users and intelligent jammers is modeled as a Markov game. Based on the spectrum waterfall observed by the intelligent jammer at the current hop, the intelligent jammer decision model is used to output the jamming decision for the current hop; Online training and updating of intelligent anti-interference channel access decision model; Based on the spectrum waterfall observed by the legitimate user in the current hop, the coarse-grained spectrum and Q function value are predicted using the trained and parameter-updated intelligent anti-interference channel access decision model to output the channel access decision of the legitimate user in the current hop, thereby realizing real-time access of the legitimate user to the anti-interference channel.

2. The fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction according to claim 1 is characterized in that: The intelligent interference decision model adopts the standard deep reinforcement learning model DQN; the intelligent interference decision model performs online training and parameter update after outputting the interference decision in each hop; the legitimate user includes a transmitter and a receiver, and the transmitter and the receiver communicate in a frequency hopping manner; during the frequency hopping communication of the legitimate user, the intelligent jammer attacks the corresponding channel according to the output of the intelligent interference decision model; the basic time slots of the legitimate user and the intelligent jammer are aligned.

3. The fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction according to claim 1 is characterized in that: The intelligent interference decision model and the intelligent anti-interference channel access decision model are constructed respectively, and the interaction process between the legitimate user and the intelligent jammer is modeled as a Markov game. The specific process is: Construct an intelligent interference decision model and an intelligent anti-interference channel access decision model respectively; The interaction process between legitimate users and intelligent jammers is represented by a tuple Description, where represents the set of environmental states, and denote the action sets of legitimate users and intelligent jammers respectively, is the state transfer function, and denote the reward functions of the legitimate user and the intelligent jammer, respectively, γ∈(0,1] is the discount factor; The goal of the intelligent anti-interference channel access decision model is to find the maximum cumulative reward The optimal strategy π * The goal of the intelligent interference decision model is to find the maximum cumulative reward The optimal strategy μ * ; The goal of legitimate users is to find The best response strategy π b .

4. The fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction according to claim 1 is characterized in that: The intelligent anti-interference channel access decision model includes a feature extraction module, a coarse-grained spectrum prediction module, a Q function estimation module and a joint decision function module; The feature extraction module includes a series of convolutional layers for extracting features from the spectrum waterfall currently observed by the legitimate user. The coarse-grained spectrum prediction module includes a series of fully connected layers, wherein the first fully connected layer uses the features extracted by the feature extraction module As input, and from the features Extract features from The Q function estimation module includes a series of fully connected layers, wherein the first fully connected layer uses the features extracted by the feature extraction module As input, and from the features Extract features from The characteristics and Features Connect to construct hidden representation Hide the representation As the input of the second fully connected layer in the coarse-grained spectrum prediction module and the Q function estimation module, and processed separately by the corresponding subsequent fully connected layers, so that the coarse-grained spectrum prediction module outputs the predicted coarse-grained spectrum The Q function estimation module outputs the estimated Q function value The joint decision function module is used to construct a joint decision function based on the output of the coarse-grained spectrum prediction model and the Q function estimation model. The transmitter of the legitimate user accesses the corresponding channel for transmission according to the output of the joint decision function.

5. The fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction according to claim 1 is characterized in that: The specific process of online training and updating the intelligent anti-interference channel access decision model is as follows: Collect coarse-grained spectrum sample-label data from the interaction between legitimate users and intelligent jammers, build a coarse-grained spectrum prediction model, and calculate the coarse-grained spectrum prediction loss; Collect state transfer experience data from the interaction between legitimate users and intelligent jammers, build a Q function estimation model, and calculate the Q function value estimation loss; The aggregation loss is calculated based on the coarse-grained spectrum prediction loss and the Q-function estimation loss, and the aggregation loss is used to update the parameters of the coarse-grained spectrum prediction model and the parameters of the Q-function estimation model.

6. The fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction according to claim 5 is characterized in that: The coarse-grained spectrum sample-label data is collected from the interaction process between the legitimate user and the intelligent jammer, a coarse-grained spectrum prediction model is constructed, and the coarse-grained spectrum prediction loss is calculated. The specific process is: Construct a coarse-grained spectrum prediction model F(·; ψ), where ψ is the parameter of the coarse-grained spectrum prediction model, the model input is the spectrum waterfall observed by the legitimate user at the beginning of the current hop, and the model output is the predicted coarse-grained spectrum of the current hop; At the beginning of the nth hop, the legitimate user calculates the coarse-grained spectrum during the (n-1)th hop Among them, the sample value in the coarse-grained spectrum S n-1 (f+f L ) is the power spectral density function estimated by the P-Welch algorithm using the time domain signal sampled at the (n-1)th hop, Δf′=B / M=B is the resolution of the coarse-grained spectrum analysis, and B is the total bandwidth; Building a dynamically refreshed cache cache The capacity is And follow the first-in-first-out principle when storing data; The legitimate user adds the coarse-grained spectrum sample label pair (S n-1 ,c n-1 ) is stored in the cache In cache The latest sample-label pairs; among them, S n-1 is the spectrum waterfall observed before the (n-1)th jump, c n-1 is the coarse-grained spectrum during the (n-1)th jump; from Randomly select a batch of sample label pairs Calculate the coarse-grained spectrum prediction loss L C (ψ): in, represents the output of the coarse-grained spectrum prediction model, Represents sample S i The real label.

7. The fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction according to claim 5 is characterized in that: The state transfer experience data is collected from the interaction process between the legitimate user and the intelligent jammer, a Q function estimation model is constructed, and the Q function value estimation loss is calculated. The specific process is as follows: Construct a Q function estimation model Q(·; θ), where θ is the parameter of the Q function estimation model, the model input is the spectrum waterfall observed by the legitimate user at the beginning of the current hop and the available anti-interference actions, and the model output is the estimated Q function value; Construct the target Q network. The structure of the target Q network is the same as the Q function estimation model, and its parameter θ - The Q function is used to estimate the model with each N u The frequency of jump update is dynamically updated; At the beginning of the nth hop, the legitimate user takes the anti-interference action taken during the (n-1)th hop. Calculate the immediate reward obtained at the (n-1)th hop in, is the signal-to-interference-noise ratio at the receiver of the legitimate user in the kth time slot, N h The number of basic time slots included in the duration of each hop; Build a dynamically refreshed experience replay pool Experience Replay Pool The capacity is And follow the first-in-first-out principle when storing data; Legitimate users collect transfer samples during their interactions with the environment And store it in the experience replay pool Experience Replay Pool The latest transfer samples; among them, S n-1 and S n are the spectrum waterfalls observed before the start of the (n-1)th and nth jumps, respectively; from Randomly select a batch of sample label pairs Calculate the Q function estimation loss L Q (θ): Among them, η i It is an experience replay pool The target value of the element in , θ - represents the parameters of the target network, is the immediate reward obtained at the i-th hop.

8. The fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction according to claim 5 is characterized in that: The aggregation loss is calculated based on the coarse-grained spectrum prediction loss and the Q function estimation loss, and the parameters of the coarse-grained spectrum prediction model and the parameters of the Q function estimation model are updated using the aggregation loss. The specific process is: The aggregate loss L(θ,ψ) is calculated based on the coarse-grained spectrum prediction loss and the Q-function estimation loss: L(θ,ψ)=L C (ψ)+λL Q (i) Among them, L C (ψ) is the coarse-grained spectral prediction loss, L Q (θ) is the Q function estimation loss, YesL Q Adaptive scaling factor of (θ); Update the parameters ψ of the coarse-grained spectrum prediction model: Among them, α C is the learning rate for updating the parameters of the coarse-grained spectrum prediction model; Update the Q function to estimate the model parameters θ: Among them, α Q is the learning rate at which the Q function estimates the model parameters.

9. The fast adaptive anti-interference channel access method based on coarse-grained spectrum prediction according to claim 4 is characterized in that: The joint decision function is constructed based on the output of the coarse-grained spectrum prediction model and the Q function estimation model. The transmitter of the legitimate user accesses the corresponding channel for transmission according to the output of the joint decision function. The specific process is as follows: Based on the output of the coarse-grained spectrum prediction model and the Q function estimation model, a joint decision function is constructed. in, Represents the predicted coarse-grained spectrum Middle a u elements; Q(S n ,a u ; θ) is the predicted Q function value, where θ is the parameter of the constructed Q function estimation model; According to the output of the joint decision function, the anti-interference action of the legitimate user during the nth hop is determined; During the nth hop, the transmitter acts according to the anti-jamming Select the corresponding channel to access to transfer.