A collaborative anti-jamming communication method and system
By constructing a collaborative anti-interference communication method and a deep reinforcement learning algorithm, the information interaction problem between communication users is solved, and the collaborative anti-interference of multiple users in complex spectrum environments is realized, which improves the anti-interference capability of the network and the frequency resource utilization efficiency.
Patent Information
- Application Number
- CN202410636103.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-05-22
AI Technical Summary
In complex and changeable communication and confrontation scenarios with shortage of resources, frequent and reliable information interactions between communication users are difficult to achieve, resulting in the waste of spectrum resources and the problem of interference and damage to interference.
By constructing a collaborative anti-interference communication method, spectrum perception results and information transmission results are obtained, spectrum perception results and information transmission results are trained using the communication frequency decision model, and communication time slot structure is designed so that decision results and ACK information between the transmitter and the receiver are transmitted through the service channel. Deep reinforcement learning algorithms are used to independently learn anti-interference and frequency multiplexing strategies.
It realizes independent learning of multiple users without information interaction in complex and harsh spectrum environments, solves the problem of waste of frequency resources and easy interference damage to interference, and improves the anti-interference ability of the communication network.
Smart Images

Figure CN118573306B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of anti-jamming in wireless communication, and particularly to a collaborative anti-jamming communication method and system. Background Art
[0002] With the continuous expansion of the user scale of wireless communication networks, there are often multiple frequency-using devices in the network. While communication users are combating external malicious interference, they also need to solve the problem of internal frequency coordination. Therefore, it is more practically significant to study multi-user collaborative anti-jamming methods.
[0003] However, current research on anti-jamming problems in multi-agent systems mostly focuses on communication scenarios with small-scale dense deployments. Users in the network only avoid co-channel interference by using orthogonal frequencies. In the face of wide-area and large-scale communication scenarios, there are problems such as spectrum resource waste and the vulnerability of orthogonality to interference damage. Moreover, in existing research, multi-agent systems mostly use centralized or distributed learning algorithms that rely on reliable control channels for information interaction between users. Although good learning effects can be obtained, in complex and variable communication confrontation scenarios with resource shortages, frequent and reliable information interaction between communication users is difficult to achieve. Summary of the Invention
[0004] The present invention provides a collaborative anti-jamming communication method and system to solve the technical defect that it is difficult to achieve frequent and reliable information interaction between communication users in complex and variable communication confrontation scenarios with resource shortages.
[0005] In a first aspect, the present invention provides a collaborative anti-jamming communication method, which includes:
[0006] Obtaining spectrum sensing results, where the spectrum sensing results include the state information of service channels obtained under different observation states; the observation states include combating external malicious interference and combating multi-user communication interference;
[0007] Obtaining the information transmission results of the service channels, where the information transmission results include successful information transmission and unsuccessful information transmission;
[0008] Based on a communication frequency decision model and the spectrum sensing results, obtaining decision results corresponding to different observation states output by the communication frequency decision model; the communication frequency decision model is obtained by training a to-be-trained decision model based on a to-be-trained sample set and information transmission results; the decision results include the next access service channel information under different observation states;
[0009] Based on the decision results corresponding to different observation states, determining a communication frequency strategy after coordinating different observation states.
[0010] According to the collaborative anti-interference communication method described above, the step of obtaining the decision results corresponding to different observation states output by the communication frequency decision model includes:
[0011] Taking the spectrum sensing result as the input of the communication frequency decision model, and obtaining the spectrum feature information of the spectrum sensing result output by the convolutional layer in the communication frequency decision model.
[0012] According to the collaborative anti-interference communication method described above, the steps of training the decision model to be trained based on the training sample set and the information transmission result to obtain the communication frequency decision model include:
[0013] Obtaining a training sample set; the training sample set includes a plurality of sample spectrum sensing data;
[0014] For any sample spectrum sensing data in the plurality of sample spectrum sensing data, taking the sample spectrum sensing data as the input of the decision model to be trained, and obtaining the decision results corresponding to different observation states output by the decision model to be trained;
[0015] Based on the decision result and the information transmission result, adjusting the parameters of the decision model to be trained, and determining the communication decision model.
[0016] According to the collaborative anti-interference communication method described above, the step of adjusting the parameters of the decision model to be trained based on the decision result and the information transmission result includes:
[0017] Updating the reward function based on the information transmission result.
[0018] According to the collaborative anti-interference communication method described above, the step of adjusting the parameters of the decision model to be trained based on the decision result and the information transmission result further includes:
[0019] Adjusting the parameters of the decision model to be trained by adopting an experience replay mechanism.
[0020] According to the collaborative anti-interference communication method described above, the step of determining the communication frequency strategy after coordinating different observation states based on the decision results corresponding to different observation states includes:
[0021] Fitting based on the decision results corresponding to different observation states to obtain a composite action value function;
[0022] Based on the composite action value function, determining the weighted average value of the decision results corresponding to different observation states,
[0023] A communication strategy after coordinating different observation states is determined according to the weighted average value, and the communication strategy includes the next access service channel information after coordinating different observation states.
[0024] According to the collaborative anti-jamming communication method described above, the spectrum sensing result includes spectrum waterfall diagrams in different observation states.
[0025] In a second aspect, the present invention provides a collaborative anti-jamming communication system, which includes a receiver and a transmitter. The transmitter executes the steps of the collaborative anti-jamming communication method described in any one of the above;
[0026] The receiver is used to determine whether the transmission information of multiple service channels is successfully transmitted. If the transmission information is successfully transmitted, feedback information is sent to the transmitter. If the transmission information is not successfully transmitted, no feedback information is sent to the transmitter;
[0027] The transmitter generates information transmission results of multiple service channels according to the feedback information sent by the receiver and determines a communication frequency usage strategy after coordinating different observation states based on the obtained spectrum sensing result.
[0028] According to the collaborative anti-jamming communication system described above, the receiver and the transmitter communicate in units of time slots, and each time slot includes three stages: data transmission, acknowledgement character (ACK) transmission, and learning decision.
[0029] In a fourth aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. The computer program, when executed by a processor, implements the steps of the collaborative anti-jamming communication method described in any one of the above.
[0030] Compared with the prior art, the collaborative anti-jamming communication method and system provided by the present invention have the following advantages:
[0031] (1) The wide-area multi-user communication frequency usage decision model proposed by the present invention is complete in construction and fully considers the frequency reuse characteristics brought by channel fading and user location distribution, solving the problems of frequency resource waste and easy interference damage to orthogonality that exist in traditional multi-user anti-jamming problems where users only rely on orthogonal frequencies to avoid co-frequency interference;
[0032] (2) Innovatively modeling the communication user transmitter as an intelligent agent with learning and decision-making capabilities, designing a communication time slot structure, enabling the decision-making results and control information such as ACK between the transmitter and the receiver to be transmitted through the service channel together with the data information, solving the problem of unreliable control channels in complex and harsh spectrum environments. At the same time, the communication system proposed by the present invention can achieve multi-user independent learning to reach a collaborative effect without information interaction and a reliable control channel.
[0033] (3) Design of innovative deep reinforcement learning algorithms. For the collaborative anti-interference method based on intelligent frequency reuse, a network model and a composite reward function are designed, enabling communication users to independently learn anti-interference and frequency reuse strategies under different observation states and effectively solve the proposed model. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] To more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0035] Figure 1 It is a schematic diagram of the application scenario of the collaborative anti-interference communication system provided by the present invention;
[0036] Figure 2 It is a schematic diagram of the structure of the user communication time slot provided by the present invention;
[0037] Figure 3 It is a schematic flowchart of a collaborative anti-interference communication method provided by the present invention;
[0038] Figure 4 It is a schematic diagram of the structure of a frequency decision-making model for communication provided by the present invention;
[0039] Figure 5 It is a comparison chart of the network normalized communication throughput varying with the simulation time under the condition that the jammer releases swept-frequency interference provided by the present invention;
[0040] Figure 6 It is a comparison chart of the network normalized communication throughput varying with the simulation time under the condition that the jammer releases comb-shaped interference provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] To make the objectives, technical solutions, and advantages of the present invention clearer, the following clearly and completely describes the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention fall within the scope of protection of the present invention.
[0042] It should be noted that in the description of the embodiments of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0043] The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are usually of the same category, and do not limit the number of objects. For example, the first object can be one or multiple.
[0044] The following is combined with Figures 1 - 6 to describe the collaborative anti-jamming communication method and system provided by the embodiments of the present invention.
[0045] For the convenience of understanding, this article first introduces the system scenario applicable to the collaborative anti-jamming communication method provided by this application. Figure 1 is a schematic diagram of the application scenario of the collaborative anti-jamming communication system provided by the present invention. The system includes: a receiver and a transmitter. In the Figure 1 shown application scenario, multiple users in the communication network are scattered. Considering the channel fading characteristics brought by path loss, there is a certain interference and multiplexing relationship between users. A communication user includes a receiver and a transmitter that executes the collaborative anti-jamming communication method with learning and decision-making capabilities. Multiple malicious jammers attack the physical transmission link in the network by sending radio interference, and the degree of interference threat depends on the interference mode, interference signal power, and the position of the jammers, etc.
[0046] Figure 2 is a schematic diagram of the structure of the user communication time slot provided by the embodiments of the present invention. Users in the network communicate in units of time slots. Each communication time slot includes three stages: data transmission, acknowledgement character (ACK) transmission, and learning and decision-making.
[0047] In one embodiment, at the beginning of the communication process, the transmitter randomly selects a transmission channel to transmit data information to the receiver. The receiver adopts broadband multi-channel receiving technology to scan the entire frequency band of the communication network, simultaneously receive and process information on multiple channels, and determine whether the information transmission with the destination address being the local end is successful according to the signal-to-interference-plus-noise ratio (SINR) (data transmission stage); if the information is successfully received, the receiver feeds back ACK information to the transmitter through this service channel. If the information is interfered during the transmission process or the received SINR is lower than the demodulation threshold due to frequency reuse, no ACK information is fed back (ACK transmission stage).
[0048] Based on the results of spectrum sensing and the feedback information of the receiver, the transmitter conducts algorithm learning and updating through a communication frequency decision model, decides the service channel to access in the next communication time slot, changes the working parameters of relevant transmission devices (learning and decision-making stage), and sends the decision result together with the data information through this service channel to enter the data transmission stage of the next communication time slot.
[0049] Combining the above commonalities, this embodiment shows a schematic flowchart of a cooperative anti-jamming communication method. The following takes an application such as Figure 3 as shown, including but not limited to the following steps:
[0050] Step 101: Obtain the results of spectrum sensing.
[0051] Among them, the results of spectrum sensing include the status information of service channels obtained under different observation states, including the on-state of service channels. Different observation states include the observation state of resisting external malicious interference and the observation state of resisting multi-user communication interference.
[0052] Preferably, the results of spectrum sensing include spectrum waterfall diagrams under different observation states.
[0053] Step 102: Obtain the information transmission results of the service channel.
[0054] Among them, the information transmission results include successful information transmission on the service channel and unsuccessful information transmission on the service channel.
[0055] Step 103: Based on the communication frequency decision model and the results of spectrum sensing, obtain the corresponding decision results under different observation states output by the communication frequency decision model.
[0056] Among them, the communication frequency decision model is obtained by training the to-be-trained decision model based on the to-be-trained sample set and the information transmission results. The decision results include the information of the next access service channel under different observation states, and the information of the next access service channel is the information of the service channel that the communication user accesses in the next communication time slot.
[0057] Based on the content of the above embodiments, as an alternative embodiment, in the collaborative anti-interference communication method provided by the present invention, the step of obtaining the decision results corresponding to different observation states output by the communication frequency decision model includes:
[0058] Taking the spectrum sensing result as the input of the communication frequency decision model, and obtaining the spectrum feature information of the spectrum sensing result output by the convolutional layer in the communication frequency decision model.
[0059] Based on the content of the above embodiments, as an alternative embodiment, the steps for training the communication frequency decision model in the collaborative anti-interference communication method provided by the present invention based on a training sample set and information transmission results include:
[0060] Obtaining a training sample set; the training sample set includes a plurality of sample spectrum sensing data.
[0061] For any one of the plurality of sample spectrum sensing data, taking the sample spectrum sensing data as the input of the training decision model, and obtaining the decision results corresponding to different observation states output by the training decision model;
[0062] Based on the decision results and information transmission results, adjusting the parameters of the training decision model, and determining a communication decision model.
[0063] In the embodiments of the present invention, the training sample set may include a plurality of sample spectrum sensing data, and the sample spectrum sensing data includes spectrum sensing results in different observation states, such as a spectrum waterfall diagram in an external malicious interference resistance observation state or a spectrum waterfall diagram in a multi-user communication interference resistance observation state.
[0064] Step 104: Based on the decision results corresponding to different observation states, determine a communication frequency strategy after coordinating different observation states.
[0065] Among them, the communication frequency strategy is the next access service channel information after users coordinate different observation states.
[0066] Based on the content of the above embodiments, as an alternative embodiment, in the collaborative anti-interference communication method provided by the present invention, the step of determining a communication frequency strategy after coordinating different observation states based on the decision results corresponding to different observation states includes:
[0067] Fitting based on the decision results corresponding to different observation states to obtain a composite action value function;
[0068] Based on the composite action value function, determining the weighted average value of the decision results corresponding to different observation states,
[0069] Determine the communication strategy after coordinating different observation states using a pure greedy strategy according to the weighted average value. The communication strategy includes the next access service channel information after coordinating different observation states.
[0070] The technical solution of the present invention will be described again below in conjunction with specific embodiments.
[0071] In one embodiment, system simulation uses the Python language and constructs a frequency decision model for communication based on the deep learning framework of TensorFlow. The parameter settings do not affect generality.
[0072] This embodiment verifies the effectiveness of the proposed model and algorithm. The parameters can be set as follows: the number of communication users N in the network = 20, the number of available channels K = 10, and the environmental noise σ = -80 dBm. The transmission power p of the communication user transmitter n = 50 mW, the length T of each communication time slot c = 10 ms, and the communication service quality threshold β th = 56 dBm. The historical duration Φ of the communication user's perception of the spectrum environment information is 100 ms, and the size of the spectrum waterfall diagram is 10×100.
[0073] Set the deep Q-network reuse reward function adopted by the frequency decision model for communication as:
[0074]
[0075] where β t is the communication service quality of the service channel, the learning rate α = 0.5, the discount factor γ = 0.8, and the replay buffer capacity Batch = 128.
[0076] There is a jammer in the communication network that can release swept interference and comb interference. In the swept interference mode, the jammer continuously interferes with any channel for a time T j = 5 ms, that is, it interferes with two channels in each communication time slot; in the comb interference mode, the jammer continuously interferes with any two channels in the network. Set the simulation time T sim = 10000 ms.
[0077] The communication network contains multiple jammers, and the jammer set is The jammer releases complex dynamic interference to attack the physical transmission link of the communication network.
[0078] In this embodiment, the entire communication network shares a section of spectrum, and the channel set is K < N. For each channel k (k ∈ {1, 2,..., K}), its frequency range is [f k -b / 2, f k +b / 2], where f kis the center frequency of channel k, and b is the channel bandwidth.
[0079] Considering the channel fading characteristics in the communication network, the channel gain coefficient for communication users n and m to transmit information on channel k is defined, which consists of path loss, shadow fading, and small-scale fading. Among them, d n,m is the distance between users n and m, and ω is the path loss factor.
[0080] During the communication user data transmission phase, since the network contains both external malicious interference and co-channel interference between users, the SINR when user n communicates on channel k is expressed as:
[0081]
[0082] where σ is the power of additive white Gaussian noise;
[0083] is the transmission power of the user transmitter, defined as:
[0084]
[0085] where U(f) is the power spectral density (PSD) equation of the transmitter.
[0086] is the malicious interference power received by user n when communicating on channel k, defined as:
[0087]
[0088] is the total co-channel interference power of the transmitters of other communication users in the network received by user n when communicating on channel k, defined as:
[0089]
[0090] where, is the set of other users in the network except communication user n.
[0091] The SINR in the communication user ACK transmission phase is defined similarly to that in the data transmission phase, so it is omitted for the sake of simplicity. Considering the transmission rates of data information and ACK information under the communication service quality constraints, the normalized communication throughput of user n on channel k can be defined as:
[0092]
[0093] where, are the SINRs of the data information and ACK information received by user n respectively.
[0094] Each transmitter of a communication user includes a sensing module that can sense the spectrum information of the entire communication system bandwidth in real time to obtain the spectrum waterfall diagram under different observation states. Considering the existence of signals of other users and external malicious interference signals in the communication network, at communication time slot t, the sensing result of user n's transmitter on channel k can be expressed as:
[0095]
[0096] Then the sensing result of user n on the entire spectrum at communication time slot t is
[0097] While avoiding external malicious interference, communication users need to coordinate the frequency usage strategy within the communication system through frequency reuse within the limited spectrum bandwidth.
[0098] For communication user n, to determine the frequency usage strategy for coordinating different observation states, its mathematical model is to maximize the long-term cumulative discounted normalized throughput, that is, to find an optimal communication strategy Satisfying the following formula:
[0099]
[0100] where γ ∈ (0, 1] is the discount factor; is the normalized throughput of user n at time t.
[0101] It can be seen from the above formula that the frequency usage strategies of other users in the multi-user network will also affect the normalized throughput of communication user n.
[0102] In the communication network, there is no central control device and reliable control channel. Communication users can only rely on the sensing function of the physical layer to independently detect the environment, and its detection result only depends on its own sensing ability and the wireless environment where it is located. Therefore, the environmental information is partially observable to communication users.
[0103] Therefore, the frequency usage decision-making problem of cooperative anti-interference communication of multiple users in a wide-area multi-user communication system can be modeled as a distributed partially observable Markov decision process (Dec-POMDP), which is Described by a quadruple, and the specific definition is:
[0104] State space Used to describe the overall structure configuration information of the communication network, including the positions, transmission frequencies, transmit powers of communication users and interferers, and basic working parameters such as the interference patterns of interferers.
[0105] Expressed as: S(t) = {L u (t), L j (t), P u (t), Pj (t), P u (t), P j (t)}, where L u (t), L j (t) are the sets of position coordinates of the communication user and the jammer respectively; P u (t), P j (t) are the sets of transmission powers of the communication user and the jammer respectively; P u (t), P j (t) are the sets of frequency usage strategies of the communication user and the jammer respectively.
[0106] Observation space Used to describe the result of the communication user's distributed independent spectrum sensing of the environment.
[0107] Adopt a time-expanded model, that is, the spectrum waterfall diagram, as the spectrum state perceived by the communication user within a period of historical time to satisfy the Markov property of the decision-making process. The result of the communication user n's sensing of the spectrum environment at time slot t can be expressed as where Φ is the historical duration.
[0108] Action space Is the set of available channels in the communication network.
[0109] The action a n (t) taken by the communication user n at time slot t is the transmission frequency selected by the user.
[0110] Reward function As the feedback result of the environment on the user's communication situation, it is used to describe the quality of the action selected by the user.
[0111] Specifically, for the multi-user anti-jamming communication network scenario, design a composite reward function The composite reward function consists of a multiplexing reward function and an interference reward function The two parts are both composed of a positive and negative reward mechanism based on the achievement of qualitative goals, that is, when the user's communication is successful this time, the user receives a positive reward, otherwise a negative penalty is given to the user. The two reward functions are respectively used to learn the frequency usage strategies corresponding to the observation state of anti-multi-user communication interference, and the frequency usage strategies corresponding to the observation state of anti-external malicious interference. Depending on the specific tasks of each, the difference lies in the magnitude relationship of the positive and negative reward values, defined as:
[0112]
[0113] where λ reward and λ penalty are the reward and penalty factors respectively, satisfying and
[0114] Therefore, in one embodiment, the cooperative anti-interference communication method may include the following steps:
[0115] Step 1, initialization: For any user n, generate a convolutional neural network network weights and randomly assign values, define the replay buffer size Batch, learning rate α, discount factor γ, and weight coefficient.
[0116] Specifically, Figure 4 FIG. is a schematic structural diagram of a frequency decision model for communication provided in this example. As shown in the figure, the frequency decision model for communication of each communication user transmitter has two convolutional neural networks with the same structure and transform the calculation of the Q value into the calculation of the parameters of the convolutional neural network, and abstract the update of the value table as the training of the network model. The communication user uses the spectral waterfall diagram obtained by independently observing the environment as the input of the network, and then extracts the spectral state features through the convolutional layer. For example, the convolutional layer uses 16 convolutional kernels of a certain size, and the stride is, for example, 2. After the convolutional layer processes, the extracted feature information passes through the fully connected layers 1 and 2 respectively. Among them, the number of neurons in the fully connected layer 2 is the same as the size of the action space. The value output by the neurons of the convolutional neural network corresponds to the action value function under the input observation state respectively
[0117] wherein, the convolutional neural network is specifically as follows: Each communication user transmitter has two convolutional neural networks with the same structure, transforms the calculation of the Q value into the calculation of the parameters of the convolutional neural network, and abstracts the update of the Q value table as the training of the network model.
[0118] Define the loss function of the convolutional neural network of communication user n as:
[0119]
[0120]
[0121] wherein, and represent the weight parameters predicted by communication user n. Target Q u and Target Q j represent the optimization target values of each iteration of the network respectively, and are defined as:
[0122]
[0123]
[0124] Among them, and are the weight parameters of the target Q network.
[0125] Step 2, for any communication time slot t, user n executes the action and obtains the composite reward The environmental state transfers to S t+1 .
[0126] Step 3, user n observes the environmental state puts the experience data into the replay buffer.
[0127] Specifically, putting the experience data into the replay buffer is as follows: The convolutional neural network both adopt the experience replay mechanism, that is, maintaining a replay buffer with a capacity of Batch to store the experience data sampled by the perception module from the environment each time. Among them, the experience data is a five-tuple data (O t , a t , R u,t , R j,t , O t+1 ).
[0128] Among them, at the initial state, the replay buffer does not contain the experience data in the training sample set to be trained. As the communication user and the environment continuously interact, when the amount of experience data reaches Batch, the communication user randomly samples Batch data from the replay buffer, calculates the loss function of the convolutional neural network and updates the network parameters, so as to break the correlation between the data and improve the sample application efficiency. When the amount of experience data reaches the maximum capacity of the replay buffer, the experience data with the longest storage time is deleted to ensure the storage of new experience.
[0129] Step 4, when the amount of data in the replay buffer is greater than or equal to Batch, user n randomly extracts Batch sample data from the replay buffer, updates the target Q value, calculates the loss function of the convolutional neural network and updates the network weights to
[0130] Specifically, updating the target Q value is as follows: At the communication time slot t, for each group of Batch sample data randomly extracted from the replay buffer by communication user n, the following operations are performed: Taking the observed state as the input of the convolutional neural network to obtain the target Q value Taking the observed state as the convolutional neural network Input to obtain the predicted Q value
[0131] Update the target Q value according to the following formula:
[0132]
[0133]
[0134] Step 5, User n uses the current observation state As the input of the updated convolutional neural network To fit and obtain the composite action value function
[0135] Step 6, User n selects an action using a pure greedy policy based on the weighted average of the composite action value function End the algorithm
[0136] Among them, selecting an action using a pure greedy policy based on the weighted average of the composite action value function is as follows: The Q network of communication user n outputs the composite action value function for each input of the observation state The output is the composite action value function
[0137] Comprehensively considering the influence of the composite action value function on the decision-making result of the communication user, calculate And The weighted average of
[0138] And make a decision on the channel accessed by the communication user in the next time slot using a pure greedy policy based on this average value, as shown in the following formula:
[0139]
[0140] Among them, η ∈ [0, 1] is the weight coefficient, indicating the influence degree of the interference action value function on the decision-making result of the communication user. When η approaches 1, the user considers more the influence of external malicious interference on its communication quality when learning the anti-interference strategy; when η approaches 0, the user is more inclined to learn the frequency usage strategy among users inside the network
[0141] To sum up, compared with the prior art, the significant advantages of the present invention are as follows:
[0142] (1) The proposed wide-area multi-user communication frequency usage decision model is completely constructed and fully considers the frequency reuse characteristics brought by channel fading and user location distribution, solving the problems of frequency resource waste and the vulnerability of orthogonality to interference damage existing in the traditional multi-user anti-interference problem where users only rely on orthogonal frequencies to avoid co-frequency interference
[0143] (2) Innovatively model the communication user transmitter as an intelligent agent with learning and decision-making capabilities, design the communication time slot structure, and enable the decision-making results and control information such as ACK between the transmitter and the receiver to be transmitted together with the data information through the service channel, solving the problem of unreliable control channels in complex and harsh spectrum environments.
[0144] (3) Innovate the design of the deep reinforcement learning algorithm, propose a collaborative anti-jamming method based on intelligent frequency reuse, design the network model and the composite reward function, enable communication users to independently learn anti-jamming and frequency reuse strategies under different observation states, and achieve effective solution of the proposed model.
[0145] Furthermore, Figure 5 It is a comparison graph of the network normalized communication throughput varying with the simulation time under different weight coefficients η of the multi-user collaborative anti-jamming algorithm based on intelligent frequency reuse when the jammer releases sweep jamming provided by the embodiment of the present invention. When η = 0.8, the proposed algorithm shows good learning effects, and can quickly converge the normalized throughput of the entire communication network to 1, that is, find the optimal spectrum access strategy for each communication user to coordinate internal frequency usage while combating external interference, and achieve anti-jamming communication for the entire network.
[0146] The case of the weight coefficient η = 0, 1 is the independent DQN algorithm under different reward function settings. It can be seen from the simulation results that when η = 0, the transmitter only based on the output result of the convolutional neural network when making a decision on accessing the channel in the next time slot, that is, the reuse action value function Although the algorithm has a relatively fast convergence speed at this time, the normalized throughput of the network after convergence is about 0.8, that is, the algorithm falls into a local optimal solution. As the weight coefficient η increases, the transmitter will consider more the influence of the interference action value function on the decision-making result, and the normalized throughput of the network after the algorithm converges gradually approaches 1. When η = 1, although the algorithm has a tendency to converge to the optimal solution, it requires a long simulation time.
[0147] Figure 6In the embodiment of the present invention, in the case where the jammer releases comb-shaped interference, it is a comparison graph of the network normalized communication throughput varying with the simulation time of the multi-user cooperative anti-jamming algorithm based on intelligent frequency reuse under different weight coefficients η. When the weight coefficient η = 0.2, the algorithm has the best convergence performance, and as η increases, the algorithm convergence speed gradually decreases. It should be noted that since the interfering channels in the comb-shaped interference mode are relatively fixed, in contrast, the co-channel interference problem among communication users brings more instability and unpredictability to the variation law of the environment. Therefore, compared with the frequency-sweeping interference mode, by reducing the magnitude of the weight coefficient, when the agent makes a decision to access the channel, it considers the composite action value function more for the influence on the decision result, which can make the algorithm achieve a better convergence effect.
[0148] In summary, the frequency decision-making model for wide-area multi-user communication proposed by the present invention constructs a communication scenario where there are both user communication resource conflicts and malicious interference attacks, and fully considers the unreliability of control information transmission in a complex and harsh spectrum environment, as well as the frequency reuse characteristics brought by channel fading and user location distribution in the communication network. It is more practically significant than the traditional communication anti-jamming model; the proposed cooperative anti-jamming method based on intelligent frequency reuse makes decisions to access channels based on the composite action value function, enabling communication users to comprehensively consider the impact of malicious interference in the environment and co-channel interference within the network on their own communication quality when learning the frequency usage strategy, and can adapt to different interference environments by adjusting the magnitude of the weight coefficient η, ensuring the convergence of the algorithm.
[0149] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute a cooperative anti-jamming communication method provided in each of the above embodiments. The method includes: obtaining a spectrum sensing result, where the spectrum sensing result includes the state information of the service channel obtained under different observation states; the observation states include anti-external malicious interference and anti-multi-user communication interference; obtaining the information transmission result of the service channel, where the information transmission result includes successful information transmission and unsuccessful information transmission; based on the communication frequency decision-making model and the spectrum sensing result, obtaining the decision result corresponding to different observation states output by the communication frequency decision-making model; the communication frequency decision-making model is obtained by training a to-be-trained decision model based on a to-be-trained sample set and the information transmission result; the decision result includes the next access service channel information under different observation states; based on the decision results corresponding to different observation states, determining a communication frequency usage strategy after coordinating different observation states.
[0150] The device embodiments described above are merely illustrative, where the units described as separate components may or may not be physically separated. Some or all of the modules can be selected according to actual needs to achieve the objectives of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A collaborative anti-jamming communication method, characterized in that The method includes: Obtaining a spectrum sensing result, where the spectrum sensing result includes the status information of multiple service channels; Obtaining the information transmission result of the service channel, where the information transmission result includes successful information transmission and unsuccessful information transmission; Based on a communication frequency usage decision model and the spectrum sensing result, obtaining the decision results corresponding to different observation states output by the communication frequency usage decision model; the communication frequency usage decision model is obtained by training a decision model to be trained based on a training sample set and an information transmission result; the decision results include the next access service channel information under different observation states; the observation states include anti-external malicious interference and anti-multi-user communication interference; Fitting based on the decision results corresponding to different observation states to obtain a composite action value function; Determining the weighted average value of the decision results corresponding to different observation states based on the composite action value function; According to the weighted average value, adopting a pure greedy strategy to determine a communication strategy after coordinating different observation states, where the communication strategy includes the next access service channel information after coordinating different observation states.
2. The collaborative anti-interference communication method according to claim 1, wherein The step of obtaining the decision results corresponding to different observation states output by the communication frequency usage decision model includes: Taking the spectrum sensing result as the input of the communication frequency usage decision model, and obtaining the spectrum feature information of the spectrum sensing result output by the convolutional layer in the communication frequency usage decision model; Obtaining the reward functions of different observation states in the communication frequency usage decision model; Determining the decision results corresponding to different observation states based on the spectrum feature information and the reward function.
3. The collaborative anti-interference communication method according to claim 1, wherein The step that the communication frequency usage decision model is obtained by training a decision model to be trained based on a training sample set and an information transmission result includes: Obtaining a training sample set; the training sample set includes multiple sample spectrum sensing data; For any sample spectrum sensing data in the multiple sample spectrum sensing data, taking the sample spectrum sensing data as the input of the decision model to be trained, and obtaining the decision results corresponding to different observation states output by the decision model to be trained; Based on the decision results and the information transmission result, adjusting the parameters of the decision model to be trained, and determining a communication decision model.
4. The collaborative anti-interference communication method according to claim 3, wherein The step of adjusting the parameters of the decision model to be trained based on the decision results and the information transmission result includes: Updating the reward function based on the information transmission result.
5. The collaborative anti-interference communication method according to claim 4, wherein The step of adjusting the parameters of the decision model to be trained based on the decision results and the information transmission result further includes: Adopting an experience replay mechanism to adjust the parameters of the decision model to be trained.
6. The collaborative anti-interference communication method according to claim 1, wherein The spectrum sensing result includes spectrum waterfall diagrams under different observation states.
7. A cooperative anti-jamming communication system, characterized in that, The system includes a receiver and a transmitter, and the transmitter executes the steps of the collaborative anti-interference communication method according to any one of claims 1-6 above; The receiver is used to determine whether the transmission information of multiple service channels is successfully transmitted. If the transmission information is successfully transmitted, feedback information is sent to the transmitter. If the transmission information is not successfully transmitted, no feedback information is sent to the transmitter; The transmitter generates information transmission results of multiple service channels according to the feedback information sent by the receiver, and determines a communication frequency usage strategy after coordinating different observation states based on the obtained spectrum sensing results.
8. The collaborative anti-interference communication system according to claim 7, wherein The receiver and the transmitter communicate in time slots, and each time slot includes three stages: data transmission, acknowledgement character (ACK) transmission, and learning decision.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the cooperative anti-jamming communication method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-user communication anti-interference intelligent decision-making method based on deep reinforcement learning
CN115103446A