An anti-reactive interference method based on perception behavior prediction

CN122554048APending Publication Date: 2026-08-11BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]本发明的目的是提供一种基于感知行为预判的抗反应式干扰方法,能够利用深度强化学习从通信交互中自主挖掘并学习反应式干扰机潜藏的感知行为规律,据此预判其干扰行为并前瞻性地制定抗干扰策略,解决现有方案难以自主预判反应式干扰机感知行为规律,无法兼顾通信可靠性与传输效率的问题

Benefits of technology

本发明通过深度强化学习算法自主挖掘并学习反应式干扰机潜藏的感知行为规律,据此预判其干扰行为并前瞻性地制定抗干扰策略,有效解决了传统抗干扰技术难以自主预判反应式干扰机“感知-干扰”内在行为规律、无法兼顾通信可靠性与传输效率的难题,实现了对反应式干扰信号的主动式动态规避,显著提升了通信链路的可靠性与传输效率;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554048A_ABST
    Figure CN122554048A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of wireless communication information security technology, specifically relating to an anti-reactive jamming method based on perception behavior prediction. The method includes constructing a wireless anti-jamming communication system model containing a reactive jammer; modeling the decision-making process of wireless anti-jamming communication as a semi-Markov decision process; constructing an anti-jamming decision model based on soft-actor commentators; defining the online interaction process of the anti-jamming decision model; and designing an iterative training process for the model, iteratively updating network parameters based on collected interaction samples until a preset maximum number of training steps is reached. This invention can utilize deep reinforcement learning to autonomously mine and learn the latent perception behavior patterns of reactive jammers from communication interactions, thereby predicting their jamming behavior and proactively formulating anti-jamming strategies. This solves the problem that existing solutions struggle to autonomously predict the perception behavior patterns of reactive jammers and cannot simultaneously ensure communication reliability and transmission efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication information security technology, specifically relating to an anti-reactive interference method based on perception behavior prediction. Background Technology

[0002] The open broadcast nature of wireless communication makes its links extremely vulnerable to malicious interference attacks. Attackers can significantly reduce the signal-to-interference-plus-noise ratio (SNR) of legitimate receivers and hinder the correct decoding of information by injecting interference signals into the target frequency band. With the popularization of software-defined radio platforms, the deployment threshold for intelligent interference attacks has been significantly lowered. Among them, reactive jammers that adopt the "sensing before jamming" mode have become one of the most serious threats to wireless communication due to their high energy efficiency, difficulty in detection, and ease of deployment. There is an urgent need to develop targeted intelligent anti-interference countermeasures. At present, research on countering reactive jamming has mainly formed four technical paradigms: adversarial, covert, escape, and deception. However, each of these solutions has inherent limitations, such as leading to a surge in energy consumption, being unable to withstand high-sensitivity jammers, lacking dynamic adaptability, or experiencing a sharp decrease in effectiveness under multi-channel jamming scenarios. The core problem is that existing solutions generally ignore the difficult-to-observe perception behavior characteristics of reactive jammers, lack the ability to learn autonomously from their inherent "perception-jamming" behavior patterns, and are unable to formulate forward-looking anti-jamming strategies. Therefore, how to autonomously predict the sensing patterns of jammers and formulate targeted anti-jamming strategies to ensure communication reliability while taking into account transmission efficiency has become a key issue that urgently needs to be addressed. Summary of the Invention

[0003] The purpose of this invention is to provide an anti-reactive jamming method based on perception behavior prediction. This method can use deep reinforcement learning to autonomously mine and learn the hidden perception behavior patterns of reactive jammers from communication interactions, thereby predicting their jamming behavior and proactively formulating anti-jamming strategies. This solves the problem that existing solutions are unable to autonomously predict the perception behavior patterns of reactive jammers and cannot balance communication reliability and transmission efficiency.

[0004] The specific technical solution adopted by this invention is as follows: A method for resisting reactive interference based on perception-based behavior prediction includes the following steps: S1: Construct a wireless anti-jamming communication system model that includes a reactive jammer, and set the working mechanism of the reactive jammer; The wireless anti-jamming communication system model in S1 includes: Transmitter, receiver, reactive jammer, and anti-jamming decision-making agent co-located with the receiver; The transmitter via Each of the four orthogonal channels communicates with the receiver, and the bandwidth of each orthogonal channel is [missing information]. ; In the In each decision-making step, the anti-interference decision-making agent simultaneously selects the orthogonal channel and transmission duration, and transmits the instruction to the transmitter via the control link. The transmitter then selects the center carrier frequency of the corresponding channel accordingly. Duration Signal transmission.

[0005] The jamming strategy of the reactive jammer in S1 is as follows: Each duration is Within the sensing interval, the reactive interference opportunity collects signal samples for energy detection, based on For each sample, the cumulative energy of the detected signal is calculated. When a signal is present, the cumulative energy is used as the basis for... and noise power Normalized energy measurement of the structure Obeying the degree of freedom Non-central parameters are The non-central chi-square distribution is denoted as . ; Therefore, the probability of the reactive jammer detecting the signal is... Represented as: ; in Let be the complementary cumulative distribution function of this distribution. The preset threshold for energy detection, non-central parameter The signal-to-noise ratio at the reactive jamming terminal is determined to satisfy the following conditions: ; Based on the above detection probability The reactive interference opportunity randomly decides whether to interfere with the current sensing channel or switch to another channel.

[0006] S2: Model the decision-making process of wireless anti-jamming communication as a semi-Markov decision process, and define the system state space, system action space and instantaneous reward function; The semi-Markov decision process in S2 is implemented through a quintuple. definition; in, For the system state space; For system action space; Let be the joint probability distribution, representing the probability distribution given the current state. With action At that time, the duration of state transition After transitioning to state The probability of, where Indicates the system status. This represents the combined action output by the intelligent agent; For instant reward functions; The set of all possible state transition times between decision steps; As a discount factor, and .

[0007] The system state space From the recent A sequence defined as a series of consecutive single-step interactive observations; The system action space By joint action Define, where The index of the communication channel selected for the current decision step. for The index set corresponding to each discrete orthogonal channel; The selected transmission duration for the current decision step. It is a discrete set of values ​​for transmission duration; The instant reward function The rate of successful symbol transmission at the receiver's receiving end is constructed.

[0008] S3: Construct an interference-resistant decision-making model based on soft actor critics; The interference-resistant decision-making model based on soft actor critics in S3 includes an actor network, two critic networks with completely identical structures, two target critic networks with completely identical structures to the critic network, and a temperature parameter. ; The actor network is in state The input is in the action space of the discrete joint system. probability distribution on ,in The parameters of the actor network are defined as follows, and the actor network is continuously optimized during training to approximate the optimal random strategy. ; The optimal stochastic strategy aims to maximize the expected sum of future cumulative rewards and strategy entropy, and is defined as follows: (5) in , , indicating up to the The cumulative system time for each decision step Representing state The entropy value of the next strategy, The temperature parameter is used to regulate the balance between exploration and utilization. The update is performed by continuously minimizing the objective function as follows: (6) in, , Scaling factor Represents the total number of optional discrete joint actions; The parameters of the two critic networks are respectively and They are used to estimate the soft Q value. The soft Q-value characterizes the performance of a given state. Next action The weighted expected sum of the accumulated discounted reward and policy entropy is used to evaluate the comprehensive value of the action under the dual objectives of maximizing reward and maintaining exploration ability. To approximate the true soft Q value, the critic network is iteratively updated by minimizing the temporal difference error. The training loss function used to update the critic network is defined as: (7) in, This represents the soft Q objective value, while The soft-state value evaluated by the target network is expressed in mathematical form as: (8) in Indicates the first The network parameters of the dual-objective commentator are calculated deterministically using the formula described above, unlike the method of approximating the soft-state value function by sampling actions in a continuous action space, thus eliminating the variance in objective estimation. Furthermore, the operation of minimizing the output of the dual-objective network effectively alleviates the problem of overestimation of the soft Q-value and enhances the stability of the policy learning process. Based on this, the actor network updates its parameters by minimizing the following objective function according to the soft Q-values ​​provided by the dual commentators. In order to optimize the strategy: (9) S4: Define the online interaction process of the interference-resistant decision-making model; The online interaction process in S4 is executed cyclically according to decision steps. The complete process of a single decision step is as follows: In the At the start of each decision step, the single-step interactive observation value of the agent acquisition system. , will recently The system state is obtained by concatenating consecutive single-step interactive observations. The system state sequence The input is fed into the current actor network to obtain the joint action for the current decision step. ; Furthermore, the aforementioned joint action The message is sent to the transmitter to control the transmitter in the channel. On the duration The signal transmission is collected, and the communication results at the receiver are used to calculate the instant reward value corresponding to this decision. Simultaneously collect single-step interactive observations The system state at the next moment is obtained by splicing the data together. The experience generated from this interaction Store in the experience replay pool middle.

[0009] S5: Design the iterative training process of the model, and iteratively update the network parameters based on the collected interaction samples until the preset maximum number of training steps is reached to obtain the optimal anti-interference decision model.

[0010] The iterative training process in S5 is performed in the experience replay pool. Once the amount of experience stored in the database reaches the preset training start threshold, the training process begins. The training process is executed synchronously and in parallel with the online interactive data collection. The specific process is as follows: From the experience replay pool Randomly select a small batch of experience; Next state Input the actor network respectively First Target Commentator Network Second Target Commentator Network The action probability distribution output by the actor network is obtained. and the first target commentator network Second Target Commentator Network The soft corresponding to all optional actions output value and ; Calculate the soft state value of the next state based on equation (8). And calculate the soft based on this. target value ,in To obtain the target value for the actual time taken for state transitions in this experience. Then, the system status will be... Input into the first target commentator network respectively Second Target Commentator Network Extract the actual actions performed. Corresponding estimation soft value and Combined with software target value The gradient descent method is used to minimize the loss function defined in formula (7) and update the commentator network; System status The inputs are fed into the actor network and the updated critic network to obtain the action probability distribution. The minimum estimated soft Q value is obtained, and gradient descent is applied to Equation (9) to update the actor network parameters. Subsequently, gradient descent is used to minimize the loss function defined in Equation (6) to update the temperature parameters. .

[0011] The target commentator network parameters are updated using a soft update method, with the following update formula: (10) in This is a preset smoothing coefficient used to control the degree of conservatism in the updating of the target network parameters; In the above process, the parameter update based on gradient descent follows a unified rule, that is, for any parameter to be updated... Its gradient descent update takes the form of: (11) in, The learning rate for the corresponding parameter. For the corresponding loss function For parameters The gradient.

[0012] The technical effects achieved by this invention are as follows: This invention autonomously mines and learns the latent perception behavior patterns of reactive jammers through deep reinforcement learning algorithms, thereby predicting their jamming behavior and proactively formulating anti-jamming strategies. This effectively solves the problem that traditional anti-jamming technologies cannot autonomously predict the inherent "perception-jamming" behavior patterns of reactive jammers and cannot balance communication reliability and transmission efficiency. It achieves proactive dynamic avoidance of reactive jamming signals and significantly improves the reliability and transmission efficiency of communication links. Furthermore, by modeling the anti-interference decision-making process as a semi-Markov decision-making process and innovatively introducing time anchor parameters in the state representation, the agent can effectively capture the time offset characteristics brought about by the perception cycle of the jammer, overcoming the limitation of traditional Markov decision-making processes in being unable to handle asynchronous interactions between legitimate users and jammers.

[0013] Based on the aforementioned effects, a wireless anti-jamming communication system model incorporating a reactive jammer is constructed, and the jammer's operating mechanism is defined. The anti-jamming communication decision-making process is modeled as a semi-Markov decision process, defining a state space including time anchor parameters, an action space combining channel selection and transmission duration, and a reward function based on the successful transmission symbol rate. A mathematical mapping between anti-jamming performance indicators and system variables is established. An anti-jamming decision-making model based on soft-actor commentators is constructed, and the model's online interaction process is defined to collect empirical data. An iterative training process for the model is designed, continuously updating the dual-commentator network, actor network, and temperature parameters until the preset maximum number of training steps is reached, obtaining the optimal anti-jamming decision-making model. Based on the output of the optimal anti-jamming decision-making model, forward-looking intelligent communication anti-jamming under dynamic interference environments is achieved. Attached Figure Description

[0014] Figure 1 This is a diagram illustrating the specific steps in this invention. Detailed Implementation

[0015] To make the objectives and advantages of this invention clearer, the invention will be specifically described below with reference to embodiments. It should be understood that the following text is merely used to describe one or more specific embodiments of the invention and does not strictly limit the scope of protection specifically claimed by the invention.

[0016] like Figure 1 As shown, a method for resisting reactive interference based on perception behavior prediction includes the following steps: S1: Construct a wireless anti-jamming communication system model that includes a reactive jammer, and set the working mechanism of the reactive jammer; The wireless anti-jamming communication system includes a legitimate transmitter, a legitimate receiver, a reactive jammer, and an anti-jamming decision-making agent co-located with the legitimate receiver. The transmitter... The receiver communicates via four orthogonal channels, the center frequencies of which are: The bandwidth of each channel is This communication process takes place over a series of discrete decision steps, and... As an index for decision steps. In the... In each decision-making step, the anti-interference decision-making agent simultaneously selects the channel and transmission duration, and then transmits the instructions to the transmitter via the control link. The transmitter then selects the center carrier frequency of the corresponding channel accordingly. Duration The signal transmission is designed to avoid detection and interference by reactive jammers. During signal transmission, if not interfered with by a reactive jammer, the legitimate receiver receives the signal... It can be represented as: (1) If interfered with by a reactive jammer, the received signal will be... Represented as: (2) In the above formula, The transmission power of a legal transmitter, For the channel coefficients of the communication link, The baseband signal transmitted by the transmitter. This refers to the transmit power of the reactive jammer. The channel coefficients of the interference link. The jamming signal emitted by the jammer. This is the additive white Gaussian noise at the receiver. Based on this, the signal-to-interference-plus-noise ratio (SNR) at the receiver... Represented as: (3) in This represents noise power.

[0017] The reactive jammer employs a "sensing first, then jamming" operating mode, and its jamming strategy is as follows: In each time interval... Interference opportunity acquisition within the sensing interval One signal sample is used for energy detection, where To perceive duration, This is the sensing bandwidth of the reactive jammer. This is for floor function. Based on this... From the samples, the cumulative energy of the detected signal can be calculated as follows: ,in The first one collected by the jammer One received signal sample. Normalized energy measurement when a valid signal is present. Obeying the degree of freedom Non-central parameters are The noncentral chi-square distribution is denoted as . Therefore, the probability of the jammer detecting the signal is... It can be represented as ,in Let be the complementary cumulative distribution function of this distribution. The preset threshold for energy detection, non-central parameter Determined by the signal-to-noise ratio at the jammer end, satisfying Based on this detection probability, the interference chance random decision determines whether to perform a duration of [duration missing] on the current sensing channel. Interference, or switching to another channel, repeats the above sensing and detection process. Both switching behaviors will generate mode switching delays. or channel switching delay .

[0018] The system operating frequency band used in this embodiment is 300–380MHz, which is evenly divided into 16 orthogonal frequency channels, i.e. Bandwidth of each sub-channel Legitimate transmitter transmission power reactive jammer transmit power Noise power spectral density The jammer's sensing duration Interference duration Mode switching latency Channel switching delay The communication link adopts the Rayleigh fading channel model and the modulation method is quadrature phase shift keying.

[0019] S2: Model the decision-making process of wireless anti-jamming communication as a semi-Markov decision process, and define the system state space, system action space and instantaneous reward function; The semi-Markov decision process uses a quintuple. Definition. Wherein, For the system state space; For system action space; Let be the joint probability distribution, representing the probability distribution given the current state. With action At that time, the duration of state transition After transitioning to state The probability of, where Indicates the system status. This represents the combined action output by the intelligent agent; For instant reward functions; The set of all possible state transition times between decision steps; Discount factor and This is used to balance the importance of immediate rewards and future rewards.

[0020] The system state space is composed of the most recent The sequence definition consists of a series of continuous single-step interactive observations; specifically, the system state. Represented as Among them, single-step interactive observations In the observed values middle, Select an index for the channel. For transmission duration, For instant reward value, The time anchor parameter enables the agent to capture the time offset characteristics caused by the sensing cycle of the interfering machine, and its expression is: ,in Indicates the system's cumulative runtime. These are predefined time block constants.

[0021] The system action space consists of joint actions Define, where The index of the communication channel selected for the current decision step. for The index set corresponding to each discrete orthogonal channel; The selected transmission duration for the current decision step. Let be the discrete set of values ​​for transmission duration, with an interval of ,in and These are the minimum and maximum allowed transmission times, respectively. The value of directly determines the first in the semi-Markov decision process. State transition time corresponding to the decision step .

[0022] Instant reward function Constructed based on the successful transmission symbol rate at the legitimate receiver, the expression is: (4) in, The actual signal-to-interference-plus-noise ratio at the legitimate receiver end. This is a preset signal demodulation threshold, representing the minimum signal-to-interference-plus-noise ratio (SNR) limit required for the receiver to correctly demodulate the signal. For the symbol transmission rate of the system, Symbol error rate, A reward scaling constant is used to shrink the reward value to an appropriate order of magnitude. A fixed penalty value corresponding to transmission failure, when the condition is met. The transmission was deemed to have failed.

[0023] The historical observation sequence length used in this embodiment Time block constants Transmission duration range The discrete value step size is 1ms, and the signal demodulation threshold is... The system's bit transmission rate Symbol transmission rate For Reward scaling constant Penalty value for transmission failure Discount factor .

[0024] S3: Construct an interference-resistant decision-making model based on soft actor critics; The interference-resistant decision-making model based on soft actor critics includes an actor network, two critic networks with identical structures, two target critic networks with identical structures to the critic networks, and a temperature parameter. Among them, the actor network is in status The input is in the action space of the discrete joint system. probability distribution on ,in These are the parameters for the actor network. This probability distribution is then sampled to output actions used for interference-resistant communication. During training, the actor network is continuously optimized to approximate the optimal random strategy. This strategy aims to maximize the expected sum of future cumulative rewards and policy entropy, and is defined as follows: (5) in , (and ) indicates up to the The cumulative system time for each decision step Representing state The entropy value of the next strategy, Temperature parameters are used to regulate the balance between exploration and utilization. These parameters are also used to allow the system to autonomously adapt to this balance during training. The update is performed by minimizing the objective function as follows: (6) In the formula, , Scaling factor This represents the total number of optional discrete joint actions.

[0025] The parameters of the two critic networks are respectively and They are used to estimate the soft Q value. This value represents the state. Next action The weighted expected sum of the accumulated discounted reward and policy entropy is used to evaluate the comprehensive value of the action under the dual objectives of maximizing reward and maintaining exploration capability. To approximate the true soft Q value, the critic network is iteratively updated by minimizing the temporal difference error. Specifically, the training loss function of the critic network is defined as: (7) in, Indicates the soft Q objective value. This represents the soft state value evaluated by the target network. Given that this method uses a discrete action space, the target commentator network can target the state... Output A soft Q value. Based on the definition of the soft-state value function. It can be explicitly interpreted as: (8) in Indicates the first The network parameters of the dual-objective commentator are calculated deterministically using the formula described above, unlike the method of approximating the soft-state value function by sampling actions in a continuous action space, thus eliminating the variance in objective estimation. Furthermore, the operation of minimizing the output of the dual-objective network effectively alleviates the problem of overestimation of the soft Q-value and enhances the stability of the policy learning process. Based on this, the actor network updates its parameters by minimizing the following objective function according to the soft Q-values ​​provided by the dual commentators. In order to optimize the strategy: (9) In this embodiment, both the actor network and the critic network employ a three-layer fully connected neural network structure, with 128 and 256 hidden layer neurons respectively, and the ReLU activation function is used. The input dimension of the actor network is... The output dimension is (That is, a combination of 16 channels and 30 transmission durations); the critic network also has an input dimension of 20 and an output dimension of 480. Entropy scaling factor The learning rate of the actor network is... The learning rate of the critic network The learning rate of the temperature parameter is .

[0026] S4: Define the online interaction process of the interference-resistant decision-making model; The online interaction process is executed cyclically in decision steps. The complete process of a single decision step is as follows: In the first step... At the start of each decision step, the single-step interactive observation value of the agent acquisition system. , will recently The system state is obtained by concatenating consecutive single-step interactive observations. Subsequently, the system state sequence will be... The input is fed into the current actor network, and the actor network is based on the state of the input. Output the execution probability distribution of all optional actions in the joint action space. Based on this probability distribution, random sampling is performed to obtain the joint action of the current decision step. Next, joint actions will be carried out. The message is sent to the transmitter to control the transmitter in the channel. On the duration The signal transmission is completed. After this transmission, the communication results from the legitimate receiver are collected, and the immediate reward value corresponding to this decision is calculated. Simultaneously collect single-step interactive observations The system state at the next moment is obtained by splicing the data together. And determine the state transition duration corresponding to this decision. The experience generated from this interaction Store in the experience replay pool Among them Used to store historical interaction data for offline training.

[0027] The storage capacity of the experience playback pool used in this embodiment is set to a maximum of 20,000 experience data entries.

[0028] S5: Design the iterative training process of the model, and iteratively update the network parameters based on the collected interaction samples until the preset maximum number of training steps is reached to obtain the optimal anti-interference decision model.

[0029] The iterative training process in the experience replay pool Once the amount of experience stored in the database reaches the preset training start threshold, the training process begins. The training process is executed synchronously and in parallel with the online interactive data collection. The specific process is as follows: First, from the experience replay pool A small batch of experiences is randomly selected. Then, one of these experiences is used... For example, the next state Enter actor network respectively and two target commentator networks and The probability distribution of the actor's network output motion is obtained. And the soft Q values ​​corresponding to all optional actions output by the two target commentator networks. and Subsequently, the soft state value of the next state is calculated according to formula (8). And calculate the soft Q target value accordingly. ,in This represents the actual time taken for state transitions in this experience. Obtain the target value. Then, the system status will be... The inputs are fed into two separate critic networks to extract the actual actions performed. The corresponding estimated soft Q value and Combined with soft Q target value Gradient descent is used to minimize the loss function defined in formula (7) and update the commentator network.

[0030] System status The inputs are fed into the actor network and the updated critic network to obtain the action probability distribution. The minimum estimated soft Q value is obtained, and gradient descent is applied to Equation (9) to update the actor network parameters. Subsequently, gradient descent is used to minimize the loss function defined in Equation (6) to update the temperature parameters. .

[0031] After updating the temperature parameters, based on the updated parameters of the two critic networks, the parameters of the two target critic networks are updated separately using a soft update method. The soft update formula is as follows: (10) in This is a preset smoothing coefficient. It is used to control the degree of conservatism in the updating of the target network parameters.

[0032] In the above process, the parameter update based on gradient descent follows a unified rule, that is, for any parameter to be updated... Its gradient descent update takes the form of: (11) in, The learning rate for the corresponding parameter. For the corresponding loss function For parameters The gradient.

[0033] After updating all the above parameters, the single batch of model training is completed. This process is repeated until the number of training steps reaches the preset maximum number of training steps. Training is then stopped, and the actor network and dual critic network with the updated parameters are determined as the optimal anti-interference decision model.

[0034] The number of mini-batch samples extracted for each training session in this embodiment is... Target smoothing coefficient The training start threshold is set to 1000 empirical data points, and the maximum training steps are set to 50,000 steps.

[0035] This invention autonomously mines and learns the latent perception behavior patterns of reactive jammers through deep reinforcement learning algorithms, thereby predicting their jamming behavior and proactively formulating anti-jamming strategies. This effectively solves the problem that traditional anti-jamming technologies cannot autonomously predict the inherent "perception-jamming" behavior patterns of reactive jammers and cannot balance communication reliability and transmission efficiency. It achieves proactive dynamic avoidance of reactive jamming signals and significantly improves the reliability and transmission efficiency of communication links. Furthermore, by modeling the anti-interference decision-making process as a semi-Markov decision-making process and innovatively introducing time anchor parameters in the state representation, the agent can effectively capture the time offset characteristics brought about by the perception cycle of the jammer, overcoming the limitation of traditional Markov decision-making processes in being unable to handle asynchronous interactions between legitimate users and jammers.

[0036] Based on the aforementioned effects, a wireless anti-jamming communication system model incorporating a reactive jammer is constructed, and the jammer's operating mechanism is defined. The anti-jamming communication decision-making process is modeled as a semi-Markov decision process, defining a state space including time anchor parameters, an action space combining channel selection and transmission duration, and a reward function based on the successful transmission symbol rate. A mathematical mapping between anti-jamming performance indicators and system variables is established. An anti-jamming decision-making model based on soft-actor commentators is constructed, and the model's online interaction process is defined to collect empirical data. An iterative training process for the model is designed, continuously updating the dual-commentator network, actor network, and temperature parameters until the preset maximum number of training steps is reached, obtaining the optimal anti-jamming decision-making model. Based on the output of the optimal anti-jamming decision-making model, forward-looking intelligent communication anti-jamming under dynamic interference environments is achieved.

[0037] The above description is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained in this invention are implemented according to conventional methods in the art unless otherwise specified or limited.

Claims

1. An anti-reactive interference method based on perception behavior prediction, characterized in that, Includes the following steps: S1: Construct a wireless anti-jamming communication system model that includes a reactive jammer, and set the working mechanism of the reactive jammer; S2: Model the decision-making process of wireless anti-jamming communication as a semi-Markov decision process, and define the system state space, system action space and instantaneous reward function; S3: Construct an interference-resistant decision-making model based on soft actor critics; S4: Define the online interaction process of the interference-resistant decision-making model; S5: Design the iterative training process of the model, and iteratively update the network parameters based on the collected interaction samples until the preset maximum number of training steps is reached to obtain the optimal anti-interference decision model.

2. The method for resisting reactive interference based on perception behavior prediction according to claim 1, characterized in that: The wireless anti-jamming communication system model in S1 includes: Transmitter, receiver, reactive jammer, and anti-jamming decision-making agent co-located with the receiver; The transmitter via Each of the four orthogonal channels communicates with the receiver, and the bandwidth of each orthogonal channel is [missing information]. ; In the In each decision-making step, the anti-interference decision-making agent simultaneously selects the orthogonal channel and transmission duration, and transmits the instruction to the transmitter via the control link. The transmitter then selects the center carrier frequency of the corresponding channel accordingly. Duration Signal transmission.

3. The method for resisting reactive interference based on perception behavior prediction according to claim 2, characterized in that: The jamming strategy of the reactive jammer in S1 is as follows: Each duration is Within the sensing interval, the reactive interference opportunity collects signal samples for energy detection, based on For each sample, the cumulative energy of the detected signal is calculated. When a signal is present, the cumulative energy is used as the basis for... and noise power Normalized energy measurement of the structure Obeying the degree of freedom Non-central parameters are The non-central chi-square distribution is denoted as . ; Therefore, the probability of the reactive jammer detecting the signal is... Represented as: ; in Let be the complementary cumulative distribution function of this distribution. The preset threshold for energy detection, non-central parameter The signal-to-noise ratio at the reactive jamming terminal is determined to satisfy the following conditions: ; Based on the above detection probability The reactive interference opportunity randomly decides whether to interfere with the current sensing channel or switch to another channel.

4. The method for resisting reactive interference based on perception behavior prediction according to claim 3, characterized in that: The semi-Markov decision process in S2 is implemented through a quintuple. definition; in, For the system state space; For system action space; Let be the joint probability distribution, representing the probability distribution given the current state. With action At that time, the duration of state transition After transitioning to state The probability of, where Indicates the system status. This represents the combined action output by the intelligent agent; For instant reward functions; The set of all possible state transition times between decision steps; As a discount factor, and .

5. The method for resisting reactive interference based on perception behavior prediction according to claim 4, characterized in that: The system state space From the recent A sequence defined as a series of consecutive single-step interactive observations; The system action space By joint action Define, where The index of the communication channel selected for the current decision step. for The index set corresponding to each discrete orthogonal channel; The selected transmission duration for the current decision step. It is a discrete set of values ​​for transmission duration; The instant reward function The rate of successful symbol transmission at the receiver's receiving end is constructed.

6. The method for resisting reactive interference based on perception behavior prediction according to claim 5, characterized in that: The interference-resistant decision-making model based on soft actor critics in S3 includes an actor network, two critic networks with completely identical structures, two target critic networks with completely identical structures to the critic network, and a temperature parameter. ; The actor network is in state The input is in the action space of the discrete joint system. probability distribution on ,in The parameters of the actor network are defined as follows, and the actor network is continuously optimized during training to approximate the optimal random strategy. ; The optimal stochastic strategy aims to maximize the expected sum of future cumulative rewards and strategy entropy, and is defined as follows: (5) in , , indicating up to the The cumulative system time for each decision step, Representing state The entropy value of the next strategy, The temperature parameter is used to regulate the balance between exploration and utilization. The update is performed by continuously minimizing the objective function as follows: (6) in, , Scaling factor Represents the total number of optional discrete joint actions; The parameters of the two critic networks are respectively and They are used to estimate the soft Q value. The soft Q-value characterizes the performance of a given state. Next action The weighted expected sum of the cumulative discounted reward and policy entropy is used to evaluate the comprehensive value of the action under the dual objectives of maximizing reward and maintaining exploration ability. To approximate the true soft Q value, the critic network is iteratively updated by minimizing the temporal difference error. The training loss function used to update the critic network is defined as: (7) in, This represents the soft Q objective value, while The soft-state value evaluated by the target network is expressed in mathematical form as: (8) in Indicates the first The operation of minimizing the output of the dual-objective network by taking the network parameters of the dual-objective commentator effectively alleviates the problem of overestimation of the soft Q value and enhances the stability of the policy learning process. Based on this, the actor network updates its parameters by minimizing the following objective function according to the soft Q value provided by the dual commentators. In order to optimize the strategy: (9)。 7. The method for resisting reactive interference based on perception behavior prediction according to claim 6, characterized in that: The online interaction process in S4 is executed cyclically according to decision steps. The complete process of a single decision step is as follows: In the At the start of each decision step, the single-step interactive observation value of the agent acquisition system. , will recently The system state is obtained by concatenating consecutive single-step interactive observations. The system state sequence The input is fed into the current actor network to obtain the joint action for the current decision step. ; Furthermore, the aforementioned joint action The message is sent to the transmitter to control the transmitter in the channel. On the duration The signal transmission is collected, and the communication results at the receiver are used to calculate the instant reward value corresponding to this decision. Simultaneously collect single-step interactive observations The system state at the next moment is obtained by splicing the data together. The experience generated from this interaction Store in the experience replay pool middle.

8. The method for resisting reactive interference based on perception behavior prediction according to claim 7, characterized in that: The iterative training process in S5 is performed in the experience replay pool. Once the amount of experience stored in the database reaches the preset training start threshold, the training process begins. The training process is executed synchronously and in parallel with the online interactive data collection. The specific process is as follows: From the experience replay pool Randomly select a small batch of experience; Next state Input the actor network respectively First Target Commentator Network Second Target Commentator Network The action probability distribution output by the actor network is obtained. and the first target commentator network Second Target Commentator Network The soft Q values ​​for all optional actions output and ; Calculate the soft state value of the next state based on equation (8). And calculate the soft Q target value accordingly. ,in To obtain the target value for the actual time taken for state transitions in this experience. Then, the system status will be... Input into the first target commentator network respectively Second Target Commentator Network Extract the actual actions performed. The corresponding estimated soft Q value and Combined with soft Q target value The gradient descent method is used to minimize the loss function defined in formula (7) and update the commentator network; System status The inputs are fed into the actor network and the updated critic network to obtain the action probability distribution. The minimum estimated soft Q value is obtained, and gradient descent is applied to Equation (9) to update the actor network parameters. Subsequently, gradient descent is used to minimize the loss function defined in Equation (6) to update the temperature parameters. ; The target commentator network parameters are updated using a soft update method, with the following update formula: (10) in This is a preset smoothing coefficient used to control the degree of conservatism in the updating of the target network parameters; In the above process, the parameter update based on gradient descent follows a unified rule, that is, for any parameter to be updated... Its gradient descent update takes the form of: (11) in, The learning rate for the corresponding parameter. For the corresponding loss function For parameters The gradient.