Anti-nonlinear frequency sweeping interference communication method and system based on proximal policy optimization
By modeling the anti-nonlinear frequency sweeping interference process of a wireless communication system as a Markov decision process, and utilizing deep neural networks and near-end policy optimization algorithms, precise suppression of dynamic nonlinear frequency sweeping interference is achieved, thereby improving the transmission reliability of wireless communication.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-28
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies are insufficient to effectively suppress dynamic nonlinear frequency sweep interference, leading to a decrease in the transmission reliability of wireless communication systems in complex electromagnetic environments.
The anti-nonlinear frequency sweeping interference communication process is modeled as a Markov decision process. A deep neural network is used for the policy network and value network, combined with a near-end policy optimization algorithm. The model weights are updated through an adaptive optimization algorithm to achieve accurate channel selection and interference avoidance.
It achieves efficient interference suppression in wireless communication systems under complex electromagnetic environments, significantly improving transmission reliability.
Smart Images

Figure CN121585296B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of wireless communication, in particular to a nonlinear sweep jamming resistant communication method and system based on near-end policy optimization. BACKGROUND
[0002] Wireless communication uses electromagnetic waves as the carrier of information transmission, and has a naturally open transmission channel. This open characteristic makes wireless communication systems vulnerable to serious threats from various unintentional or malicious interference. As a typical malicious interference pattern, sweep jamming implements periodic and rapid scanning on each frequency point in the target frequency band to form a blocking interference effect covering a wide frequency band with a small power cost, and becomes a widely used malicious interference pattern. Existing sweep jamming suppression methods usually adjust communication parameters such as carrier frequency based on interference pattern recognition and parameter extraction to achieve interference suppression. Although this method is effective for linear sweep jamming with fixed parameters, its performance significantly decreases under dynamic nonlinear sweep jamming with varying key parameters such as sweep direction and sweep rate. This is mainly because the time-varying characteristics of interference make it difficult for traditional methods to accurately model and extract interference feature parameters.
[0003] With the continuous improvement of the performance of digital communication devices, dynamic nonlinear sweep jamming has gradually been implemented, posing a significant threat to the reliable transmission of wireless communication systems. In recent years, reinforcement learning has shown unique advantages in the field of communication jamming suppression due to its outstanding dynamic policy optimization capability. It continuously optimizes transmission strategies through interaction with the interference environment, providing a new technical path for dealing with dynamic nonlinear sweep jamming problems. Therefore, how to use reinforcement learning technology to address the problems of unpredictable time-frequency state, excessively large policy space dimension, and insufficient model generalization ability under dynamic nonlinear sweep jamming, and to achieve efficient suppression of nonlinear sweep jamming, thereby improving the reliability of wireless communication transmission in complex electromagnetic environments, has become a key technical problem that needs to be solved. SUMMARY
[0004] Therefore, it is necessary to provide a nonlinear sweep jamming resistant communication method and system based on near-end policy optimization to solve the above technical problems.
[0005] A nonlinear sweep jamming resistant communication method based on near-end policy optimization, the method comprising:
[0006] The anti-nonlinear frequency sweeping jamming communication process is modeled as a Markov decision process, wherein, at each time step, the state includes a historical jamming trajectory, a current transmission channel, a distance between a current jamming channel and the transmission channel, the action includes a transmission channel selection of a next time slot, and the reward value includes a product of a penalty term corresponding to a jamming avoidance effect and a penalty term corresponding to a channel switching cost, and the jamming avoidance effect penalty term is valued according to a distance gradient between the current jamming channel and the transmission channel;
[0007] a deep neural network model including a policy network and a value network is constructed; the policy network is used to output selection probabilities of each transmission channel according to a current state, and the value network is used to evaluate an estimated value of the current state;
[0008] under guidance of the current policy network, the transmitter and the receiver synchronously perform a transmission channel selection decision of a previous time slot, switch to a corresponding channel to perform data communication, and store interactive data of the receiver and the jammer as experience samples into an experience pool; the interactive data includes a previous state, an action, a reward value and a current state;
[0009] experience samples are sampled from the experience pool, parameters of the policy network and the value network are optimized by using a proximal policy optimization algorithm and the experience samples, and iterative updating of model weights is completed by using an adaptive optimization algorithm;
[0010] the updated policy network is used to select a transmission channel of a next time slot and fed back to the transmitter.
[0011] An anti-nonlinear frequency sweeping jamming communication system based on a proximal policy optimization, the system comprising:
[0012] a model construction module configured to model an anti-nonlinear frequency sweeping jamming communication process as a Markov decision process, wherein, at each time step, the state includes a historical jamming trajectory, a current transmission channel, a distance between a current jamming channel and the transmission channel, the action includes a transmission channel selection of a next time slot, and the reward value includes a product of a penalty term corresponding to a jamming avoidance effect and a penalty term corresponding to a channel switching cost, and the jamming avoidance effect penalty term is valued according to a distance gradient between the current jamming channel and the transmission channel;
[0013] a network construction module configured to construct a deep neural network model including a policy network and a value network; the policy network is used to output selection probabilities of each transmission channel according to a current state, and the value network is used to evaluate an estimated value of the current state;
[0014] The sample obtaining module is configured to, under the guidance of the current policy network, synchronize the transmitter and the receiver to execute the transmission channel selection decision of the previous time slot, switch to the corresponding channel for data communication, and store the interaction data between the receiver and the jammer as experience samples into an experience pool; the interaction data includes the previous state, the action, the reward value, and the current state;
[0015] The network updating module is configured to sample experience samples from the experience pool, optimize the parameters of the policy network and the value network by using a proximal policy optimization algorithm and the experience samples, and complete iterative updating of model weights by using an adaptive optimization algorithm;
[0016] The result output module is configured to select the transmission channel of the next time slot by using the updated policy network and feed back to the transmitter.
[0017] A computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:
[0018] The anti-nonlinear sweep jamming communication process is modeled as a Markov decision process, wherein, at each time step, the state includes a historical jamming trajectory, a current transmission channel, a distance between a current jamming channel and the transmission channel, the action includes transmission channel selection of the next time slot, the reward value includes a product of a penalty item corresponding to a jamming avoidance effect and a penalty item corresponding to a channel switching cost, and the jamming avoidance effect penalty item is valued according to a distance gradient between the current jamming channel and the transmission channel;
[0019] A deep neural network model comprising a policy network and a value network is constructed; the policy network is configured to output selection probabilities of each transmission channel according to a current state, and the value network is configured to evaluate an estimated value of the current state;
[0020] Under the guidance of the current policy network, the transmitter and the receiver synchronize to execute the transmission channel selection decision of the previous time slot, switch to the corresponding channel for data communication, and store the interaction data between the receiver and the jammer as experience samples into an experience pool; the interaction data includes the previous state, the action, the reward value, and the current state;
[0021] Experience samples are sampled from the experience pool, the parameters of the policy network and the value network are optimized by using a proximal policy optimization algorithm and the experience samples, and iterative updating of model weights is completed by using an adaptive optimization algorithm;
[0022] The transmission channel of the next time slot is selected by using the updated policy network and fed back to the transmitter.
[0023] A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the following steps:
[0024] The anti-nonlinear frequency-sweep interference communication process is modeled as a Markov decision process, wherein, at each time step, the state includes a historical interference trajectory, a current transmission channel, a distance between the current interference channel and the transmission channel, the action includes a transmission channel selection of a next time slot, the reward value includes a product of a penalty term corresponding to an interference avoidance effect and a penalty term corresponding to a channel switching cost, and the interference avoidance effect penalty term is valued according to a distance gradient between the current interference channel and the transmission channel;
[0025] A deep neural network model including a policy network and a value network is constructed; the policy network is used to output selection probabilities of each transmission channel according to a current state, and the value network is used to evaluate an estimated value of the current state;
[0026] Under the guidance of the current policy network, the transmitter and the receiver synchronously perform a transmission channel selection decision of a previous time slot, switch to a corresponding channel to perform data communication, and store interactive data of the receiver and the jammer into an experience pool as experience samples; the interactive data includes a previous state, an action, a reward value and a current state;
[0027] Experience samples are sampled from the experience pool, and parameters of the policy network and the value network are optimized by using a proximal policy optimization algorithm and the experience samples, and iterative updates of model weights are completed by using an adaptive optimization algorithm;
[0028] The updated policy network is used to select a transmission channel of a next time slot and is fed back to the transmitter.
[0029] The anti-nonlinear frequency-sweep interference communication method and system based on the proximal policy optimization can comprehensively capture a dynamic change trend and a real-time position correlation of the frequency-sweep interference by modeling the anti-nonlinear frequency-sweep interference communication process as a Markov decision process with multi-dimensional states, fusing a historical interference trajectory, a current transmission channel and a distance between an interference channel and a transmission channel, and realizing a precise channel selection decision by combining a gradient-valued reward mechanism of two objectives of interference avoidance and switching cost, a policy network outputting selection probabilities of each channel, a value network evaluating a long-term benefit of a state, avoiding blind switching or passive interference, and storing interactive data in an experience pool, optimizing double-network parameters by using a proximal policy optimization algorithm, and completing weight iteration by combining adaptive optimization. The wireless communication system can autonomously learn an optimal anti-interference strategy, efficiently suppresses dynamic nonlinear frequency-sweep interference, and significantly improves the reliability of wireless communication transmission in a complex electromagnetic environment. BRIEF DESCRIPTION OF DRAWINGS
[0030] Figure 1 A flowchart of the anti-nonlinear frequency-sweep interference communication method based on the proximal policy optimization in one embodiment is shown;
[0031] Figure 2 a time slot in an embodiment;
[0032] Figure 3 a structural diagram of a policy network and a value network in an embodiment, wherein, Figure 3 (a) a structural diagram of a policy network, Figure 3 (b) a structural diagram of a value network;
[0033] Figure 4 a flow diagram of a method for anti-nonlinear sweep jamming communication based on proximal policy optimization in a specific embodiment;
[0034] Figure 5 a structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION
[0035] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0036] First, the following assumptions are made:
[0037] 1. The wireless communication system includes a pair of transceivers, and the transceivers are synchronized according to a communication time slot (hereinafter referred to as "time slot"). The time slot number is represented as . The system uses a frequency band , which is divided into non-overlapping channels with equal intervals . In order to realize anti-jamming communication, the system dynamically switches the communication channel to avoid interference signals.
[0038] 2. The receiver has a wideband spectrum sensing and learning and decision-making capability, can sense the interference signals in the available frequency band in real time, and can feed back the channel selection instructions to the transmitter through a reliable control channel. The receiver is deployed with three function modules that can work in parallel, i.e., data communication, spectrum sensing and intelligent learning, so that the transceiver can complete data transmission, wideband spectrum sensing and learning and decision-making in parallel in each time slot, as shown in the accompanying Figure 2 .
[0039] 3. The jammer transmits dynamic nonlinear sweep jamming in the frequency band range, and changes the sweep rate and direction in each time slot with a certain probability, injects malicious interference signals into the scanned communication channel, forces the signal-to-jam-and-noise ratio (SJNR) of the receiver end of the wireless communication system to be less than the demodulation threshold, and thus blocks the communication.
[0040] In an embodiment, asFigure 1 As shown, a nonlinear sweep jamming interference communication method based on proximal policy optimization is provided, comprising the following steps:
[0041] Step 102, model the nonlinear sweep jamming interference communication process as a Markov decision process, wherein at each time step, the state includes the historical interference trajectory, the current transmission channel, the distance between the current interference channel and the transmission channel, the action includes the transmission channel selection of the next time slot, the reward value includes the product of the interference avoidance effect corresponding penalty term and the channel switching cost corresponding penalty term, and the interference avoidance effect penalty term is valued according to the distance gradient between the current interference channel and the transmission channel.
[0042] The environmental state is defined as the vector representation of the historical interference environment state and the current communication state observed by the system, and the state of the th time slot is represented as:
[0043] (1)
[0044] wherein, is the interference channel discretization value of the th time slot, , is the historical length parameter, is the transmission channel used in the current time slot, is the distance feature component between the current interference channel and the transmission channel. The historical interference trajectory component , wherein is the interference signal frequency of the th time slot, is the channel discretization function:
[0045] (2)
[0046] It is noted that when the historical record is sufficient (N , the interference frequency of the corresponding time slot is discretized according to formula (2); when the historical record is insufficient, a discrete uniform random distribution is used to complete: The distance feature component between the current interference channel and the communication channel , wherein is the communication channel selection made in the last time slot (i.e. action).
[0047] The action is defined as the next time slot transmission channel that can be selected by the system, and the action of the th time slot is represented as:
[0048] (3)
[0049] The reward function is designed to guide the system in making optimal channel selection decisions under dynamic nonlinear frequency sweeping interference environments. The reward value comprehensively considers two key factors: interference avoidance effectiveness and channel handover cost, employing a product form to achieve multi-objective balance. The system, based on the... The action selected in the previous time slot is executed in this time slot. Combined with the state of the previous time slot , can be in the first The reward for each time slot is calculated using the following formula. :
[0050] (4)
[0051] in, For channel distance penalty term, The channel handover penalty term is calculated using the following two formulas:
[0052] (5)
[0053] (6)
[0054] Among them, the transmission channel of the previous time slot and distance feature components From the state The current time slot transmission channel is directly derived from this. From the action The current time slot distance feature component is obtained. It can be based on the frequency of the sensed interference signal. The result calculated using equation (2) , and then combine Calculation .
[0055] Step 104: Construct a deep neural network model including a policy network and a value network. The policy network is used to output the selection probability of each transmission channel based on the current state, and the value network is used to evaluate the estimated value of the current state.
[0056] like Figure 3 As shown, a schematic diagram of a policy network and a value network is provided, wherein, Figure 3 (a) is a schematic diagram of the policy network structure. Figure 3 (b) is a schematic diagram of the value network structure, showing the construction of both a policy network and a value network. Both networks consist of an input layer, two hidden layers, and an output layer. The policy network input state vector... Output probability distribution , indicating state The probability of selecting each action is given. The number of nodes in the input layer of the policy network is... The number of output layer nodes is The number of hidden layer nodes is set to 128 according to the empirical rule of deep neural networks. , represents a rounding operation function. The weight matrix of the policy network has a size of and is initialized with He uniform distribution: sampled from a uniform distribution with , the corresponding bias matrix has a size of and is initialized to a small constant 0.1. has a size of and is initialized with He uniform distribution: sampled from a uniform distribution with , the corresponding bias matrix has a size of and is initialized to a small constant 0.1. has a size of and is initialized with a reduced Xavier uniform distribution: sampled from a uniform distribution with , the corresponding bias matrix has a size of and is initialized to a small negative constant -0.01. Similarly, the value network input state vector , the output state value , the number of input layer nodes is , the number of output layer nodes is 1, and the number of hidden layer nodes is set to . The weight matrix of the value network has a size of and is initialized with He uniform distribution, and the corresponding bias matrix has a size of and is initialized to a small constant 0.1. has a size of and is also initialized with He uniform distribution, and the corresponding bias matrix has a size of and is initialized to a small constant 0.1. has a size of and is initialized with Xavier uniform distribution: sampled from a uniform distribution with , the corresponding bias matrix has a size of and is initialized to 0.
[0057] Step 106, under the guidance of the current policy network, the transmitter and the receiver execute the transmission channel selection decision of the last time slot synchronously, switch to the corresponding channel for data communication, and store the interaction data between the receiver and the jammer as experience samples into the experience pool.
[0058] The interaction data includes the previous state, the action, the reward value and the current state. The previous time slot state recorded by the spectrum sensing module and the current action performed by the data communication module are calculated according to formula (4) , so as to obtain a complete set of experience samples , which are stored in the experience pool.
[0059] In step 108, experience samples are sampled from the experience pool, the parameters of the policy network and the value network are optimized using the proximal policy optimization algorithm and the experience samples, and the iterative update of the model weight is completed using the adaptive optimization algorithm.
[0060] The steps of sampling from the experience pool, calculating the value error, calculating the advantage function, calculating the ratio of the new and old policies, calculating the clipped policy loss, calculating the value function loss, and updating the network parameters by back propagation are completed by the receiver intelligent learning module to complete one update of the policy network and the value network of the proximal policy optimization algorithm.
[0061] In step 110, the updated policy network is used to select the transmission channel of the next time slot and fed back to the transmitter.
[0062] The receiver intelligent learning module inputs the current state into the policy network of the proximal policy optimization algorithm to obtain the selection probability of each channel in the next time slot , and the channel is selected according to the probability to obtain the transmission channel of the next time slot , which is fed back to the transmitter through a reliable control channel, and the transmitter and the receiver jointly perform data communication according to the action in the next time slot.
[0063] In the above anti-nonlinear sweep jamming communication method based on proximal policy optimization, by modeling the anti-nonlinear sweep jamming communication process as a Markov decision process with multiple dimensions of state, the historical jamming trajectory, the current transmission channel, and the distance between the jamming and the transmission channel are fused to comprehensively capture the dynamic change trend and real-time position correlation of the sweep jamming, with the help of the dual-network architecture of the policy network and the value network, the policy network outputs the selection probability of each channel, the value network evaluates the long-term return of the state, and the gradient assignment reward mechanism combining the interference avoidance and switching cost double targets can realize accurate channel selection decision, avoid blind switching or passive jamming, store the interaction data in the experience pool, optimize the parameters of the dual networks using the proximal policy optimization algorithm, and complete the iterative update of the weight using the adaptive optimization, which can quickly adapt to the nonlinear changes of the sweep direction and rate, and ensure the stability and real-time performance of the policy update. The embodiments of the present application can enable the wireless communication system to autonomously learn the optimal anti-jamming strategy, efficiently suppress dynamic nonlinear sweep jamming, and significantly improve the reliability of wireless communication transmission in complex electromagnetic environments.
[0064] In one embodiment, the method further comprises: if the communication is not ended, entering the next time slot and returning to the transmitter and the receiver synchronizing the transmission channel selection decision of the previous time slot to execute the iteration optimization strategy network and the value network; and if the communication is ended, stopping the iteration.
[0065] In one embodiment, the step of calculating the reward value comprises: determining a penalty item of interference avoidance effect according to the distance between the current interference channel and the transmission channel according to a preset gradient; determining a penalty item of channel switching cost according to whether the current transmission channel is consistent with the transmission channel of the previous time slot and whether the distance between the interference channel and the transmission channel after switching increases; and multiplying the penalty item of interference avoidance effect and the penalty item of channel switching cost to obtain the reward value.
[0066] In one embodiment, the step of obtaining the historical interference trajectory comprises: discretizing the interference signal frequency of the historical time slot, mapping the interference signal frequency to the corresponding channel number, and obtaining the historical interference trajectory according to the discretized values of the interference channels of each time slot; and when the number of historical time slot records is less than a preset historical length parameter, using a discrete uniform random distribution to extract values from all available channel numbers to supplement the historical interference trajectory. In this embodiment, by constructing the historical interference trajectory, the interference channel information of continuous time slots can be integrated, and the direction trend and rate change of the nonlinear frequency sweep can be deduced. This design, on the one hand, realizes more accurate capture of historical interference state, avoiding one-sidedness of single interference information; on the other hand, it can effectively predict the nonlinear change trend of the interference, providing more comprehensive interference dynamic support for the channel selection decision of the subsequent strategy network, which is conducive to the strategy network outputting channel selection probability more adaptive to the interference change characteristics, and improving the decision accuracy of anti-nonlinear frequency sweep interference.
[0067] In one embodiment, the strategy network and the value network each comprise an input layer, a double hidden layer, and an output layer; the number of input layer nodes of the strategy network is consistent with the state dimension, the number of output layer nodes is consistent with the number of available transmission channels, and the number of hidden layer nodes is the square root of the product of the number of input layer nodes and the number of output layer nodes rounded up; the number of input layer nodes of the value network is consistent with the state dimension, the number of output layer nodes is 1, and the number of hidden layer nodes is the same as that of the strategy network.
[0068] In one embodiment, optimizing the parameters of the policy network and value network using a proximal policy optimization algorithm and empirical samples includes: calculating the target value and value error based on empirical samples using the proximal policy optimization algorithm, and obtaining the advantage function based on the value error; limiting the ratio of new to old policies within a preset range to restrict the update amplitude of the policy network, and calculating the policy loss using the advantage function and the restricted ratio of new to old policies; calculating the value loss based on the target value and the estimated value of the value network; minimizing the policy loss and value loss respectively through gradient backpropagation, and adjusting the parameters of the policy network and value network; calculating the target value and value error based on empirical samples using the proximal policy optimization algorithm includes: inputting the current state from the empirical samples into the value network to obtain the estimated value of the current state, calculating the target value of the current state based on the estimated value, a preset future reward discount coefficient, and the reward value in the empirical samples; and calculating the value error by the difference between the target value of the current state and the estimated value of the previous state.
[0069] Specifically, a batch of experience samples is extracted from the experience pool P. The batch size is set to 64 in this invention.
[0070] For each sample, Input the value network and get , bring in ,in This is a discount factor reflecting the importance of future rewards; in near-term policy optimization algorithms, this hyperparameter is typically set to 0.99; then input into the value network... ,get ; Calculate the value error .
[0071] This invention employs online learning of proximal strategy optimization, thus simplifying the calculation of the advantage function. Advantage function Indicates action Better than average. Indicates action It is worse than the average level.
[0072] For each sample, Input the policy network to obtain the probability of choosing each action in this state. This serves as a benchmark for subsequent comparisons. K The optimization iterations are performed 4 to 10 times (5 in this invention), and each iteration first calculates the ratio of the old and new strategies using the following formula:
[0073] (7)
[0074] in, The parameters representing the current policy network, including all weight matrices and bias matrices, are updated in each iteration; representing the current parameters , the policy network input outputs all action probability distributions.
[0075] The clipped policy loss is calculated as follows:
[0076] (8)
[0077] where, is a hyperparameter clipping factor, which is set to 0.1 according to experience, used to limit the update amplitude of the policy; represents limiting to the interval ; represents averaging all samples in the current batch.
[0078] The value loss is calculated according to the standard mean square error:
[0079] (9)
[0080] where, represents the parameters of the current value network, including all weight matrices and bias matrices, which are updated in each iteration; representing the current parameters , the value network input outputs the estimated value; represents averaging all samples in the current batch.
[0081] The gradients of the policy loss with respect to the policy network parameters and the value loss with respect to the value network parameters are calculated by backpropagation, and then the parameters and are updated using two independent Adam optimizers:
[0082] (10)
[0083] (11)
[0084] where, and are the learning rates of the policy network and the value network, respectively, used to control the amplitude of each parameter update, which are set to 0.0001 and 0.003, respectively; the above updates of the policy network and the value network are calculated by backpropagation to calculate the gradients and use the Adam optimizer to update the network parameters.
[0085] In one embodiment, as shown in Figure 4 Fig. 1 is a flowchart of a method for anti-nonlinear sweep jamming communication based on proximal policy optimization, the method being applied to a wireless communication system including a transmitter and a receiver, the transmitter and the receiver being synchronized in time slots, the method comprising the following steps:
[0086] Step 1, modeling the anti-nonlinear sweep jamming communication problem as a Markov decision process;
[0087] Step 2, constructing a policy network and a value network of the proximal policy optimization algorithm and initializing parameters;
[0088] Step 3, the transmitter and the receiver performing the selected action in the last time slot and switching to the corresponding channel for data communication;
[0089] Step 4, the receiver observing the nonlinear sweep jamming environment, obtaining the current state and recording;
[0090] Step 5, the receiver calculating the reward corresponding to the state and action in the last time slot to obtain experience samples and storing the experience samples in an experience pool;
[0091] Step 6, the receiver updating the parameters of the policy network and the value network;
[0092] Step 7, the receiver selecting a transmission channel for the next time slot by using the policy network and transmitting the selected channel back to the transmitter;
[0093] Step 8, if the communication has not ended, entering the next time slot and returning to Step 3 until the communication ends.
[0094] It can be understood that step 1 models the communication problem as a Markov decision process by defining a multi-dimensional state including the historical interference trajectory, the current transmission channel, and the distance between the interference and the transmission channel, breaking through the limitation of traditional methods that only rely on real-time interference parameters, realizing comprehensive perception of the sweeping direction and rate change trend, and providing complete environment information for policy learning, step 2 builds a double-network architecture of the policy network and the value network, laying the foundation for dynamic policy optimization, wherein the policy network is responsible for outputting the channel selection probability, and the value network evaluates the long-term return of the state, avoiding the shortsightedness of single network decision, steps 3-5 perform channel switching by the transceiver, real-time observation of the interference environment by the receiver, calculation of the multi-objective balance reward, and storage of experience samples, so that the network can continuously obtain training data under the real interference environment, ensuring that the decision strategy is consistent with the actual scene, step 6 updates the network parameters by using the proximal policy optimization algorithm, limiting the policy update amplitude to avoid decision shock caused by interference mutation, and combining experience pool batch sampling to improve the parameter update efficiency, balancing the real-time and stability of policy optimization, steps 7-8 enable the policy network to continuously learn and optimize the channel selection strategy, so that the system can autonomously adapt to the nonlinear changes of the sweeping direction and rate, efficiently suppress the sweeping interference without modeling the interference characteristics in advance, and significantly improve the reliability of wireless communication transmission in complex electromagnetic environments.
[0095] It should be understood that, although Figure 1 The steps in the flowchart of the method of the present application are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, Figure 1 At least part of the steps in the method of the present application can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or sub-steps or stages of other steps.
[0096] In one embodiment, a nonlinear sweeping interference resistant communication system based on proximal policy optimization is provided, comprising:
[0097] A model construction module is configured to model the nonlinear sweeping interference resistant communication process as a Markov decision process, wherein at each time step, the state includes the historical interference trajectory, the current transmission channel, and the distance between the current interference channel and the transmission channel, the action includes the transmission channel selection in the next time slot, the reward value includes the product of the interference avoidance effect penalty term and the channel switching cost penalty term, and the interference avoidance effect penalty term is valued according to the distance gradient between the current interference channel and the transmission channel;
[0098] a network construction module, configured to construct a deep neural network model comprising a policy network and a value network, the policy network being configured to output selection probabilities of transmission channels according to a current state, and the value network being configured to evaluate an estimated value of the current state;
[0099] a sample acquisition module, configured to, under guidance of the current policy network, execute transmission channel selection decisions of a previous time slot by the transmitter and the receiver synchronously, switch to corresponding channels for data communication, and store interactive data of the receiver and the jammer as experience samples into an experience pool, the interactive data comprising a previous state, an action, a reward value, and a current state;
[0100] a network updating module, configured to sample experience samples from the experience pool, optimize parameters of the policy network and the value network by using a proximal policy optimization algorithm and the experience samples, and complete iterative updating of model weights by using an adaptive optimization algorithm;
[0101] a result output module, configured to select transmission channels of a next time slot by using the updated policy network and feed back to the transmitter.
[0102] Specific definitions of the anti-nonlinear frequency-sweep jamming communication system based on proximal policy optimization can be found in the definitions of the anti-nonlinear frequency-sweep jamming communication method based on proximal policy optimization in the foregoing, which will not be repeated here. The various modules in the anti-nonlinear frequency-sweep jamming communication system based on proximal policy optimization can be realized by software, hardware, or a combination thereof, in whole or in part. The various modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in the computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the various modules.
[0103] In an embodiment, a computer device is provided, which can be a terminal, and an internal structure diagram of the computer device can be as shown in Figure 5 The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with external terminals through a network connection. The computer program is executed by the processor to implement an anti-nonlinear frequency-sweep jamming communication method based on proximal policy optimization. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball, or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0104] Those skilled in the art can understand that Figure 5 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0105] In one embodiment, a computer device is provided, comprising a memory and a processor, the memory storing a computer program, and the processor implementing the steps of the method in the above embodiments when executing the computer program.
[0106] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps of the method in the above embodiments.
[0107] Those skilled in the art can understand that all or part of the processes in the above embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0108] Each technical feature of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.
[0109] The above embodiments only express several implementation ways of the present application, and the description is more specific and detailed, but it should not be understood as a limitation to the scope of the application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, several modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A communication method for resisting nonlinear frequency sweeping interference based on near-end strategy optimization, characterized in that, The method includes: The anti-nonlinear frequency sweeping interference communication process is modeled as a Markov decision process. At each time step, the state includes the historical interference trajectory, the current transmission channel, and the distance between the current interference channel and the transmission channel. The action includes the selection of the transmission channel in the next time slot. The reward value includes the product of the penalty term corresponding to the interference avoidance effect and the penalty term corresponding to the channel switching cost. The penalty term for the interference avoidance effect is assigned a value according to the distance gradient between the current interference channel and the transmission channel. A deep neural network model is constructed, comprising a policy network and a value network; the policy network is used to output the selection probability of each transmission channel based on the current state, and the value network is used to evaluate the estimated value of the current state. Under the guidance of the current policy network, the transmitter and receiver synchronously execute the transmission channel selection decision of the previous time slot, switch to the corresponding channel for data communication, and store the interaction data between the receiver and the jammer as experience samples in the experience pool; the interaction data includes the previous state, action, reward value and current state; Experience samples are sampled from the experience pool, and the parameters of the policy network and value network are optimized using the near-end policy optimization algorithm and the experience samples. An adaptive optimization algorithm is then used to iteratively update the model weights. The updated policy network is used to select the transmission channel for the next time slot and the feedback is sent back to the transmitter.
2. The method according to claim 1, characterized in that, The method further includes: If the communication has not ended, proceed to the next time slot and return to the transmitter and receiver to synchronously execute the transmission channel selection decision steps of the previous time slot, continuously iterating and optimizing the policy network and value network; If communication ends, the iteration stops.
3. The method according to claim 1, characterized in that, The steps for calculating the reward value include: Based on the distance between the current interference channel and the transmission channel, the penalty term for interference avoidance effect is determined according to a preset gradient; The channel handover cost penalty is determined based on whether the current transmission channel is consistent with the previous time slot transmission channel and whether the distance between the interference channel and the transmission channel increases after the handover. The reward value is obtained by multiplying the interference avoidance effect penalty term with the channel handover cost penalty term.
4. The method according to claim 1, characterized in that, The steps for obtaining the historical interference trajectory include: The interference signal frequency of the historical time slot is discretized and mapped to the corresponding channel number. The historical interference trajectory is obtained based on the discretized value of the interference channel of each time slot. When the number of historical time slot records is insufficient to meet the preset historical length parameter, a discrete uniform random distribution is used to extract values from all available channel numbers to supplement the historical interference trajectory.
5. The method according to claim 1, characterized in that, Both the policy network and the value network include an input layer, two hidden layers, and an output layer. The number of input layer nodes in the policy network is consistent with the state dimension, the number of output layer nodes is consistent with the number of available transmission channels, and the number of hidden layer nodes is the square root of the product of the number of input layer nodes and the number of output layer nodes, rounded to the nearest whole number. The number of nodes in the input layer of the value network is the same as the number of nodes in the state dimension, the number of nodes in the output layer is 1, and the number of nodes in the hidden layer is the same as that in the policy network.
6. The method according to claim 1, characterized in that, The parameters of the policy network and value network optimized using the proximal policy optimization algorithm and the empirical samples include: The near-end strategy optimization algorithm is used to calculate the target value and value error based on the empirical samples, and the advantage function is obtained based on the value error. The ratio of new to old policies is limited to a preset range to restrict the update range of the policy network. The policy loss is calculated using the advantage function and the limited ratio of new to old policies. Calculate value loss based on the estimated value of the target value and the value network; The policy loss and value loss are minimized respectively through gradient backpropagation, and the parameters of the policy network and value network are adjusted.
7. The method according to claim 6, characterized in that, The near-end strategy optimization algorithm calculates the target value and value error based on the empirical samples, including: Input the current state from the experience sample into the value network to obtain the estimated value of the current state. Calculate the target value of the current state based on the estimated value, the preset future reward discount coefficient, and the reward value in the experience sample. The value error is calculated by the difference between the target value of the current state and the estimated value of the previous state.
8. A communication system for resisting nonlinear frequency sweeping interference based on near-end strategy optimization, characterized in that, The system includes: The model building module is used to model the anti-nonlinear frequency sweeping interference communication process as a Markov decision process. At each time step, the state includes the historical interference trajectory, the current transmission channel, and the distance between the current interference channel and the transmission channel. The action includes the selection of the transmission channel in the next time slot. The reward value includes the product of the penalty term corresponding to the interference avoidance effect and the penalty term corresponding to the channel switching cost. The penalty term for the interference avoidance effect is assigned a value according to the distance gradient between the current interference channel and the transmission channel. A network construction module is used to construct a deep neural network model including a policy network and a value network; the policy network is used to output the selection probability of each transmission channel according to the current state, and the value network is used to evaluate the estimated value of the current state. The sample acquisition module is used to enable the transmitter and receiver to synchronously execute the transmission channel selection decision of the previous time slot under the guidance of the current policy network, switch to the corresponding channel for data communication, and store the interaction data between the receiver and the jammer as experience samples in the experience pool; the interaction data includes the previous state, action, reward value and current state. The network update module is used to sample experience samples from the experience pool, optimize the parameters of the policy network and the value network using the near-end policy optimization algorithm and the experience samples, and use an adaptive optimization algorithm to complete the iterative update of the model weights. The result output module is used to select the transmission channel for the next time slot using the updated policy network and feed it back to the transmitter.
Citation Information
Patent Citations
Method for dispersing bird behavior migration in humanization airport based on anti-adaptive continuous ultrasonic neural interference
CN121128703A
Military 5G user equipment
US12425120B1