Wireless communication anti-interference strategy optimization method and system based on reinforcement learning

By using a reinforcement learning-based method to optimize anti-interference strategies in wireless communication, and by employing an anti-interference decision model that cascades a representation learning model and a policy mapping model, anti-interference strategies can be acquired and optimized in real time. This solves the problem of adaptability in dynamic interference environments in wireless communication and improves the anti-interference effect.

CN121841544APending Publication Date: 2026-04-10CHANGSHU INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing wireless communication anti-interference technologies lack adaptability and intelligent decision-making capabilities in dynamic and complex interference environments, making it difficult to adjust anti-interference strategies in real time, resulting in limited anti-interference effectiveness.

Method used

By employing a reinforcement learning-based approach, the system acquires real-time communication environment status information and utilizes an anti-interference decision model composed of a cascaded representation learning model and a policy mapping model to output the optimal anti-interference action at the current moment. The action is then executed by an anti-interference action executor, and response information is collected synchronously for collaborative model updates, thereby achieving adaptive optimization.

Benefits of technology

Continuous learning and optimization in a dynamically changing interference environment improves the anti-interference effect of wireless communication systems, enhances their adaptability and intelligent decision-making capabilities, and improves communication quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121841544A_ABST
    Figure CN121841544A_ABST
Patent Text Reader

Abstract

The invention discloses a wireless communication anti-interference strategy optimization method and system based on reinforcement learning, and relates to the technical field of wireless communication, and the method comprises the steps: obtaining the communication environment state information of a target scene in real time; inputting the communication environment state information into a pre-trained anti-interference decision model, and obtaining an anti-interference action at the current moment; and executing an anti-interference action through an anti-interference action actuator, synchronously collecting anti-interference action response information, and carrying out collaborative updating on the anti-interference decision model according to the anti-interference action response information. The technical problems that an existing wireless communication anti-interference technology lacks self-adaptability and intelligent decision-making capacity to a dynamic complex interference environment, an anti-interference strategy is difficult to adjust in real time, and the anti-interference effect is limited are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of wireless communication technology, specifically to a method and system for optimizing wireless communication anti-interference strategies based on reinforcement learning. Background Technology

[0002] With the rapid development of wireless communication technology, its application scenarios are becoming increasingly complex and diverse, ranging from traditional voice communication to data transmission and Internet of Things (IoT) connections, which places higher demands on communication quality and reliability.

[0003] However, wireless communication environments typically face various forms of interference, affecting the performance of communication systems and leading to degraded signal transmission quality, increased bit error rate, and reduced throughput. Traditional anti-interference techniques often rely on fixed rules or preset interference models, lacking adaptability and intelligent decision-making capabilities in dynamic and complex interference environments. They are also difficult to adjust anti-interference strategies in real time, resulting in limited anti-interference effectiveness. Summary of the Invention

[0004] This application provides a method and system for optimizing anti-interference strategies for wireless communication based on reinforcement learning. This solves the technical problem that existing anti-interference technologies lack adaptability and intelligent decision-making capabilities in dynamic and complex interference environments, making it difficult to adjust anti-interference strategies in real time and resulting in limited anti-interference effects.

[0005] The technical solution to the above-mentioned technical problems in this application is as follows: In a first aspect, this application provides a method for optimizing anti-interference strategies in wireless communication based on reinforcement learning, the method comprising: Real-time acquisition of communication environment status information of the target scene, wherein the communication environment status information is high-dimensional status data; The communication environment state information is input into a pre-trained anti-interference decision model to obtain the anti-interference action at the current moment. The anti-interference decision model is constructed by cascading a representation learning model and a policy mapping model. The anti-interference action is executed by the anti-interference action actuator, the anti-interference action response information is collected synchronously, and the anti-interference decision model is updated collaboratively based on the anti-interference action response information.

[0006] Secondly, this application provides a wireless communication anti-interference strategy optimization system based on reinforcement learning, including: The information acquisition module is used to acquire the communication environment status information of the target scene in real time, wherein the communication environment status information is high-dimensional status data; The model training module is used to input the communication environment state information into the pre-trained anti-interference decision model to obtain the anti-interference action at the current moment. The anti-interference decision model is composed of a representation learning model and a policy mapping model cascaded together. The model update module is used to execute the anti-interference action through the anti-interference action executor, synchronously collect the anti-interference action response information, and collaboratively update the anti-interference decision model based on the anti-interference action response information.

[0007] This application provides one or more technical solutions, which have at least the following technical effects or advantages: This application provides a method and system for optimizing anti-interference strategies in wireless communication based on reinforcement learning. First, it acquires real-time communication environment state information of the target scenario, providing an environmental awareness foundation for anti-interference decision-making. Second, it inputs high-dimensional state data into a pre-trained anti-interference decision-making model composed of a representation learning model and a policy mapping model, outputting the optimal anti-interference action at the current moment, thus achieving intelligent mapping from environmental awareness to decision output. Third, it executes the acquired anti-interference action through an anti-interference action executor, simultaneously collecting anti-interference action response information, including feedback rewards and communication environment state information at the next moment. During the update process, the policy mapping model calculates the policy gradient based on reinforcement learning methods and empirical data in the response information, updating its network parameters with the goal of maximizing the expected cumulative reward in the future. The representation learning model not only considers its own reconstruction loss but also receives the policy gradient from the backpropagation of the policy mapping model, updating the encoder parameters through weighted calculation of the joint loss, making representation learning more aligned with the needs of the anti-interference decision-making task. The collaborative update mechanism ensures that the anti-interference decision model can continuously learn and optimize in a dynamically changing interference environment, thereby continuously improving its adaptability and intelligent decision-making capabilities, effectively coping with complex and ever-changing interference situations, and improving the anti-interference effect of wireless communication systems.

[0008] Through the above technical solution, this application utilizes a pre-trained anti-interference decision model, composed of a cascaded representation learning model and a policy mapping model, based on the high-dimensional state information of the sensing communication environment, to output the current optimal anti-interference action. After executing the anti-interference action, response information is actively collected, and the model is collaboratively updated, achieving deep coupling between feature extraction and decision optimization. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a flowchart illustrating the wireless communication anti-interference strategy optimization method based on reinforcement learning provided in an embodiment of this application. Figure 2 This is a schematic diagram of the structure of the wireless communication anti-interference strategy optimization system based on reinforcement learning provided in the embodiments of this application.

[0011] The components represented by each number in the attached diagram are explained below: Information acquisition module 11, model training module 12, model update module 13. Detailed Implementation

[0012] This application provides a method and system for optimizing anti-interference strategies for wireless communication based on reinforcement learning. This method addresses the technical problem that existing anti-interference technologies for wireless communication lack adaptability and intelligent decision-making capabilities in dynamic and complex interference environments, making it difficult to adjust anti-interference strategies in real time and resulting in limited anti-interference effects.

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0014] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0015] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid unnecessarily obscuring the description of this application. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0016] Example 1, as Figure 1 As shown in the embodiments of this application, a method for optimizing wireless communication anti-interference strategies based on reinforcement learning is provided, including: S10: Real-time acquisition of communication environment status information of the target scene, wherein the communication environment status information is high-dimensional status data; The communication environment status information includes at least the following: Channel quality state dimension, which is used to characterize the channel fading characteristics of the target communication link; Interference situation dimension, which is used to characterize the intensity distribution and variation characteristics of external interference signals; The communication resource occupancy dimension is used to characterize the current usage status of communication spectrum and power resources. The communication performance feedback dimension is used to characterize at least one of the communication error rate, throughput, and latency status.

[0017] In this embodiment of the application, firstly, multiple sensors and monitoring devices deployed in the target scene are used to collect multi-dimensional and real-time data on the communication environment.

[0018] In terms of channel quality status, the large-scale and small-scale fading characteristics of the target communication link are captured by receiving parameters such as signal strength indicators and channel status information.

[0019] In terms of interference situation, a spectrum analyzer is used to scan the interference signals within the working frequency band to obtain the intensity distribution information such as the center frequency, bandwidth, and power spectral density of the interference signals. By continuously monitoring and analyzing the characteristics of their changes over time, the type of interference can be determined, such as narrowband interference, broadband interference, and impulse interference.

[0020] In terms of communication resource occupancy, the network controller collects real-time data on the current communication system's spectrum resource occupancy and power resource allocation and usage status, such as transmit power, receive power, and power margin.

[0021] In terms of communication performance feedback, performance indicators such as communication error rate, real-time throughput, and data transmission latency are obtained from communication terminals or base stations as a basis for measuring the current communication quality.

[0022] The aforementioned multi-dimensional, high-dimensional state data collectively constitute a comprehensive perception of the target scene's communication environment, providing input information for the anti-interference decision-making model.

[0023] Specifically, step S10 in the method includes: The raw wireless signal containing the target frequency band is collected by deploying a radio frequency front-end and analog-to-digital converter at the communication receiver. The original wireless signal is preprocessed, wherein the preprocessing includes at least down-conversion, filtering and synchronization operations; Statistical analysis and dimensional filtering are performed on the preprocessing results to obtain multi-dimensional key state information related to anti-interference decision-making, and the communication environment state information in the form of multi-dimensional feature vectors is generated.

[0024] In this embodiment of the application, firstly, at the communication receiving end, the radio frequency front end amplifies, filters and down-converts the received wireless signal to convert the high-frequency signal into an intermediate frequency or baseband signal, and then converts the analog signal into a digital signal through an analog-to-digital converter to complete the acquisition of the original wireless signal.

[0025] Secondly, the acquired digital signals are preprocessed. Down-conversion further adjusts the signals to a frequency band suitable for subsequent processing, while filtering removes out-of-band noise and interference. Synchronization operations include carrier synchronization, bit synchronization, and frame synchronization to ensure the accuracy of the signals in time and frequency.

[0026] Secondly, after preprocessing, statistical analysis is performed on the signal, such as calculating the power spectral density, signal-to-noise ratio, and characteristic parameters of the interference signal. At the same time, dimensional filtering is performed in combination with the needs of anti-interference decision-making to remove redundant or irrelevant information and retain state information that affects the decision, such as channel fading coefficient, main parameters of the interference signal, spectrum resource occupancy rate, bit error rate, etc. The multi-dimensional key state information is integrated into a multi-dimensional feature vector, which is used as the communication environment state information input into the subsequent anti-interference decision-making model.

[0027] S20: Input the communication environment state information into the pre-trained anti-interference decision model to obtain the anti-interference action at the current moment, wherein the anti-interference decision model is constructed by cascading a representation learning model and a policy mapping model; In this embodiment, the communication environment state information collected and processed in the above steps is input into a pre-trained anti-interference decision model to obtain the anti-interference action at the current moment. The anti-interference decision model is composed of a cascaded representation learning model and a policy mapping model.

[0028] Specifically, the representation learning model, as the front-end processing unit of the anti-interference decision-making model, maps the original, potentially redundant and noisy, high-dimensional state data into a low-dimensional, compact, and highly representative latent space feature vector by reducing the dimensionality and extracting features from the input high-dimensional communication environment state information. During the pre-training phase, the representation learning model utilizes a large amount of historically collected communication environment state data for unsupervised learning.

[0029] The policy mapping model receives latent space feature vectors from the representation learning model and outputs the anti-interference action for the current time step based on these vectors. The policy mapping model employs a policy network structure from deep reinforcement learning. The input layer receives the latent space feature vectors, and through nonlinear transformations in intermediate hidden layers, the output layer corresponds to the probability distribution or specific action value across all possible anti-interference action spaces. During the pre-training phase, the policy mapping model utilizes historical experience data with reward signals for supervised or reinforcement learning pre-training to learn the initial mapping relationship from environmental features to anti-interference actions.

[0030] The training steps of the anti-interference decision model include: Construct and train the representation learning model, which is used to encode the input high-dimensional state data into a low-dimensional feature vector; Construct and train a policy mapping model based on reinforcement learning, wherein the policy mapping model takes the low-dimensional feature vector as input and outputs the anti-interference action; The policy mapping model and the representation learning model, after cascading training, generate the anti-interference decision model; The representation learning model includes an encoder and decoder with an asymmetric configuration.

[0031] In this embodiment, firstly, a representation learning model is constructed, employing an asymmetric encoder and decoder structure. The encoder uses a deep neural network, with the number of layers and neurons determined based on the dimensionality of the high-dimensional state data. The decoder uses a relatively shallow network structure, taking the low-dimensional feature vector output by the encoder as input. The number of neurons in the output layer matches the dimensionality of the original high-dimensional state data. The decoder reconstructs the low-dimensional feature vector, and the reconstruction error, such as the mean squared error, is used as the training loss of the representation learning model. During training, a large amount of historically collected unlabeled communication environment state data is used to optimize the network parameters of the encoder and decoder through backpropagation. This enables the encoder to learn the key structural features contained in the data, achieving data dimensionality reduction while retaining the information required for anti-interference decision-making.

[0032] For example, for an input containing hundreds of dimensions of state data, the encoder sets 3-5 hidden layers, with the number of neurons in each layer gradually decreasing. The ReLU activation function is used to achieve a non-linear transformation, mapping the high-dimensional input to a low-dimensional latent space feature vector.

[0033] Secondly, a policy mapping model based on reinforcement learning is constructed. The policy mapping model adopts a deep deterministic policy gradient reinforcement learning algorithm framework, with the policy network at its core. The input layer of the policy network receives the low-dimensional feature vector output by the representation learning model, the intermediate hidden layers perform feature processing through fully connected or convolutional operations, and the output layer determines the output format according to the type of anti-interference action space.

[0034] Specifically, if the anti-interference action space is discrete, such as frequency hopping selection and modulation mode switching, the output layer uses the softmax activation function to output the probability distribution of each action; if it is continuous, such as transmit power adjustment and bandwidth allocation, the output layer uses the tanh activation function to output the specific value of the action. During the pre-training phase, historical experience data with reward signals—that is, trajectory data of state-action-reward-next state—is used to train the policy mapping model through reinforcement learning algorithms. This enables the policy network to output anti-interference actions that maximize cumulative rewards based on the input low-dimensional feature vector.

[0035] Finally, the trained representation learning model and the policy mapping model are concatenated to form a complete anti-interference decision model. During the concatenation process, the encoder part of the representation learning model acts as a pre-processing module for the policy mapping model. It converts the real-time high-dimensional communication environment state information into low-dimensional feature vectors, which are then directly input into the policy network of the policy mapping model. The policy network then outputs the anti-interference action for the current moment.

[0036] Furthermore, before model deployment, the cascaded anti-interference decision models are jointly fine-tuned by using a small amount of labeled data or online interactive data to adjust model parameters, ensuring that the features extracted by the representation learning model are highly matched with the decision requirements of the policy mapping model, thereby improving the overall decision performance of the model.

[0037] Specifically, constructing and training the representation learning model includes: Raw state data from historical communication scenarios are collected to form a pre-training dataset, and the pre-training dataset is organized into batch training data according to the dimensional configuration of the communication environment state information. Construct an original representation learning model with an encoder and a decoder, and initialize the network parameters of the original representation learning model; The original representation learning model is iteratively trained with the goal of minimizing the reconstruction loss between the output data of the decoder and the original state data input to the encoder. When the reconstruction loss is lower than a preset threshold, training is completed and the original representation learning model is output as the representation learning model.

[0038] In this embodiment, firstly, raw state data under different interference types, channel conditions, and communication loads are collected from historical communication scenarios using the same data acquisition and preprocessing method as in step S10, to construct a pre-training dataset. Based on the dimensions of channel quality, interference situation, resource usage, and performance feedback included in the communication environment state information, each sample data point in the pre-training dataset is organized into a structured multidimensional array and divided into multiple training batches according to a preset batch size to accommodate the batch gradient descent requirements during model training.

[0039] Secondly, an original representation learning model is constructed, which consists of an encoder and a decoder. The encoder adopts a deep neural network structure, such as a multilayer perceptron, with the number of neurons in its input layer matching the dimension of the high-dimensional state data. It compresses the high-dimensional input into a low-dimensional latent space through layer-by-layer nonlinear transformation. The decoder adopts a neural network structure, but its network size is smaller than that of the encoder. Its input is the low-dimensional feature vector output by the encoder, and the number of neurons in its output layer matches the dimension of the original high-dimensional state data. It is used to reconstruct the low-dimensional feature vector into a vector in the original data space.

[0040] During the model initialization phase, Xavier initialization is used to randomly initialize the network weights and biases of the encoder and decoder to ensure the stability and convergence speed of model training.

[0041] Then, the original representation learning model is iteratively trained with the training objective of minimizing the reconstruction loss between the decoder output data and the encoder input original state data. In each training batch, the batch of original state data is input into the encoder to obtain low-dimensional feature vectors; then the low-dimensional feature vectors are input into the decoder to obtain reconstructed data; the mean squared error (MSE) between the reconstructed data and the original state data is calculated as the reconstruction loss.

[0042] Furthermore, the network parameters of the encoder and decoder are updated using the backpropagation algorithm and gradient descent optimizer, such as the Adam optimizer, to gradually reduce the reconstruction loss. During training, an early stopping mechanism is implemented. When the reduction in reconstruction loss is less than a preset threshold or the reconstruction loss reaches a preset target value within several consecutive training epochs, training is stopped. The original representation learning model at this point serves as the final representation learning model, and its encoder part will be used for subsequent cascading with the policy mapping model.

[0043] For example, for an input containing 1000-dimensional state data, the encoder can have four hidden layers with 512, 256, 128, and 64 neurons respectively. A ReLU activation function is used to perform a non-linear transformation, mapping the 1000-dimensional input to a 64-dimensional latent space feature vector. The decoder has two hidden layers with 128 and 256 neurons respectively. The output layer is 1000-dimensional, also using the ReLU activation function, and finally outputs the reconstructed data through a linear activation function. During training, the batch size is set to 64, the learning rate is initialized to 0.001, and the Adam optimizer is used. Training stops when the decrease in reconstruction loss is less than 1e-5 for 10 consecutive training epochs.

[0044] Furthermore, a policy mapping model based on reinforcement learning is constructed and trained, including: Construct a reinforcement learning-based original policy network and initialize the network parameters, while simultaneously defining the reward calculation function; During the interaction between the wireless communication system and the environment in the target scenario, sequential experience data consisting of states, actions, rewards, and new states is collected. The reward function is configured with the goal of maximizing the expected future cumulative reward, and the original policy network and the reward calculation function are iteratively optimized and trained by combining the sequence experience data and reinforcement learning methods. When the reward function converges, training is complete and the policy mapping model consisting of the trained original policy network is output.

[0045] In this embodiment, firstly, a primary policy network based on reinforcement learning is constructed. This network adopts a deep neural network structure, specifically in the form of alternating fully connected layers and batch normalization layers.

[0046] Specifically, the number of neurons in the input layer is consistent with the dimension of the low-dimensional feature vector output by the representation learning model. For example, when the low-dimensional feature vector is 64-dimensional, the input layer has 64 neurons. The intermediate hidden layers have 3-5 layers, and the number of neurons in each layer is determined according to the complexity of the action space, such as 256, 128, and 64 neurons. Each hidden layer is connected to a batch normalization layer to accelerate training convergence and alleviate overfitting. The activation function is LeakyReLU to enhance nonlinear expression. The number of neurons in the output layer corresponds to the dimension of the anti-interference action space. If it is a discrete action space, such as 8 frequency hopping options, the output layer has 8 neurons and uses the softmax activation function to output the probability distribution. If it is a continuous action space, such as the transmission power is continuously adjusted in the range of 0-30dBm, the output layer has 1 neuron and uses the tanh activation function to map the output to the [-1,1] interval before scaling it to the actual power range through a linear transformation.

[0047] Simultaneously, the network parameters are initialized, with weights initialized using the He normal distribution and biases initialized to 0.01. A target network is configured for the policy network, with the structure of the target network being consistent with the original policy network, and is used to stabilize the calculation of the target Q value during the reinforcement learning training process.

[0048] Furthermore, a reward calculation function is defined synchronously. This function takes the indicators of communication performance feedback dimension as the core and combines the changes in interference situation and resource utilization efficiency to construct a multi-factor weighted reward mechanism.

[0049] Specifically, the reward value is obtained by weighted summation of three parts: basic reward, interference suppression reward, and resource optimization reward. The basic reward is related to the communication bit error rate, throughput, and latency. For example, a +20 reward is given when the bit error rate is below a preset threshold; a deduction of 5 reward is made for every 1% decrease in throughput; and a deduction of 3 reward is made for every 10ms increase in latency. The interference suppression reward is calculated based on the change in the power spectral density of the interference signal. A +15 reward is given if the interference power decreases by more than 20%, and an additional +10 reward is given if the interference type changes from pulse interference to narrowband interference. The resource optimization reward is related to power margin; a +8 reward is given when the transmit power is in the 30%-70% power margin range to avoid excessive waste or insufficient power resources.

[0050] For example, the mathematical expression of the reward function is: R = α × basic reward + β × interference suppression reward + γ × resource optimization reward, where α, β, and γ are weight coefficients, and the optimal values ​​are determined by grid search according to the actual scenario requirements, such as α = 0.5, β = 0.3, and γ = 0.2.

[0051] Secondly, during the interaction between the wireless communication system and the environment in the target scenario, a semi-physical simulation platform is built to collect sequential experience data. This platform uses software-defined radio equipment to simulate the communication transmitter and receiver, generating controllable narrowband, broadband, and pulse interference types through programmable interference sources, and constructing a multi-user communication scenario using a network simulator. At each time step, such as 100ms, the current communication environment state information is first collected by sensors and processed into a low-dimensional feature vector by a representation learning model. The original policy network outputs anti-interference actions based on this and applies them to the communication system, such as switching to a specified frequency hopping or adjusting the transmission power. Subsequently, the new communication environment state and communication performance indicators after the action are executed are collected, and the reward value at the current moment is calculated. Finally, the quadruple of "current state feature vector - anti-interference action - reward value - new state feature vector" is stored in an experience replay buffer with a capacity of 1 million entries. A priority experience replay mechanism is used to dynamically adjust the sample sampling probability according to the absolute value of the reward value.

[0052] Furthermore, after configuring the reward function with the goal of maximizing the expected future cumulative reward, the original policy network and reward calculation function are iteratively optimized and trained by combining sequential experience data with deep deterministic policy gradient reinforcement learning methods.

[0053] Specifically, at fixed intervals, a batch of samples is randomly sampled from the experience replay buffer. For example, 256 samples are randomly collected every 200 steps. The current state feature vector is input into the original policy network to obtain the action prediction value, and the new state feature vector is input into the target policy network to obtain the target action. The target Q value, i.e., the expected value of the future cumulative reward, is calculated through the target Q network. The formula for calculating the target Q value is Qt=r+γ×Q'(s',a'|θ^-), where r is the current reward, γ is the discount factor set to 0.95, Q' is the target Q network, θ^- is the target network parameter, s' is the new state feature vector, and a' is the action output by the target policy network.

[0054] Furthermore, the original Q-network calculates the current Q-value based on the current state and action, and updates the parameters of the original Q-network by minimizing the mean squared error loss function between the current Q-value and the target Q-value. Then, the parameters of the original policy network are updated using the policy gradient ascent method, where the policy gradient is obtained by taking the expectation of the gradient of the Q-value with respect to the policy network parameters. Every 1000 training steps, the parameters of the original policy network and Q-network are softly updated to the target network. During training, the change in the average cumulative reward value is monitored in real time. Each epoch contains 10,000 steps. When the fluctuation range of the average cumulative reward value is less than 5% and tends to stabilize over 50 consecutive training epochs, the reward function value is considered to have converged, and training is stopped. The trained original policy network is then solidified into a policy mapping model, and its network parameters and structure serve as the core decision-making unit of the anti-interference decision model.

[0055] S30: The anti-interference action is executed by the anti-interference action actuator, the anti-interference action response information is collected synchronously, and the anti-interference decision model is updated collaboratively based on the anti-interference action response information.

[0056] In this embodiment, firstly, the anti-interference action instructions output by the trained policy mapping model are transmitted to the anti-interference action executor. The anti-interference action executor transforms the abstract action instructions into specific hardware operations or protocol configurations. Simultaneously, anti-interference action response information is collected through a multi-dimensional data acquisition device. Furthermore, based on the collected anti-interference action response information, the anti-interference decision model is collaboratively updated.

[0057] The anti-interference action is executed by an anti-interference action actuator, and the anti-interference action response information is collected synchronously, including: The anti-interference action is executed periodically with a preset decision time slot. The feedback reward generated after the anti-interference action is executed is collected, wherein the feedback reward is obtained based on the reward calculation function associated with the policy mapping model; Obtain the communication environment status information at the next moment after the anti-interference action is executed; The current communication environment status information, the anti-interference action, the feedback reward, and the next communication environment status information are structured into a single piece of empirical data and stored. Multiple pieces of empirical data constitute the anti-interference action response information.

[0058] In this embodiment, anti-interference actions are first executed periodically using preset decision time slots. The length of the decision time slot is determined based on the dynamic characteristics of the communication system. For example, in a fast-time-varying interference environment, the decision time slot is set to 50ms to achieve a rapid response to interference changes; in a slow-time-varying environment, the decision time slot can be extended to 200ms to reduce unnecessary computational resource consumption. At the beginning of each decision time slot, the anti-interference decision model receives real-time communication environment state information, encodes it into a low-dimensional feature vector through a representation learning model, and then outputs an anti-interference action command through a policy mapping model. This command is sent to the anti-interference action executor through the control interface of the communication system.

[0059] Secondly, the feedback rewards generated after executing anti-interference actions are collected. The calculation of feedback rewards follows the reward calculation function defined during the training phase of the policy mapping model, which is based on a quantitative evaluation of changes in communication performance indicators, the degree of mitigation of interference, and resource utilization efficiency after the action is executed.

[0060] For example, if the anti-interference action selects a new frequency hopping frequency, after the actuator completes the frequency switching, the data acquisition device immediately measures the bit error rate, throughput, and time delay after the switching, and calculates the basic reward by comparing it with the baseline value before the switching; at the same time, the power spectral density of the interference signal is obtained through the spectrum monitoring module. If the interference power at the new frequency point is reduced by 30% compared with the original frequency, the corresponding score is assigned according to the interference suppression reward rule.

[0061] Next, the communication environment state information at the next moment after the anti-interference action is executed is obtained. The state information acquisition at the next moment adopts the same dimensional configuration and preprocessing process as the current moment's state acquisition, including channel quality parameters, interference situation parameters, resource occupancy parameters, and performance feedback parameters, such as the current bit error rate and real-time throughput. After filtering, normalization, and other preprocessing, the acquired raw data is organized into a structured multidimensional array with the same format as the pre-training dataset, which serves as the input state for the next decision slot.

[0062] Finally, the current communication environment state information, anti-interference actions, feedback rewards, and the next communication environment state information are structurally integrated into a complete set of empirical data, formatted as a quadruple (s, a, r, s'), where s is the current state feature vector, a is the executed anti-interference action, r is the feedback reward, and s' is the next state feature vector. This empirical data is stored in real-time in an experience replay buffer, which uses a first-in, first-out (FIFO) strategy to manage the data. When the storage capacity reaches a preset limit, such as 1 million records, the oldest empirical data is automatically discarded. Simultaneously, to improve the efficiency of subsequent model updates, newly added empirical data is prioritized based on the absolute value of its reward. Samples with larger absolute reward values, i.e., those potentially contributing more to policy optimization, are assigned higher sampling priority.

[0063] Furthermore, the anti-interference decision model is collaboratively updated based on the anti-interference action response information, including: When the amount of experience data stored in the anti-interference action response information reaches a preset activation threshold, a collaborative update is activated. Randomly select small batches of data samples from the stored multiple sets of empirical data; The small batch of data samples are transmitted to the representation learning model and the policy mapping model in the anti-interference decision model respectively, and the parameters are updated collaboratively.

[0064] In this embodiment, firstly, a threshold is set for the number of experience data entries stored in the anti-interference action response information storage. This threshold is determined based on a combination of the capacity of the experience replay buffer and the model update efficiency. For example, when 5,000 new experience data entries are added to the buffer, a collaborative update process is triggered to ensure that enough new samples participate in model optimization.

[0065] Secondly, a priority sampling strategy is used to randomly extract small batches of data samples from the stored experience data, such as 256 samples each time. The sampling probability of high-priority samples is three times that of low-priority samples, so as to improve the utilization rate of key experience.

[0066] Next, the extracted small batches of data samples are transmitted to the representation learning model and the policy mapping model respectively for collaborative parameter updates. For the representation learning model, the current state feature vector s in the sample and the original state data are re-input into the encoder. The original state data is the original high-dimensional state data corresponding to the low-dimensional vector recovered from s through the inverse encoding process.

[0067] Furthermore, the reconstruction loss is calculated, and the encoder parameters are fine-tuned using gradient descent to better adapt the model to the state feature distribution in the new environment. For the policy mapping model, the same deep deterministic policy gradient method as in the training phase is used, updating the Q-network and policy network parameters using the (s,a,r,s') quadruplets in the samples. After every 10 mini-batch updates, a soft update of the target network is performed, with the update coefficient set to 0.01 to balance model stability and learning speed. Through the collaborative update mechanism of the representation learning model and the policy mapping model, the anti-interference decision model achieves continuous adaptive optimization to the dynamic communication environment.

[0068] Specifically, transmitting the small batch of data samples to the representation learning model and the policy mapping model in the anti-interference decision model for collaborative parameter updates includes: The policy mapping model calculates the policy gradient based on the communication environment state information in the mini-batch data samples, the anti-interference action, the feedback reward, and the communication environment state information at the next moment, combined with reinforcement learning methods. With the goal of maximizing the expected future cumulative reward, the network parameters of the policy mapping model are updated according to the policy gradient; The representation learning model calculates the reconstruction loss by forward propagation based on the communication environment state information in the mini-batch data samples, and receives the policy gradient by backpropagation from the policy mapping model. The joint loss is calculated by weighting the reconstruction loss and the policy gradient, and the network parameters of the encoder in the representation learning model are updated by backpropagation based on the joint loss. Based on the updated anti-interference decision model, the anti-interference strategy optimization is performed in the next cycle.

[0069] In this embodiment, the policy mapping model first uses the (s,a,r,s') quadruple from a mini-batch data sample to calculate the policy gradient using a deep deterministic policy gradient algorithm. Specifically, the current state feature vector s is input into the original policy network to obtain the probability distribution or specific value of action a. Combined with the Q-value estimation of (s,a) from the original Q-network, the policy gradient is obtained by taking the expectation of the gradient of the Q-value with respect to the policy network parameters. This gradient direction points to the parameter update direction that can increase the expected cumulative reward in the future.

[0070] Secondly, with the goal of maximizing the expected cumulative reward in the future, the policy mapping model updates its network parameters, including the weights and biases of each hidden layer, along the calculated policy gradient direction using stochastic gradient ascent. The learning rate is set to 0.001, and the parameters are iterated using the Adam optimizer to balance convergence speed and stability.

[0071] Simultaneously, the representation learning model receives raw communication environment state information from small batches of data samples, such as the original channel impulse response, interference power spectrum data, and their corresponding low-dimensional state feature vectors. The encoder of the representation learning model forward propagates the raw state information to generate new low-dimensional feature vectors. By calculating the mean square error between the new low-dimensional feature vector and the original low-dimensional feature vector in the sample, the reconstruction loss is obtained, which reflects the encoder's ability to compress and reconstruct state information.

[0072] Furthermore, the representation learning model not only updates its parameters based on its own reconstruction loss, but also receives the policy gradient from the backpropagation of the policy mapping model during the update process. The policy gradient reflects the impact of the current state representation on the quality of policy decisions. The representation learning model constructs a joint loss function by weighted summation of the reconstruction loss and the policy gradient. The policy gradient is obtained by differentiating the policy gradient with respect to the encoder parameters. The weight coefficient λ is used to balance reconstruction accuracy and policy relevance; for example, setting λ=0.3 ensures that the policy gradient contributes 30% to the joint loss.

[0073] Subsequently, the representation learning model updates the encoder's network parameters using the backpropagation algorithm based on the joint loss function, so that the low-dimensional state feature vector generated by the encoder can not only reconstruct the original state information, but also more effectively support the policy mapping model to make optimal anti-interference decisions.

[0074] Finally, after completing the collaborative parameter update of the representation learning model and the policy mapping model, the updated anti-interference decision model will be used to execute the anti-interference policy optimization process in the next decision slot, including state acquisition, feature encoding, action decision, action execution and response information acquisition, forming a continuous closed-loop learning and optimization process to adapt to the dynamic changes in the wireless communication environment and continuously improve anti-interference performance.

[0075] In summary, compared with existing technologies, this application achieves continuous adaptive optimization of the anti-interference decision model for dynamic communication environments by deeply integrating reinforcement learning with deep deterministic policy gradient methods and introducing a collaborative update mechanism between representation learning models and policy mapping models.

[0076] In summary, the embodiments of this application have at least the following technical effects: This application provides a reinforcement learning-based method for optimizing anti-interference strategies in wireless communication. First, it acquires real-time communication environment state information of the target scenario, providing an environmental awareness foundation for anti-interference decision-making. Second, it inputs high-dimensional state data into a pre-trained anti-interference decision-making model composed of a representation learning model and a policy mapping model, outputting the optimal anti-interference action for the current moment, thus achieving intelligent mapping from environmental awareness to decision output. Third, it executes the acquired anti-interference action through an anti-interference action executor, simultaneously collecting anti-interference action response information, including feedback rewards and communication environment state information for the next moment. During the update process, the policy mapping model calculates the policy gradient based on reinforcement learning methods and empirical data in the response information, updating its network parameters with the goal of maximizing the expected cumulative reward in the future. The representation learning model not only considers its own reconstruction loss but also receives the policy gradient from the backpropagation of the policy mapping model, updating the encoder parameters through weighted calculation of the joint loss, making representation learning more aligned with the needs of the anti-interference decision-making task. The collaborative update mechanism ensures that the anti-interference decision model can continuously learn and optimize in a dynamically changing interference environment, thereby continuously improving its adaptability and intelligent decision-making capabilities, effectively coping with complex and ever-changing interference situations, and improving the anti-interference effect of wireless communication systems.

[0077] Through the above technical solution, this application utilizes a pre-trained anti-interference decision model, composed of a cascaded representation learning model and a policy mapping model, based on the high-dimensional state information of the sensing communication environment, to output the current optimal anti-interference action. After executing the anti-interference action, response information is actively collected, and the model is collaboratively updated, achieving deep coupling between feature extraction and decision optimization.

[0078] Example 2, as Figure 2 As shown, based on the same inventive concept as the reinforcement learning-based wireless communication anti-interference strategy optimization method provided in Embodiment 1, this application also provides a reinforcement learning-based wireless communication anti-interference strategy optimization system, including: The information acquisition module 11 is used to acquire the communication environment status information of the target scene in real time, wherein the communication environment status information is high-dimensional status data; The model training module 12 is used to input the communication environment state information into the pre-trained anti-interference decision model to obtain the anti-interference action at the current moment. The anti-interference decision model is composed of a representation learning model and a policy mapping model cascaded together. The model update module 13 is used to execute the anti-interference action through the anti-interference action executor, synchronously collect the anti-interference action response information, and collaboratively update the anti-interference decision model based on the anti-interference action response information.

[0079] In one embodiment, the information acquisition module 11 is specifically used for: The raw wireless signal containing the target frequency band is collected by deploying a radio frequency front-end and analog-to-digital converter at the communication receiver. The original wireless signal is preprocessed, wherein the preprocessing includes at least down-conversion, filtering and synchronization operations; Statistical analysis and dimensional filtering are performed on the preprocessing results to obtain multi-dimensional key state information related to anti-interference decision-making, and the communication environment state information in the form of multi-dimensional feature vectors is generated.

[0080] Furthermore, the communication environment status information includes at least: Channel quality state dimension, which is used to characterize the channel fading characteristics of the target communication link; Interference situation dimension, which is used to characterize the intensity distribution and variation characteristics of external interference signals; The communication resource occupancy dimension is used to characterize the current usage status of communication spectrum and power resources. The communication performance feedback dimension is used to characterize at least one of the communication error rate, throughput, and latency status.

[0081] Furthermore, in one embodiment of the application, the training step of the anti-interference decision model includes: Construct and train the representation learning model, which is used to encode the input high-dimensional state data into a low-dimensional feature vector; Construct and train a policy mapping model based on reinforcement learning, wherein the policy mapping model takes the low-dimensional feature vector as input and outputs the anti-interference action; The policy mapping model and the representation learning model, after cascading training, generate the anti-interference decision model; The representation learning model includes an encoder and decoder with an asymmetric configuration.

[0082] Furthermore, in one embodiment of the application, constructing and training the representation learning model includes: Raw state data from historical communication scenarios are collected to form a pre-training dataset, and the pre-training dataset is organized into batch training data according to the dimensional configuration of the communication environment state information. Construct an original representation learning model with an encoder and a decoder, and initialize the network parameters of the original representation learning model; The original representation learning model is iteratively trained with the goal of minimizing the reconstruction loss between the output data of the decoder and the original state data input to the encoder. When the reconstruction loss is lower than a preset threshold, training is completed and the original representation learning model is output as the representation learning model.

[0083] Furthermore, a policy mapping model based on reinforcement learning is constructed and trained, including: Construct a reinforcement learning-based original policy network and initialize the network parameters, while simultaneously defining the reward calculation function; During the interaction between the wireless communication system and the environment in the target scenario, sequential experience data consisting of states, actions, rewards, and new states is collected. The reward function is configured with the goal of maximizing the expected future cumulative reward, and the original policy network and the reward calculation function are iteratively optimized and trained by combining the sequence experience data and reinforcement learning methods. When the reward function converges, training is complete and the policy mapping model consisting of the trained original policy network is output.

[0084] Furthermore, in one embodiment, the anti-interference action is executed by an anti-interference action actuator, and anti-interference action response information is collected synchronously, including: The anti-interference action is executed periodically with a preset decision time slot. The feedback reward generated after the anti-interference action is executed is collected, wherein the feedback reward is obtained based on the reward calculation function associated with the policy mapping model; Obtain the communication environment status information at the next moment after the anti-interference action is executed; The current communication environment status information, the anti-interference action, the feedback reward, and the next communication environment status information are structured into a single piece of empirical data and stored. Multiple pieces of empirical data constitute the anti-interference action response information.

[0085] Furthermore, in one embodiment, the anti-interference decision model is collaboratively updated based on the anti-interference action response information, including: When the amount of experience data stored in the anti-interference action response information reaches a preset activation threshold, a collaborative update is activated. Randomly select small batches of data samples from the stored multiple sets of empirical data; The small batch of data samples are transmitted to the representation learning model and the policy mapping model in the anti-interference decision model respectively, and the parameters are updated collaboratively.

[0086] Further, the small batch of data samples are transmitted to the representation learning model and the policy mapping model in the anti-interference decision model, respectively, for collaborative parameter updates, including: The policy mapping model calculates the policy gradient based on the communication environment state information in the mini-batch data samples, the anti-interference action, the feedback reward, and the communication environment state information at the next moment, combined with reinforcement learning methods. With the goal of maximizing the expected future cumulative reward, the network parameters of the policy mapping model are updated according to the policy gradient; The representation learning model calculates the reconstruction loss by forward propagation based on the communication environment state information in the mini-batch data samples, and receives the policy gradient by backpropagation from the policy mapping model. The joint loss is calculated by weighting the reconstruction loss and the policy gradient, and the network parameters of the encoder in the representation learning model are updated by backpropagation based on the joint loss. Based on the updated anti-interference decision model, the anti-interference strategy optimization is performed in the next cycle.

[0087] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0088] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0089] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.

Claims

1. A method for optimizing anti-interference strategies in wireless communication based on reinforcement learning, characterized in that, include: Real-time acquisition of communication environment status information of the target scene, wherein the communication environment status information is high-dimensional status data; The communication environment state information is input into a pre-trained anti-interference decision model to obtain the anti-interference action at the current moment. The anti-interference decision model is constructed by cascading a representation learning model and a policy mapping model. The anti-interference action is executed by the anti-interference action actuator, the anti-interference action response information is collected synchronously, and the anti-interference decision model is updated collaboratively based on the anti-interference action response information.

2. The wireless communication anti-interference strategy optimization method based on reinforcement learning as described in claim 1, characterized in that, Real-time acquisition of communication environment status information of the target scene, wherein the communication environment status information is high-dimensional status data, including: The raw wireless signal containing the target frequency band is collected by deploying a radio frequency front-end and analog-to-digital converter at the communication receiver. The original wireless signal is preprocessed, wherein the preprocessing includes at least down-conversion, filtering and synchronization operations; Statistical analysis and dimensional filtering are performed on the preprocessing results to obtain multi-dimensional key state information related to anti-interference decision-making, and the communication environment state information in the form of multi-dimensional feature vectors is generated.

3. The wireless communication anti-interference strategy optimization method based on reinforcement learning as described in claim 1, characterized in that, The communication environment status information includes at least: Channel quality state dimension, which is used to characterize the channel fading characteristics of the target communication link; Interference situation dimension, which is used to characterize the intensity distribution and variation characteristics of external interference signals; The communication resource occupancy dimension is used to characterize the current usage status of communication spectrum and power resources. The communication performance feedback dimension is used to characterize at least one of the communication error rate, throughput, and latency status.

4. The wireless communication anti-interference strategy optimization method based on reinforcement learning as described in claim 1, characterized in that, The training steps of the anti-interference decision model include: Construct and train the representation learning model, which is used to encode the input high-dimensional state data into a low-dimensional feature vector; Construct and train a policy mapping model based on reinforcement learning, wherein the policy mapping model takes the low-dimensional feature vector as input and outputs the anti-interference action; The policy mapping model and the representation learning model, after cascading training, generate the anti-interference decision model; The representation learning model includes an encoder and decoder with an asymmetric configuration.

5. The wireless communication anti-interference strategy optimization method based on reinforcement learning as described in claim 4, characterized in that, Constructing and training the representation learning model includes: Raw state data from historical communication scenarios are collected to form a pre-training dataset, and the pre-training dataset is organized into batch training data according to the dimensional configuration of the communication environment state information. Construct an original representation learning model with an encoder and a decoder, and initialize the network parameters of the original representation learning model; The original representation learning model is iteratively trained with the goal of minimizing the reconstruction loss between the output data of the decoder and the original state data input to the encoder. When the reconstruction loss is lower than a preset threshold, training is completed and the original representation learning model is output as the representation learning model.

6. The wireless communication anti-interference strategy optimization method based on reinforcement learning as described in claim 4, characterized in that, Construct and train a policy mapping model based on reinforcement learning, including: Construct a reinforcement learning-based original policy network and initialize the network parameters, while simultaneously defining the reward calculation function; During the interaction between the wireless communication system and the environment in the target scenario, sequential experience data consisting of states, actions, rewards, and new states is collected. The reward function is configured with the goal of maximizing the expected future cumulative reward, and the original policy network and the reward calculation function are iteratively optimized and trained by combining the sequence experience data and reinforcement learning methods. When the reward function converges, training is complete and the policy mapping model consisting of the trained original policy network is output.

7. The wireless communication anti-interference strategy optimization method based on reinforcement learning as described in claim 1, characterized in that, The anti-interference action is executed by the anti-interference action actuator, and the anti-interference action response information is collected synchronously, including: The anti-interference action is executed periodically with a preset decision time slot. The feedback reward generated after the anti-interference action is executed is collected, wherein the feedback reward is obtained based on the reward calculation function associated with the policy mapping model; Obtain the communication environment status information at the next moment after the anti-interference action is executed; The current communication environment status information, the anti-interference action, the feedback reward, and the next communication environment status information are structured into a single piece of empirical data and stored. Multiple pieces of empirical data constitute the anti-interference action response information.

8. The wireless communication anti-interference strategy optimization method based on reinforcement learning as described in claim 1, characterized in that, The anti-interference decision model is collaboratively updated based on the anti-interference action response information, including: When the amount of experience data stored in the anti-interference action response information reaches a preset activation threshold, a collaborative update is activated. Randomly select small batches of data samples from the stored multiple sets of empirical data; The small batch of data samples are transmitted to the representation learning model and the policy mapping model in the anti-interference decision model respectively, and the parameters are updated collaboratively.

9. The wireless communication anti-interference strategy optimization method based on reinforcement learning as described in claim 8, characterized in that, The small batch of data samples is transmitted to the representation learning model and the policy mapping model in the anti-interference decision model, respectively, for collaborative parameter updates, including: The policy mapping model calculates the policy gradient based on the communication environment state information in the mini-batch data samples, the anti-interference action, the feedback reward, and the communication environment state information at the next moment, combined with reinforcement learning methods. With the goal of maximizing the expected future cumulative reward, the network parameters of the policy mapping model are updated according to the policy gradient; The representation learning model calculates the reconstruction loss by forward propagation based on the communication environment state information in the mini-batch data samples, and receives the policy gradient by backpropagation from the policy mapping model. The joint loss is calculated by weighting the reconstruction loss and the policy gradient, and the network parameters of the encoder in the representation learning model are updated by backpropagation based on the joint loss. Based on the updated anti-interference decision model, the anti-interference strategy optimization is performed in the next cycle.

10. A wireless communication anti-interference strategy optimization system based on reinforcement learning, characterized in that, The method for optimizing wireless communication anti-interference strategies based on reinforcement learning as described in any one of claims 1-9 includes: The information acquisition module is used to acquire the communication environment status information of the target scene in real time, wherein the communication environment status information is high-dimensional status data; The model training module is used to input the communication environment state information into the pre-trained anti-interference decision model to obtain the anti-interference action at the current moment. The anti-interference decision model is composed of a representation learning model and a policy mapping model cascaded together. The model update module is used to execute the anti-interference action through the anti-interference action executor, synchronously collect the anti-interference action response information, and collaboratively update the anti-interference decision model based on the anti-interference action response information.

Citation Information

Cited By

  • A method for optimizing a radio frequency front end for anti-jamming through-the-earth communication

    CN122340517A

  • A method for optimizing radio frequency front-end for strong interference-resistant ground-penetrating communication

    CN122340517B