Rapid anti-interference communication method and system based on state-action similarity weighted reward mechanism
By adopting a fast anti-interference communication method based on the state-action similarity weighted reward mechanism in the wireless communication system, the problem that traditional technology is difficult to adapt to the dynamic electromagnetic environment is solved, and the anti-interference performance and learning efficiency of the communication system are improved.
Patent Information
- Application Number
- CN202411874880.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-19
AI Technical Summary
Traditional anti-interference technology is difficult to adapt to the dynamically changing electromagnetic environment, resulting in decreased communication quality and interruption, deep reinforcement learning algorithms have high complexity and slow convergence speed, which affects the optimization and performance improvement of communication systems.
A fast anti-interference communication method based on the state-action similarity weighted reward mechanism is adopted. By optimizing the reward function and introducing the similarity measurement method of state-action pairs, a sample database is constructed using historical successful experience, and the exploration-utilization mechanism is optimized to improve the learning efficiency and decision-making quality of the agent.
Significantly improve the anti-interference performance of wireless communication systems, enhance the stability and reliability of data transmission, reduce computing complexity, improve learning efficiency, and enhance the adaptability and robustness of the system in complex electromagnetic environments.
Smart Images

Figure CN119946659A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication anti-interference technology, and in particular to a fast anti-interference communication system, method and system based on a state-action similarity weighted reward mechanism. Background Art
[0002] With the rapid development of wireless communication technology, the electromagnetic environment is becoming increasingly complex, and communication systems are facing various forms of interference, such as malicious interference, co-channel interference, and environmental noise. These interferences will seriously affect the communication quality and even cause communication interruption. Traditional anti-interference technology mainly relies on fixed spectrum allocation and preset anti-interference strategies, which are difficult to adapt to the dynamically changing electromagnetic environment in a timely manner.
[0003] In recent years, with the rapid development of artificial intelligence technology, Deep Reinforcement Learning (DRL) combines the powerful feature extraction capabilities of deep learning and the decision optimization capabilities of reinforcement learning, especially the introduction of Deep Q-Network (DQN), making it the mainstream way to solve the problem of communication anti-interference. However, such algorithms are often highly complex and take a long time to converge, which is not conducive to the optimization and performance improvement of wireless communication systems. Therefore, it is of great significance to improve the convergence speed of anti-interference algorithms, reduce the complexity of calculations, and maintain a high success rate of confrontation. Summary of the invention
[0004] In view of the above-mentioned problems, the present invention provides a fast anti-interference communication method and system based on a state-action similarity weighted reward mechanism. By optimizing the reward function, a reward function that can comprehensively consider the immediate reward and state-action similarity is designed. At the same time, a similarity measurement method for state-action pairs is introduced, and a state-action pair sample database is constructed using historical successful anti-interference experience to ensure that high-value and high-similarity experience samples can be used first. Combined with the optimized exploration-utilization mechanism, the intelligent agent can more effectively utilize historical experience, improve learning efficiency and decision-making quality, simplify calculation complexity, and enhance the anti-interference ability and transmission efficiency of wireless communication systems in complex electromagnetic environments.
[0005] The present invention provides a fast anti-interference communication method based on a state-action similarity weighted reward mechanism, which is applied to a wireless communication system and includes the following steps:
[0006] Step 1: Model the communication anti-interference problem as a Markov decision process, in which, based on the electromagnetic environment faced by the wireless communication system, the environmental state of different time slots is defined, including the instantaneous signal power observed by the receiver in each time slot, and the transmission actions selected by the wireless communication system in different time slots are defined, including the transmission channel, transmission power and transmission rate selected by the transmitter;
[0007] Step 2: construct a deep Q network and initialize network parameters, wherein the deep Q network includes a policy Q network and a target Q network. The structure of the target Q network is the same as that of the policy Q network, and is used to provide a stable target Q value during training to calculate and update the action value of the agent;
[0008] Step 3: Calculate the environment state of the previous time slot (t-1) and the instant reward of the transmission action, combine the environment state of the previous time slot (t-1), the transmission action and the environment state of the current time slot (t), and obtain a set of complete state-action samples and store them in the database;
[0009] Step 4: Use the historical samples in the state-action pair database to calculate the similarity metric of the current state-action pair, and generate a weighted reward by weighting the similarity;
[0010] Step 5: The environment state, transmission action, weighted reward of the previous time slot (t-1), the environment state of the current time slot (t) and the termination flag are used to form a set of experience samples and stored in the experience pool for subsequent training;
[0011] Step 6: Adjust the loss function of the deep Q network based on the weighted reward and update the weights of the Q network through the back-propagation algorithm.
[0012] Step 7: Use the updated strategy Q network to select the transmission action for the current time slot (t) and feed it back to the wireless communication system; if the communication is not over, return to step 3 and continue to perform the transmission action selection and system interaction for the next time slot (t+1) until the communication process ends.
[0013] In some embodiments, step 1 includes the following sub-steps:
[0014] Step 101: define the environment state of the wireless communication system by the average received power;
[0015] Step 102: The transmission channel, transmission power and transmission rate of the T+1 time slot constitute the transmission action decided by the receiver in the T time slot;
[0016] Step 103: Define a reward function for the instant reward based on the weight coefficient of the transmission rate, the cost factor of the transmission power, the average SJNR of the receiving end in the T+1 time slot, and the SJNR demodulation threshold of the receiver.
[0017] In some embodiments, the higher the transmission rate selected by the agent, the higher the weight coefficient of the transmission rate; the lower the transmission power selected by the agent, the lower the cost factor of the transmission power of the wireless wage system.
[0018] In some embodiments, step 2 includes the following sub-steps:
[0019] Step 201, defining a strategy Q network for estimating a state-action value vector Q(s,a), wherein the strategy Q network is a multi-layer feedforward neural network;
[0020] Step 202: define a target Q network, the structure of which is the same as that of the policy Q network, and is used to provide a stable target Q value during the training process to calculate and update the action value of the agent;
[0021] Step 203: Initialize the network parameters of the strategy Q network and the target Q network, including weights and bias items;
[0022] Step 204: Select an optimization algorithm and a loss function to update the network parameters of the strategy Q network;
[0023] Step 205: setting hyperparameters during training;
[0024] Step 206: Initialize the experience pool to store experience samples obtained by the agent during the interaction process;
[0025] Step 207, initializing the step counter;
[0026] Step 208: During the training process, the network parameters of the policy Q network are periodically copied to the target Q network, that is, the target network update operation is performed.
[0027] In some embodiments, step 3 includes the following sub-steps:
[0028] Step 301: In each time slot, the receiver observes the environmental state s of the current time slot (t) t ;
[0029] Step 302: The transmitter performs transmission action a in the previous time slot (t-1). t-1 ;
[0030] Step 303: According to the environmental state s t and transfer action a t-1 , use the preset reward function to calculate the instant reward r(s t ,a t-1 );
[0031] Step 304: The receiver observes the environmental state s at the current time slot (t) t+1, as the environmental state of the next time slot (t+1);
[0032] Step 305: Set the current time slot environment state s t , transmission action a t-1 、Instant reward r(s t ,a t-1 ), the environmental state s of the next time slot t+1 Combined to form a complete state-action pair sample (s t ,a t-1 ,r(s t ,a t-1 ),s t+1 );
[0033] Step 306: judge the current sample. If the immediate reward meets the requirements and there is no sample in a similar state in the database, store it in the state-action pair database.
[0034] In some embodiments, step 4 includes the following sub-steps:
[0035] Step 401, dynamically update the state space and state-action pair database;
[0036] Step 402: extract historical samples from the state-action pair sample database, including previous state-action pairs (s′, a′) and their corresponding immediate rewards r(s′, a′);
[0037] Step 403: For the current state-action pair (s, a), compare it with each historical sample. If the difference between the immediate reward r(s, a) of the current state-action pair (s, a) and the immediate reward r(s′, a′) of the historical sample state-action pair (s′, a′) is less than or equal to the similarity threshold, and the state transition probability P of the current state-action pair f (s,a) and the state transition probability P of historical samples f When (s′, a′) are equal, the two state-action pairs are judged to satisfy the similarity condition;
[0038] Step 404: For samples that successfully find similar samples in the state-action pair sample database, their immediate rewards are weighted by similarity to obtain weighted rewards.
[0039] In some embodiments, step 401 specifically includes: recording all historically observed states from the initial transmission to the current state, obtaining a statistical state set Expand possible future states to form a more complete state space based on historical data and model predictions Then, the state-action pairs that successfully resist interference and their related information are stored in the state-action pair sample database D, including the environment state s′, the transmission action a′, the immediate reward r(s′, a′), the state-action value vector Q(s′, a′) and the transition probability P f (s′, a′); record the state at each time step and update the state-action pair sample database.
[0040] In some embodiments, step 6 includes the following sub-steps:
[0041] Step 601: Randomly sample a batch of experience samples (s i ,a i-1 ,R(s i ,a i-1 ),s i+1 ,done i ), where s i is the environmental state, a i is the transmission action, s i+1 For the next environment state, done i It is the termination sign;
[0042] Step 602: For each empirical sample, calculate the target Q value:
[0043]
[0044] Among them, Q target is the target Q network, used to calculate the next environment state s i+1 All possible transfer actions a i+1 The maximum Q value of; γ is the discount factor;
[0045] Step 603: Calculate the predicted value Q(s) of the current strategy Q network i ,a i ; θ), where θ is the network parameter of the strategy Q network;
[0046] Step 604: Define the loss function L(θ) of the deep Q network as:
[0047]
[0048] Among them, N is the batch size, which means the number of samples used in each training;
[0049] Step 605: Use the back propagation algorithm to calculate the gradient of the loss function to the network parameter θ of the policy Q network.
[0050] Step 606, updating the network parameters θ of the strategy Q network according to the gradient;
[0051] Step 607: Periodically copy the updated network parameters θ of the policy Q network to the network parameters of the target Q network.
[0052] In some embodiments, step 7 includes the following sub-steps:
[0053] Step 701: In the current time slot (t), use the updated strategy Q network Q (s t ,a t ;θ), according to the current environment state s t , select the optimal transmission action
[0054] Step 702: The selected transmission action Applied to wireless communication systems, that is, the transmitter performs transmission actions in the current time slot (t)
[0055] Step 703: The transmitter sends a signal according to the selected transmission parameters, and the receiver receives the signal synchronously in the same time slot, and performs demodulation and decoding according to the current electromagnetic environment state;
[0056] Step 704: The wireless communication system obtains the transmission result of the current time slot according to the feedback from the receiver and updates the environment state s t+1 ;
[0057] Step 705: The environmental state s of the current time slot (t) t , transfer action Instant reward r(s,a), environment state s at the next time slot (t+1) t+1 and the termination symbol done to form a new experience sample And stored in the experience pool.
[0058] Step 706: If the communication is not finished, return to step 3 and continue to perform the transmission action selection and system interaction for the next time slot until the communication process is finished.
[0059] The present invention also provides a wireless communication system, comprising a transmitter and a receiver, wherein the transmitter and the receiver are synchronized according to time slots, and the above-mentioned fast anti-interference communication method based on the state-action similarity weighted reward mechanism is executed during the communication process.
[0060] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0061] 1. By introducing a reward mechanism based on similarity measurement, the present invention enables the intelligent agent to more effectively identify and utilize historical successful experiences, thereby selecting better transmission actions in complex electromagnetic environments, significantly improving the anti-interference performance of the wireless communication system, and enhancing the stability and reliability of data transmission.
[0062] 2. The present invention adopts a similarity-weighted reward function, which enables the agent to fully utilize the experience of similar state-action pairs during training, accelerate the convergence speed of DQN, reduce the time and computing resources required for training, and improve the overall learning efficiency.
[0063] 3. The present invention combines a dynamically adjusted exploration-utilization mechanism, and the intelligent agent can flexibly balance the exploration of new strategies and the utilization of known optimal strategies in different interference environments, thereby improving the intelligence level of decision-making and ensuring that the optimal transmission solution is always selected in a changing environment.
[0064] 4. By storing instant rewards and weighted rewards simultaneously, the present invention enables the system to maintain high adaptability and robustness in the face of diverse and dynamically changing interference sources, reduce communication interruptions and performance degradation caused by environmental changes, and ensure the continuous and stable operation of the communication system.
[0065] 5. The present invention stores both instant rewards and weighted rewards in the experience replay pool, simplifies the implementation logic of similarity calculation and training process, avoids data redundancy and complex synchronization problems, makes the system design more concise and efficient, and facilitates practical application and further optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 The flowchart of the fast anti-interference communication method based on the state-action similarity weighted reward mechanism in an embodiment of the present invention.
[0067] Figure 2 This is a flowchart of constructing and training a deep Q network in an embodiment of the present invention.
[0068] Figure 3 The flowchart of generating weighted rewards by similarity weighting in an embodiment of the present invention. DETAILED DESCRIPTION
[0069] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0070] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0071] like Figure 1 As shown, this embodiment proposes a fast anti-interference communication method based on a state-action similarity weighted reward mechanism, which is applied to a wireless communication system, wherein the wireless communication system includes a transmitter and a receiver, and the transmitter and the receiver are synchronized according to time slots. The method includes the following steps:
[0072] Step 1: Model the communication anti-interference problem as a Markov decision process (MDP), in which, based on the electromagnetic environment faced by the wireless communication system, the environmental state of different time slots is defined, including the instantaneous signal power observed by the receiver in each time slot, and the transmission actions selected by the wireless communication system in different time slots are defined, including the transmission channel, transmission power and transmission rate selected by the transmitter;
[0073] Step 2: construct a deep Q-network (DQN) and initialize network parameters, wherein the deep Q-network includes a policy Q-network and a target Q-network. The structure of the target Q-network is the same as that of the policy Q-network, and is used to provide a stable target Q value during training to calculate and update the action value of the agent;
[0074] Step 3: Calculate the environment state of the previous time slot (t-1) and the instant reward of the transmission action, combine the environment state of the previous time slot (t-1), the transmission action and the environment state of the current time slot (t), and obtain a set of complete state-action samples and store them in the database;
[0075] Step 4: Use the historical samples in the state-action pair database to calculate the similarity metric of the current state-action pair, and generate a weighted reward by weighting the similarity;
[0076] Step 5: The environment state, transmission action, weighted reward of the previous time slot (t-1), the environment state of the current time slot (t) and the termination flag are used to form a set of experience samples and stored in the experience pool for subsequent training;
[0077] Step 6: Adjust the loss function of the deep Q network based on the weighted reward and update the weights of the Q network through the back-propagation algorithm.
[0078] Step 7: Use the updated strategy Q network to select the transmission action for the current time slot (t) and feed it back to the wireless communication system; if the communication is not over, return to step 3 and continue to perform the transmission action selection and system interaction for the next time slot (t+1) until the communication process ends.
[0079] In some embodiments, in step 1, the communication anti-interference problem is modeled as a Markov decision process, which specifically includes the following sub-steps:
[0080] Step 101: define the environment state of the wireless communication system by the average received power, expressed as:
[0081]
[0082] Among them, S T Indicates the environmental status of the wireless communication system in time slot T, the average received power It consists of the transmitter signal power and the interference signal power received on channel m, expressed as:
[0083]
[0084] Among them, δ(·) is the indicator function, that is, when x is true, δ(x) = 1, otherwise δ(x) = 0; p T-1 is the transmitting power of the transmitter, g T-1 is the transmission channel gain, Indicates the average power of interference plus noise received.
[0085] Step 102: The action space in the Markov decision process contains the set of all possible actions that the agent may take. The transmission channel, transmission power and transmission rate of the T+1 time slot constitute the transmission action decided by the receiver in the T time slot, which is expressed as:
[0086] a T =(f T+1 ,p T+1 ,v 1+1 )∈A (3)
[0087] Among them, a T represents the transmission action of T time slot decision, f T+1 is the transmission channel of the T+1 time slot, p T+1 is the transmit power of the T+1 time slot, v T+1 is the transmission rate of the T+1 time slot;
[0088] Step 103, based on the reward function defining the immediate reward, is expressed as:
[0089]
[0090] in, is the weight coefficient of the transmission rate. The higher the transmission rate selected by the agent, the higher the weight coefficient of the transmission rate. is the cost factor of the transmission power. The lower the transmission power selected by the agent, the lower the cost factor of the transmission power of the wireless wage system. T+1 is the average SJNR of the receiving end in the T+1 time slot, and is the SJNR demodulation threshold of the receiver. The higher the transmission rate, the higher the demodulation threshold. δ(·) is the indicator function. That is, it is 1 when the demodulation is successful, otherwise it is 0.
[0091] In some embodiments, step 2 constructs a deep Q network and initializes network parameters, specifically including the following sub-steps:
[0092] Step 201: define a strategy Q network (QNetwork) for estimating a state-action value vector Q(s,a), wherein the strategy Q network is a multi-layer feedforward neural network, including:
[0093] Input layer: receives a state vector s, the dimension of which is the number of channels M;
[0094] Hidden layer: includes one or more fully connected layers, each hidden layer contains several neurons, and the activation function is a nonlinear function;
[0095] Output layer: outputs the state-action value vector Q(s,a), where the dimension of the state-action value vector is the size of the action space;
[0096] Step 202: define a target Q network (Target QNetwork), the structure of which is the same as that of the policy Q network, and is used to provide a stable target Q value during the training process to calculate and update the action value of the agent;
[0097] Step 203: Initialize network parameters of the strategy Q network and the target Q network, including weights and bias items. The network parameters may be obtained by random initialization or pre-training.
[0098] Step 204: Select an optimization algorithm and a loss function to update the network parameters of the strategy Q network;
[0099] Step 205: Set the hyperparameters in the training process, including: learning rate α, discount factor γ, initial ε value of the exploration strategy, ε decay rate ε decay , batch size and the capacity of the experience replay buffer;
[0100] Step 206: Initialize the experience pool (experience playback buffer) to store the experience samples obtained by the agent during the interaction process, including the environment state, transmission action, weighted reward, next state and termination flag;
[0101] Step 207: Initialize a step counter to record the total number of steps of the agent interacting with the environment, so as to adjust the exploration rate ε according to the number of steps in the ε-greedy strategy;
[0102] Step 208: During the training process, the network parameters of the policy Q network are copied to the target Q network at a certain step size or cycle, that is, the target network update operation is performed to ensure that the target Q network provides a stable target value and reduce the shock and instability during the training process;
[0103] Among them, the target Q network is used to provide a stable reference value when calculating the target Q value, and its network parameters are copied from the strategy Q network at a certain step size or period during the training process; by introducing the target Q network, the instability of the training process can be effectively alleviated, and the learning efficiency and decision-making performance of the intelligent agent in a dynamic interference environment can be improved.
[0104] In some embodiments, in step 3, the environment state of the previous time slot (t-1) and the immediate reward of the transmission action are calculated, and the environment state of the previous time slot (t-1), the transmission action and the environment state of the current time slot (t) are combined to obtain a set of complete state-action samples and store them in the database, which specifically includes the following sub-steps:
[0105] Step 301: In each time slot, the receiver observes the environmental state s of the current time slot (t) t , including instantaneous signal power and interference information;
[0106] Step 302: The transmitter performs transmission action a in the previous time slot (t-1). t-1 , including the selected transmission channel, transmission power and transmission rate;
[0107] Step 303: According to the environmental state s t and transfer action a t-1 , use the preset reward function to calculate the instant reward r(s t ,a t-1 ), the reward function comprehensively considers factors such as communication success rate, channel quality and energy consumption;
[0108] Step 304: The receiver observes the environmental state s at the current time slot (t) t+1 , as the environmental state of the next time slot (t+1);
[0109] Step 305: Set the current time slot environment state s t , transmission action a t-1 、Instant reward r(s t ,a t-1 ), the environmental state s of the next time slot t+1 Combined to form a complete state-action pair sample (s t ,a t-1 ,r(s t ,a t-1 ),s t+1 );
[0110] Step 306: judge the current sample. If the immediate reward meets the requirements and there is no sample in a similar state in the database, store it in the state-action pair database.
[0111] In some embodiments, in step 4, historical samples in the state-action pair database are used to calculate the similarity metric value of the current state-action pair, and similarity weighted rewards are generated by similarity weighting, which specifically includes the following sub-steps:
[0112] Step 401: Dynamically update the state space and state-action pair database:
[0113] In the case that the complete state space cannot be modeled in advance, the accuracy of the similarity measure is gradually improved by dynamically expanding the state space. All historically observed states from the initial transmission to the current state are recorded to obtain a statistical state set. Expand possible future states to form a more complete state space based on historical data and model predictions Then, the state-action pairs that successfully resist interference and their related information are stored in the state-action pair sample database D, including the environment state s′, the transmission action a′, the immediate reward r(s′, a′), the state-action value vector Q(s′, a′) and the transition probability P f (s′, a′); record the state at each time step and update the state-action pair sample database;
[0114] Step 402: extract historical samples from the state-action pair sample database, including previous state-action pairs (s′, a′) and their corresponding immediate rewards r(s′, a′);
[0115] Step 403: For the current state-action pair (s, a), compare it with each historical sample. If the difference between the immediate reward r(s, a) of the current state-action pair (s, a) and the immediate reward r(s′, a′) of the historical sample state-action pair (s′, a′) is less than or equal to the similarity threshold, and the state transition probability P of the current state-action pair f (s,a) and the state transition probability P of historical samples f When (s′, a′) are equal, it can be determined that the two state-action pairs meet the similarity condition; among which:
[0116] The similarity determination condition is expressed as:
[0117] |r(s,a)-r(s′,a′)|≤θ R (5)
[0118] Among them, θ R is the similarity threshold;
[0119] The state transition probability determination condition is expressed as:
[0120] P f (s,a)=P f (s′,a′) (6)
[0121] Step 404: For samples that successfully find similar samples in the state-action pair sample database, their immediate rewards are weighted by similarity to obtain a weighted reward, which is expressed as:
[0122] R(s,a)=r(s,a)*β+λ (7)
[0123] Among them, β represents the scaling factor of the immediate reward r(s,a). By adjusting β, the influence weight of the immediate reward can be increased or decreased according to the similarity measurement results of the state-action pair; λ represents an offset value (translation) added to the scaled reward value. By adjusting λ, the final reward value can be moved up or down from the baseline value to achieve fine-tuning and compensation of the reward signal.
[0124] Through the above steps, the wireless communication system can use historical successful experience in a dynamic and complex electromagnetic environment to calculate the similarity measurement value of the current state-action pair, and optimize the decision-making strategy of the intelligent agent by adjusting the reward function, guiding the intelligent agent to prefer proven effective strategies, improve the communication anti-interference performance and learning efficiency, ensure fairness between different state-action pairs through weighted rewards, and simplify the computational complexity.
[0125] In some embodiments, in step 5, the environment state s of the previous time slot (t-1) is t ,Transmission action a t-1 , weighted reward R(s t ,a t-1 ), the environmental state s of the current time slot (t) t+1 and the termination flag done t A set of experience samples is formed and stored in the experience pool for subsequent training, including the following:
[0126] The experience sample (s t ,a t-1 ,R(s t ,a t-1 ),s t+1 ,done t ) is stored in the experience pool (experience playback buffer) for subsequent training and learning processes;
[0127] In some embodiments, in step 6, the loss function of the deep Q network is adjusted based on the similarity weighted reward, and the weight of the Q network is updated by the back propagation algorithm, which specifically includes the following sub-steps:
[0128] Step 601: Randomly sample a batch of experience samples (s i ,a i-1 ,R(s i ,a i-1 ),s i+1 ,done i ), where s i is the environmental state, a i is the transmission action, s i+1 For the next environment state, done i It is the termination sign;
[0129] Step 602: For each empirical sample, calculate the target Q value:
[0130] y i =R(s i ,a i )+γ·max a′ Q target (s i+1 ,a i+1 )·(1-done i ) (8)
[0131] Among them, Q target is the target Q network, used to calculate the next environment state s i+1 All possible transfer actions a i+1 The maximum Q value of; γ is the discount factor;
[0132] Step 603: Calculate the predicted value Q(s) of the current strategy Q network i ,a i ; θ), where θ is the network parameter of the strategy Q network;
[0133] Step 604: Define the loss function L(θ) of the deep Q network as:
[0134]
[0135] Among them, N is the batch size, which means the number of samples used in each training;
[0136] Step 605: Use the back propagation algorithm to calculate the gradient of the loss function to the network parameter θ of the policy Q network.
[0137] Step 606: Update the network parameters θ of the strategy Q network according to the gradient:
[0138]
[0139] Where η is the learning rate;
[0140] Step 607: Periodically copy the updated network parameters θ of the policy Q network to the network parameters of the target Q network:
[0141] θ target ←θ (11)
[0142] This ensures that the target Q network provides a stable target value and reduces fluctuations during training.
[0143] Through the above steps, the loss function of the deep Q network is adjusted using similarity weighted rewards, so that the policy Q network pays more attention to the state-action pairs similar to the historical successful experience during the training process, improving the learning efficiency and anti-interference ability of the intelligent agent. At the same time, the target Q network, as a delayed copy of the policy Q network, provides a stable target Q value, assists the training of the policy Q network, and enhances the decision-making performance of the intelligent agent in a dynamic interference environment.
[0144] In some embodiments, in step 7, the updated strategy Q network is used to select the transmission action of the current time slot (t), and the updated strategy Q network is fed back to the wireless communication system, which specifically includes the following sub-steps:
[0145] Step 701: In the current time slot (t), use the updated strategy Q network Q (s t ,a t ;θ), according to the current environment state s t , select the optimal transmission action Right now:
[0146]
[0147] Among them, θ is the network parameter of the updated strategy Q network, and a is a set of feasible actions, including different combinations of transmission channels, transmission power and transmission rate.
[0148] Step 702: The selected transmission action Applied to wireless communication systems, that is, the transmitter performs transmission actions in the current time slot (t)
[0149] Step 703: The transmitter sends a signal according to the selected transmission parameters, and the receiver receives the signal synchronously in the same time slot, and performs demodulation and decoding according to the current electromagnetic environment state;
[0150] Step 704: The wireless communication system obtains the transmission result of the current time slot according to the feedback from the receiver and updates the environment state s t+1 ;
[0151] Step 705: The environmental state s of the current time slot (t) t , transfer action Instant reward r(s,a), environment state s at the next time slot (t+1) t+1 and the termination symbol done to form a new experience sample And stored in the experience pool.
[0152] Step 706: If the communication is not finished, return to step 3 and continue to perform the transmission action selection and system interaction for the next time slot until the communication process is finished.
[0153] Among them, by using the updated policy Q network, the agent can t and the learned strategy to select the optimal transmission action in real time And apply it to wireless communication systems, so as to improve the anti-interference performance and transmission efficiency of wireless communication systems in complex electromagnetic environments.
[0154] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A fast anti-interference communication method based on state-action similarity weighted reward mechanism, applied to wireless communication system, characterized in that: The following steps are involved: Step 1: Model the communication anti-interference problem as a Markov decision process, in which, based on the electromagnetic environment faced by the wireless communication system, the environmental state of different time slots is defined, including the instantaneous signal power observed by the receiver in each time slot, and the transmission actions selected by the wireless communication system in different time slots are defined, including the transmission channel, transmission power and transmission rate selected by the transmitter; Step 2: construct a deep Q network and initialize network parameters, wherein the deep Q network includes a policy Q network and a target Q network. The structure of the target Q network is the same as that of the policy Q network, and is used to provide a stable target Q value during training to calculate and update the action value of the agent; Step 3: Calculate the environment state of the previous time slot (t-1) and the instant reward of the transmission action, combine the environment state of the previous time slot (t-1), the transmission action and the environment state of the current time slot (t), and obtain a set of complete state-action samples and store them in the database; Step 4: Use the historical samples in the state-action pair database to calculate the similarity metric of the current state-action pair, and generate a weighted reward by weighting the similarity; Step 5: The environment state, transmission action, weighted reward of the previous time slot (t-1), the environment state of the current time slot (t) and the termination flag are used to form a set of experience samples and stored in the experience pool for subsequent training; Step 6: Adjust the loss function of the deep Q network based on the weighted reward and update the weights of the Q network through the back-propagation algorithm. Step 7: Use the updated strategy Q network to select the transmission action for the current time slot (t) and feed it back to the wireless communication system; if the communication is not over, return to step 3 and continue to perform the transmission action selection and system interaction for the next time slot (t+1) until the communication process ends.
2. The fast anti-interference communication method based on state-action similarity weighted reward mechanism according to claim 1 is characterized in that: Step 1 includes the following sub-steps: Step 101: define the environment state of the wireless communication system by the average received power; Step 102: The transmission channel, transmission power and transmission rate of the T+1 time slot constitute the transmission action decided by the receiver in the T time slot; Step 103: Define a reward function for the instant reward based on the weight coefficient of the transmission rate, the cost factor of the transmission power, the average SJNR of the receiving end in the T+1 time slot, and the SJNR demodulation threshold of the receiver.
3. The fast anti-interference communication method based on state-action similarity weighted reward mechanism according to claim 2 is characterized in that: The higher the transmission rate selected by the agent, the higher the weight coefficient of the transmission rate; the lower the transmission power selected by the agent, the lower the cost factor of the transmission power of the wireless wage system.
4. The fast anti-interference communication method based on state-action similarity weighted reward mechanism according to claim 1 is characterized in that: Step 2 includes the following sub-steps: Step 201, defining a strategy Q network for estimating a state-action value vector Q(s,a), wherein the strategy Q network is a multi-layer feedforward neural network; Step 202: define a target Q network, the structure of which is the same as that of the policy Q network, and is used to provide a stable target Q value during the training process to calculate and update the action value of the agent; Step 203: Initialize the network parameters of the strategy Q network and the target Q network, including weights and bias items; Step 204: Select an optimization algorithm and a loss function to update the network parameters of the strategy Q network; Step 205: setting hyperparameters during training; Step 206: Initialize the experience pool to store experience samples obtained by the agent during the interaction process; Step 207, initializing the step counter; Step 208: During the training process, the network parameters of the policy Q network are periodically copied to the target Q network, that is, the target network update operation is performed.
5. The fast anti-interference communication method based on state-action similarity weighted reward mechanism according to claim 1 is characterized in that: Step 3 includes the following sub-steps: Step 301: In each time slot, the receiver observes the environmental state s of the current time slot (t) t ; Step 302: The transmitter performs transmission action a in the previous time slot (t-1). t-1 ; Step 303: According to the environmental state s t and transfer action a t-1 , use the preset reward function to calculate the instant reward r(s t ,a t-1 ); Step 304: The receiver observes the environmental state s at the current time slot (t) t+1 , as the environmental state of the next time slot (t+1); Step 305: Set the current time slot environment state s t , transmission action a t-1 、Instant reward r(s t ,a t-1 ), the environmental state s of the next time slot t+1 Combined to form a complete state-action pair sample (s t ,a t-1 ,r(s t ,a t-1 ),s t+1 ); Step 306: judge the current sample. If the immediate reward meets the requirements and there is no sample in a similar state in the database, store it in the state-action pair database.
6. The fast anti-interference communication method based on state-action similarity weighted reward mechanism according to claim 1 is characterized in that: Step 4 includes the following sub-steps: Step 401, dynamically update the state space and state-action pair database; Step 402: extract historical samples from the state-action pair sample database, including previous state-action pairs (s′, a′) and their corresponding immediate rewards r(s′, a′); Step 403: For the current state-action pair (s, a), compare it with each historical sample. If the difference between the immediate reward r(s, a) of the current state-action pair (s, a) and the immediate reward r(s′, a′) of the historical sample state-action pair (s′, a′) is less than or equal to the similarity threshold, and the state transition probability P of the current state-action pair f (s,a) and the state transition probability P of historical samples f (s ′ ,a ′ ) are equal, then the two state-action pairs are determined to meet the similarity condition; Step 404: For samples that successfully find similar samples in the state-action pair sample database, their immediate rewards are weighted by similarity to obtain weighted rewards.
7. The fast anti-interference communication method based on state-action similarity weighted reward mechanism according to claim 6 is characterized in that: Step 401 specifically includes: recording all historically observed states from the initial transmission to the current state, and obtaining a statistical state set Expand possible future states to form a more complete state space based on historical data and model predictions Then, the state-action pairs that successfully resist interference and their related information are stored in the state-action pair sample database D, including the environment state s′, the transmission action a′, the immediate reward r(s′, a′), the state-action value vector Q(s′, a′) and the transition probability P f (s ′ ,a ′ ); record the state at each time step and update the state-action pair sample database.
8. The fast anti-interference communication method based on state-action similarity weighted reward mechanism according to claim 1 is characterized in that: Step 6 includes the following sub-steps: Step 601: Randomly sample a batch of experience samples (s i ,a i-1 ,R(s i ,a i-1 ),s i+1 ,done i ), where s i is the environmental state, a i is the transmission action, s i+1 For the next environment state, done i It is the termination sign; Step 602: For each empirical sample, calculate the target Q value: Among them, Q target is the target Q network, used to calculate the next environment state s i+1 All possible transfer actions a i+1 The maximum Q value of; γ is the discount factor; Step 603: Calculate the predicted value Q(s) of the current strategy Q network i ,a i ; θ), where θ is the network parameter of the strategy Q network; Step 604: Define the loss function L(θ) of the deep Q network as: Among them, N is the batch size, which means the number of samples used in each training; Step 605: Use the back propagation algorithm to calculate the gradient of the loss function to the network parameter θ of the policy Q network. Step 606, updating the network parameters θ of the strategy Q network according to the gradient; Step 607: Periodically copy the updated network parameters θ of the policy Q network to the network parameters of the target Q network.
9. The fast anti-interference communication method based on state-action similarity weighted reward mechanism according to claim 1, characterized in that: Step 7 includes the following sub-steps: Step 701: In the current time slot (t), use the updated strategy Q network Q (s t ,a t ;θ), according to the current environment state s t , select the optimal transmission action Step 702: The selected transmission action Applied to wireless communication systems, that is, the transmitter performs transmission actions in the current time slot (t) Step 703: The transmitter sends a signal according to the selected transmission parameters, and the receiver receives the signal synchronously in the same time slot, and performs demodulation and decoding according to the current electromagnetic environment state; Step 704: The wireless communication system obtains the transmission result of the current time slot according to the feedback from the receiver and updates the environment state s t+1 ; Step 705: The environmental state s of the current time slot (t) t , transfer action Instant reward r(s,a), environment state s at the next time slot (t+1) t+1 and the termination symbol done to form a new experience sample And stored in the experience pool. Step 706: If the communication is not finished, return to step 3 and continue to perform the transmission action selection and system interaction for the next time slot until the communication process is finished.
10. A wireless communication system, characterized in that: The invention comprises a transmitter and a receiver, wherein the transmitter and the receiver are synchronized according to time slots, and during the communication process, a fast anti-interference communication method based on a state-action similarity weighted reward mechanism as described in any one of claims 1 to 9 is executed.
Citation Information
Patent Citations
Traffic light control method based on deep reinforcement learning and inverse reinforcement learning
CN115762199A
Intelligent anti-interference decision-making method for nonlinear frequency sweeping interference
CN117595963A
Multi-domain communication anti-interference method and system based on width reinforcement learning, and medium
CN117998419A
Access network slice resource allocation method based on improved multi-agent reinforcement learning
CN118283625A
Distributed Deep Reinforcement Learning Framework for Software-Defined Unmanned Aerial Vehicle Network Control
US20240205095A1