Non-orthogonal multiple access system physical layer security communication method based on meta-reinforcement learning

By optimizing power allocation through meta-reinforcement learning networks, the problem of weak generalization ability of deep reinforcement learning in dynamic environments is solved, and secure communication in non-orthogonal multiple access systems is realized, thereby improving the security and transmission rate of the communication system.

CN116405930BActive Publication Date: 2026-02-17FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310259528.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-16
Publication Date
2026-02-17
Estimated Expiration
2043-03-16

AI Technical Summary

Technical Problem

Existing deep reinforcement learning methods have weak generalization ability in dynamic environments and are difficult to adapt to the dynamic channel characteristics in communication systems, resulting in power allocation strategies failing to effectively improve communication security.

Method used

By employing a meta-reinforcement learning network that combines meta-learning and deep reinforcement learning, a power allocation optimization objective function is constructed, and the meta-reinforcement learning network is used for physical layer security encryption of the system, thereby achieving secure communication in a non-orthogonal multiple access system.

Benefits of technology

It improves the system's communication security and transmission rate in dynamic environments, enhances the system's generalization ability, and adapts to changes in different channel environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116405930B_ABST
    Figure CN116405930B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of communication, and particularly relates to a non-orthogonal multiple access system physical layer security communication method based on meta-reinforcement learning. The application comprises the following steps: constructing a power allocation optimization objective function for maximizing system physical layer security and rate, wherein the case of multiple eavesdroppers eavesdropping information is considered; and using a meta-reinforcement learning network to perform security encryption on the system physical layer, so as to realize non-orthogonal multiple access system physical layer security communication. The application overcomes the defects of the existing power allocation method based on deep reinforcement learning, solves the problem that the prior art cannot be applied to a changing channel environment and is thus difficult to be practically applied, and improves the security of the non-orthogonal multiple access system physical layer.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of communication, and particularly relates to a non-orthogonal multiple access system physical layer secure communication method based on meta-reinforcement learning. BACKGROUND

[0002] The development of communication technology brings a variety of new transmission technologies, which not only improve the performance of legitimate users, but also exist the situation of being eavesdropped by malicious users, which brings new security challenges. In the face of potential physical layer security threats, combining physical layer security technology into the system to protect the security of information transmission has become a recent research hotspot. At the same time, in the problem of multi-user communication, power allocation has always been an important link to optimize the overall communication quality of the system, which has attracted widespread attention. With the increasing physical layer security problems of communication systems in actual situations, in order to improve the security of communication and reduce the loss caused by eavesdropping, the physical layer encryption method for the communication system is becoming more and more important. Under the modeling of the specific communication system model, the physical layer encryption method of the system can be summarized as maximizing the sum rate of the system while minimizing the rate of the eavesdropping end as much as possible. In this way, the method of encrypting the physical layer can be simplified as optimizing the power allocation or resource allocation strategy of the system to achieve the purpose of enhancing the physical layer security of the system and encrypting the physical layer. At present, the method of using deep reinforcement learning technology to allocate power in the communication system is gradually increasing. Deep reinforcement learning is a combination of deep learning and reinforcement learning. This method can learn a specific power allocation strategy based on the interaction between the agent and the environment, but the premise of this method is that the environment is constant and static. In actual situations, the channel characteristics are dynamic, which means that the static environment power allocation strategy learned by the deep reinforcement learning method is difficult to apply to the dynamic environment.

[0003] Meta-reinforcement learning is a method that combines the advantages of meta-learning and reinforcement learning. Meta-learning algorithm is an algorithm for learning the initialization parameters of deep learning model. Under the learned initialization parameters, the model can converge with a small amount of training data, adapt to new environment characteristics, and quickly deploy to a completely new environment. The reinforcement learning method combined with the meta-learning algorithm solves the shortcomings of the original method that the generalization ability is weak and a large amount of training data is required. SUMMARY

[0004] The purpose of the application is to propose a non-orthogonal multiple access system physical layer secure communication method based on meta-reinforcement learning with generalization ability and unable to adapt to dynamic environment. And an optimization objective function of power allocation is proposed.

[0005] The application provides a non-orthogonal multiple access system physical layer security communication method based on meta-reinforcement learning.

[0006] (I) constructing a power allocation optimization objective function aiming at maximizing system physical layer security and transmission rate

[0007] The non-orthogonal multiple access (wireless communication) system provided by the application comprises a wireless system sending end user, a receiving end base station and a malicious eavesdropping end, as shown in Figure 1 .

[0008] Let the sending end i-th user sending signal S i be represented as:

[0009]

[0010] wherein P total is the total transmission power of the sending end, alpha i is the power allocation factor (coefficient) of the i-th user, X i is the information signal of the i-th user, i = 1, 2... n, and n is the number of users; the channels of the user-base station, user-eavesdropping end and base station-eavesdropping end are respectively represented by channel coefficients: h sd , h se , h de The channel coefficient is a random variable obeying Rayleigh distribution, i.e. h ~ Rayleigh (sigma r 2 ), sigma r 2 is the variance of Rayleigh distribution; the noise signal is Gaussian noise obeying Gaussian distribution, i.e. n ~ N (0, sigma g 2 ), sigma g 2 is the variance of Gaussian distribution.

[0011] Let the receiving signal y l at the receiving end base station be represented as:

[0012]

[0013] wherein, respectively represent the channel coefficients from the wireless signal source to the receiving end, n d is the receiving end additive Gaussian white noise.

[0014] Let the receiving signal y e at the illegal eavesdropping end be represented as:

[0015]

[0016] where, respectively represent the channel coefficient from the wireless signal source to the illegal eavesdropping end, n d is the additive white Gaussian noise at the receiving end, n a is the artificial noise sent by the base station to the eavesdropper, which will be identified and removed at the legal receiving end.

[0017] At the receiving end, the successive interference cancellation technology is used, and the decoding order is determined according to the signal power. The signal-to-interference-and-noise ratio of the first user is:

[0018]

[0019] The signal-to-interference-and-noise ratio of the second user is:

[0020]

[0021] Similarly, the signal-to-interference-and-noise ratio of the nth user is:

[0022]

[0023] where, is the noise power at the legal receiving end. For convenience of representation, the user number in the above signal-to-interference-and-noise ratio formula is determined according to the decoding order, and the signal number of the first decoding is small, which has nothing to do with the sequence number in the signal representation at the receiving end.

[0024] Considering the illegal eavesdropping end, it is assumed that the eavesdropping ability of the eavesdropping end is strong, and it can distinguish different users and decode each user signal separately. At the same time, there are multiple eavesdropping ends in the model system, and it is assumed that there is an eavesdropping user with the strongest eavesdropping ability among the multiple eavesdropping ends. If the system can ensure the safety of information transmission when considering the strongest eavesdropping end, it can be represented that the system can perform safe information transmission under multiple eavesdropping ends. The case of the eavesdropping end with the strongest eavesdropping ability is considered below.

[0025] The signal-to-interference-and-noise ratio of the first user at the eavesdropping end is:

[0026]

[0027] The signal-to-interference-and-noise ratio of the second user at the eavesdropping end is:

[0028]

[0029] The signal-to-interference-and-noise ratio of the nth user at the eavesdropping end is:

[0030]

[0031] where, Noise power for the eavesdropping end.

[0032] In order to strengthen the physical layer security of the system, the security and rate of the system are optimized. According to the definition of security rate, the security rate of the signal is equal to the difference between the legal end rate and the illegal eavesdropping end rate:

[0033]

[0034]

[0035]

[0036] Where, R s is the legal end user rate, R e is the illegal eavesdropping end rate, [x] + = max{0, x}, when the calculation result is negative, the security rate is 0, that is, it is impossible to carry out safe and reliable communication.

[0037] The security and rate is defined as the sum of the security rates of all users in the system:

[0038]

[0039] The optimization objective under the NOMA uplink model is as follows:

[0040]

[0041]

[0042]

[0043] P min ≤ alpha i * P total ≤ P max

[0044] Where, P min , P max is the minimum transmission power and the maximum transmission power of the user in the system. The solution of the optimization function is a set of power allocation factors that maximizes the security and rate of the system. Therefore, the process of encrypting the physical layer security is converted into a power allocation method with the optimization objective of the security and rate of the system.

[0045] (II) Adopting meta-reinforcement learning network to encrypt the physical layer security of the system;

[0046] The specific steps are as follows:

[0047] S1, the meta-reinforcement learning network (i.e., the deep reinforcement learning network) adopts a DQN, DQN_target double network structure (Mnih V, Kavukcuoglu K, Silver D, et al. Playing Atari with Deep Reinforcement Learning [J]. arXiv e-prints, 2013.), two network structures are the same, and the action-behavior value function Q is realized by a fully connected layer network. The DQN network parameters are updated each time, while the DQN_target network is a target network, which is applied to the network after the final training is completed, and the parameters are updated by cloning the parameters of the DQN network every syn_num steps. The network parameters of the DQN network and the DQN_target network are randomly initialized. The initial parameters initial_param are set to θ, and the parameters to be updated temp_param are set to θ.

[0048] S2, since the power allocation is actually an optimization problem for continuous variables, i.e., the optimization problem shown in equation (12), it is necessary to determine the discretization scheme of the continuous action. The present application uses a coding method to discretize the action into three states of increase, decrease and no change of the user power allocation factor, and only changes the power allocation factors of two users each time. The smaller the unit size of increase and decrease, the smaller the error caused by quantization, but the computational complexity of model training increases, and the training time increases.

[0049] The coding discretization includes two discretization methods, one is to discretize the power value, and each action determines a fixed power value instead of determining the increase and decrease value of the power. The second is to determine the increase and decrease of the power value of all users at the moment for each user at the same time.

[0050] S3, the meta-reinforcement learning training task set is K groups of different channel distribution parameters, specifically wireless channel distributions subject to different standard deviations and expectations, and M (M

[0051] S3.1, initialize the sampled task environment, initialize the experience replay buffer. The corresponding DQN and DQN_target load the same parameters temp_param. Initialize the optimizer as the Adam optimizer, and the optimization parameter is the parameter of the DQN.

[0052] S3.2, perform training of episode rounds, each round resets the environment to get the initial state state1, when the training is not over (no done signal is received), according to the ε-greedy strategy decaying over time (A. Ray and H. Ray, "Proposing ε-greedy Reinforcement Learning Technique to Self-Optimize Memory Controllers," 2021 2nd International Conference on Secure Cyber Computing and Communications (ICSCCC), 2021, pp. 318-323), it is determined whether the current action is a randomly generated action or an action with the maximum q value output by the DQN network, and the action is marked as a1. And take the action into the environment to update the state to get state2, as well as the return r1 of the current action, and the mark done indicating whether the round is aborted, and store the obtained experience (state1, a1, r1, state2, done) in the experience replay buffer. In this way, (state n , a n , r n , state n+1 , done) is obtained until the minimum number of cache experiences is reached, and the next training can be started.

[0053] S3.3, randomly extract batch experience tuples from the experience cache according to batch_size, calculate the value of the loss function, and perform gradient back propagation. The formula of the loss function is as follows:

[0054]

[0055] Where r n is the reward of the current experience, γ is the discount factor, which is used to reduce the contribution of the next step to the overall learning direction, Q target (S n+1 ,a n+1 ) is the q value output by the target network for the next state, and Q(S n ,a n ) is the q value output by the current network in the current state.

[0056] S3.4, copy the network parameters of DQN to DQN_target every syn_num steps.

[0057] S3.5, K times of gradient descent are performed in each round, and the DQN_target parameter of the last obtained task i (i=1, 2, 3,..., M) is θ'.

[0058] S4, the meta-reinforcement learning network learning gradient update is performed, and the following formula is used:

[0059]

[0060] Wherein, ∈ is a learning update step, and in the learning of each task, it can be written as the update of the to-be-updated parameter variable temp_param, as follows:

[0061]

[0062] Advantages of the present application

[0063] In view of the shortcomings that the traditional deep reinforcement learning method has weak generalization ability and cannot adapt to dynamic environment, the learning characteristics of meta-learning are used, and a non-orthogonal multiple access system physical layer security communication method and device based on meta-reinforcement learning are proposed, which is close to the actual situation and has more application value. BRIEF DESCRIPTION OF DRAWINGS

[0064] Figure 1 It is a non-orthogonal multiple access wireless communication model diagram.

[0065] Figure 2 It is an algorithm flowchart of the present application.

[0066] Figure 3 It is a NOMA system security and rate comparison diagram under different power allocation methods in the embodiment of the present application. PREFERRED EMBODIMENT

[0067] The technical scheme of the present application is demonstrated through specific implementation cases.

[0068] The case environment is set as wireless communication under a non-orthogonal multiple access system, and physical layer security is considered, and the target function is to optimize the security and rate of the system. There are 4 users in the system. First, the preparation link, according to the Rayleigh distribution of the wireless channel, a plurality of groups of different independent user-base station, user-eavesdropping, base station-eavesdropping channel coefficients are generated and put into the meta-training task set U.

[0069] S1, randomly initialize the network parameters of the DQN network and the DQN_target network. The network structure of the DQN is a fully connected layer neural network, which has a total of 5 layers of structure, including 1 input layer, 1 output layer, and 3 hidden layers, and the input and output sizes of each layer are [input_size, 128], [128, 256], [256, 256], [256, 128], and [128, n_actions]. Among them, input_size is the dimension of the input state, and n_actions is the number of discretized actions.

[0070] S2, determine the continuous action discretization scheme. The unit size of the power allocation factor is 0.01, that is, if the action is to increase the power of user 1, the power allocation factor of user 1 is increased by 0.01, and if the action is to reduce the power of user 1, the power allocation factor of user 1 is reduced by 0.01. Since it is a 4-user system, according to the rule that each action only increases or decreases the power of 2 users, the total number of actions is 12.

[0071] S3, randomly sample M=4 tasks from the meta-training task set U.

[0072] S3.1, initialize the sampled task environment and initialize the experience replay buffer. The corresponding DQN and DQN_target are loaded with the same parameters temp_param. The optimizer is initialized as an Adam optimizer, and the optimization parameter is the parameter of the DQN.

[0073] S3.2, perform 20 rounds of training. Each round of training resets the environment to obtain an initial state state1. When the training is not completed (no done signal is received), according to the ε-greedy strategy with time decay, the initial ε is 1.0, the decay coefficient is 0.9975, and the minimum ε is 0.05. Determine whether the current action is a randomly generated action or an action with the maximum q value output by the DQN network, and mark the action as a1. Then, the action is input into the environment to update the state to obtain state2, select the return r1 of the current action, and select the mark done indicating whether the round is terminated. The obtained experience (state1, a1, r1, state2, done) is stored in the experience replay buffer. In this way, (state n ,a n ,r n ,state n+1 ,done) is obtained, and the next training can be started until the minimum cache experience number is reached. The minimum cache experience number is 100000.

[0074] S3.3, Randomly sample batch_size=100 experience tuples from the experience buffer, compute the value of the loss function, and perform gradient backpropagation. Discount factor γ=0.9.

[0075] S3.4, Copy the network parameters of DQN to DQN_target every interval syn_num=20 steps.

[0076] S3.5, Perform 12 gradient descent steps per episode, and the final DQN_target parameters of task i (i=1, 2, 3,..., M) are θ'.

[0077] S4, Perform meta-learning gradient update, as follows:

[0078]

[0079] Where ∈ is the meta-learning update step, and it can be written as the update of the to-be-updated parameter variable temp_param in the learning of each task, as follows:

[0080]

[0081] Where the step of meta-learning update ∈=0.1.

[0082] The simulation results show that the power allocation method of the non-orthogonal multiple access system based on meta-reinforcement learning can achieve good physical layer security encryption effect.

[0083] The above-described embodiments are only to better illustrate the methods and devices proposed by the present application, so as to help the reader better understand the principles of the present application. The embodiments and parameter settings should be understood as the protection scope of the present application is not limited to such specific examples and embodiments. Those skilled in the art can make other various specific modifications and combinations of the above disclosed technologies without departing from the essential scope of the present application, and these modifications and combinations still belong to the protection scope of the present application.

Claims

1.A method for physical layer security communication in a non-orthogonal multiple access system based on meta-reinforcement learning, characterized in that, The power allocation optimization objective function maximizing system physical layer security and transmission rate is constructed The power allocation optimization objective function maximizing system physical layer security and transmission rate is constructed The non-orthogonal multiple access system includes a wireless system sending end user, a receiving end base station, and a malicious eavesdropping end Let the ith user at the transmitting end send a signal S i is represented as: Among them, P total α represents the total transmit power of the transmitter. i Let X be the power allocation factor for the i-th user. i Let h be the information signal of the i-th user; i = 1, 2...n, where n is the number of users; the channels between the user and the base station, the user and the eavesdropping terminal, and the base station and the eavesdropping terminal are represented by channel coefficients: h sd h se h de The channel coefficients are random variables that follow a Rayleigh distribution; Let y be the received signal at the receiving end base station l is represented as: wherein, respectively represent channel coefficients from the wireless signal source to the receiving end, n d is the additive white Gaussian noise at the receiving end; Let y be the received signal at an illegal eavesdropping terminal e is represented as: wherein, respectively represent the channel coefficients from the wireless signal source to the illegal eavesdropping end, n d is the additive white Gaussian noise at the receiving end, n a is the artificial noise transmitted at the base station to interfere with the eavesdropper; At the receiving end, the continuous interference cancellation technology is used, and the decoding order is determined according to the signal power size, the signal-to-interference-and-noise ratio of the first user is The signal-to-interference-and-noise ratio of the second user is Similarly, the signal-to-interference-and-noise ratio of the nth user is wherein Pnoise is the noise power of the legitimate receiver It is assumed that the eavesdropping ability of the eavesdropping end is strong, and different users can be distinguished and decoded individually; meanwhile, there are multiple eavesdropping ends in the model system, and it is assumed that there is an eavesdropping user with the strongest eavesdropping ability among the multiple eavesdropping ends; if the system ensures the security of information transmission when considering the strongest eavesdropping end, it means that the system can perform secure information transmission under multiple eavesdropping ends; the case of the strongest eavesdropping end is considered below The signal-to-interference-and-noise ratio of the first user of the eavesdropping end is The signal-to-interference-and-noise ratio of the second user of the eavesdropping end is The signal-to-interference-and-noise ratio of the nth user of the eavesdropping end is wherein is the noise power at the eavesdropping end; In order to strengthen the physical layer security of the system, the security and rate of the system are used as the optimization objective, according to the definition of security rate, the security rate of the signal is equal to the difference between the legal end rate and the illegal eavesdropping end rate Wherein, R s is the legal end user rate, R e is the illegal eavesdropping end rate, [x] + = max{0, x}, when the calculation result is negative, the security rate is 0, that is, it is impossible to carry out safe and reliable communication; The security and rate is defined as the sum of the security rates of all users in the system Therefore, the optimization objective function under the NOMA uplink model is as follows P min ≤α i *P total ≤P max where P min , P max are the minimum and maximum transmit power of the users in the system; the solution of the optimization objective function is a set of power allocation factors that maximizes the rate and makes the system safe. The meta-reinforcement learning network is used to encrypt the physical layer of the system The specific steps are as follows S1, the meta-reinforcement learning network adopts the DQN and DQN_target double network structure, the two network structures are the same, and the action-behavior value function Q is realized by using the full connection layer network; the DQN network parameters are updated every time, while the DQN_target network is the target network, which is the network applied after the final training is completed, and the parameters are updated by cloning the parameters of the DQN network every syn_num steps; the network parameters of the DQN network and the DQN_target network are randomly initialized; the initialization parameter is θ, and the parameter to be updated is θ S2, for the optimization problem shown in equation (12), the continuous action is discretized, and three states of increasing, decreasing, and unchanged of the coded discrete action to the user power allocation factor are used S3, the meta-reinforcement learning training task set is K different channel distribution parameters, specifically wireless channel distributions subject to different standard deviations and expectations, M different channel environments (M S3.1, initialize the sampled task environment and initialize the experience replay buffer; the corresponding DQN and DQN_target load the same parameters temp_param; the optimizer is initialized as the Adam optimizer, and the optimization parameter is the parameter of the DQN S3.2, training is performed in episode rounds, each round resets the environment to obtain an initial state state1, when the training is not finished, according to the ε-greedy strategy decaying over time, it is determined whether the current action is a randomly generated action or an action with the maximum output q value according to the DQN network, and the action is marked as a1; and the action is brought into the environment to update the state to obtain state2; the return r1 of the current action and the mark indicating whether the round is terminated are selected, and the obtained (state1, a1, r1, state2, done) is stored in the experience replay buffer, and the above steps are repeated to obtain (state n ,a n ,r n ,state n+1 (,done), until the minimum cache experience number is reached, then the next training begins; S3.3, a batch of experience tuples is randomly extracted from the experience buffer according to batch_size, the value of the loss function is calculated, and gradient back propagation is performed; the formula of the loss function is as follows: Where, r n As a reward for current experience, γ is a discount factor used to reduce the contribution of the next step to the overall learning direction. target (S n+1 ,a n+1 Let Q(S) be the q-value of the target network's output for the next state. n ,a n This is the q-value output by the current network in its current state; S3.4, the network parameters of DQN are copied to DQN_target network every syn_num steps; S3.5, K times of gradient descent are performed in each round, and the final DQN_target parameters of the task i (i = 1, 2, 3,..., M) are θ'; S4, the meta-reinforcement learning network learning gradient update is performed, and the following formula is used: wherein, ∈ is the learning update step, and in the learning of each task, the update of the to-be-updated parameter variable temp_param can be written as follows:

Citation Information

Patent Citations

  • Dynamic spectrum access method based on layered deep reinforcement learning in NOMA system

    CN113207127A

  • Physical layer security and rate maximization method

    CN114124171A