High-energy-efficiency concealed wireless communication method based on DDPG algorithm
By applying the DDPG algorithm training intelligence body in the downlink NOMA system, the transmission strategy is automatically adjusted to improve energy efficiency and concealment, solving the challenge of achieving high-energy-efficient concealed communication when only the current time slot channel state information is available.
Patent Information
- Application Number
- CN202510334575.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-10
AI Technical Summary
In the downlink NOMA system, how to realize high-efficiency hidden communication through the DRL algorithm without only the current time slot channel state information is solved by the challenge of traditional methods in dealing with non-convex and difficult-to-solve optimization problems.
The high-energy-efficient hidden wireless communication method based on the DDPG algorithm is adopted, and the amount of information sent by the base station to the user is controlled to improve the energy efficiency and concealment of the system by training the intelligence body to automatically adjust the transmission strategy according to the channel state information, and control the amount of information sent by the base station to the user to improve the energy efficiency and concealment of the system.
It achieves the goal of improving the energy efficiency of the system while ensuring hidden performance. It is suitable for secret message transmission under limited energy conditions, and has higher energy efficiency and concealment compared with traditional Greedy algorithms.
Smart Images

Figure CN120128913A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless network information secure transmission, and particularly relates to an energy-efficient covert wireless communication method based on the DDPG algorithm. Background Art
[0002] Non-orthogonal multiple access (NOMA) has been widely studied in 5G and B5G wireless communications. Its use can significantly improve the spectral efficiency of wireless communication systems and increase the number of user connections. However, the broadcast nature of wireless transmission brings the risk of the transmitted information being eavesdropped without the user's permission. Although using traditional security means, such as encryption technology or physical layer security technology, can increase the security of information content, it may still arouse the suspicion of opponents and increase the intensity of detecting the secret message being transmitted when transmitting secret messages.
[0003] Since Bash et al. proposed the square root law in 2013, the feasibility of enhancing the security of user communications using covert transmission technology has attracted extensive research. Currently, the research on covert transmission in NOMA communication systems mainly focuses on the trade-off between covertness, effectiveness, and reliability. There is relatively little research on the problem of energy-efficient covert communication.
[0004] The literature "On the design of secure non-orthogonal multiple access systems, IEEE Journal on Selected Areas in Communications, vol. 35, no. 10, pp. 2196-2206, 2017" and "Energy-efficient resource allocation for secure NOMA-enabled mobile edge computing networks, IEEE Transactions on Communications, vol. 68, no. 1, pp. 493-505, 2019." proposed methods to improve the EE of NOMA systems under secrecy constraints. Currently, the research on energy-saving covert transmission in downlink NOMA systems mainly involves solving optimization problems to find closed-form or approximate solutions that ensure the proposed transmission scheme satisfies both the energy efficiency (EE) and covertness requirements simultaneously. However, the process of deriving these solutions often faces challenges in dealing with non-convex and intractable problems.
[0005] The emergence and development of Deep Reinforcement Learning (DRL) provide a new approach to solving these challenges. For example, the literature "Learning-based Power Control for Secure Covert Semantic Communication, arXiv preprint arXiv:2407.07475, 2024." explores the feasibility of applying DRL algorithms to achieve semantic covert communication under energy-constrained conditions.
[0006] Unfortunately, there is a lack of research on how DRL algorithms can achieve energy-efficient covert communication in a downlink NOMA system when only the current time-slot Channel State Information (CSI) is known. Therefore, it is necessary to study the potential of using the DDPG algorithm in the DRL algorithm to achieve energy-efficient covert communication in a downlink NOMA link system. Summary of the Invention
[0007] The object of the present invention is to solve the problems raised in the background technology and propose an energy-efficient covert wireless communication method based on the DDPG algorithm.
[0008] To achieve the object of the present invention, the present invention discloses an energy-efficient covert wireless communication method based on the DDPG algorithm, including a base station Alice, a common user Carl, a covert user Bob, and a detecting user Willie. Carl is designated as the remote user and Bob as the proximal user within the system; each node in the system uses a single antenna and operates in a time-division duplex mode; in the model, Alice needs to transmit a fixed total amount of data to Carl and Bob respectively within N time slots, and Willie uses power detection to check whether Alice is sending private information to Bob; in order to improve the energy efficiency of the system and hide the secret message transmitted by Alice, the DDPG algorithm is used to train an agent that automatically adjusts the transmission strategy according to the Channel State Information (CSI) to control the amount of information sent by Alcie to Carl and Bob each time.
[0009] Furthermore, an energy-efficient covert wireless communication method based on the DDPG algorithm specifically includes the following steps:
[0010] Step 1: Design an energy-efficient covert communication algorithm framework based on DDPG and train the agent; first, design the reward function, and further design the network parameters of the actor neural network, the critic neural network, the actor target neural network, and the critic target neural network; after a preset number of episode training, obtain the agent.
[0011] Step 2, CSI acquisition: The user periodically sends pilot signals to the base station so that the base station can estimate the CSI of the uplink. The base station uses the reciprocity principle to convert the uplink CSI into downlink CSI;
[0012] Step 3, power allocation and resource scheduling: According to the CSI of the current time slot, Alice inputs the CSI, the amount of data that Alice still needs to transmit, and the number of time slots that have passed into the agent. The agent outputs the number of messages c 1,i and c 2,i that Alice needs to send to Carl and Bob in the current time slot, so as to allocate appropriate transmission power P 1,i and P 2,i ;
[0013] Step 4, information encoding and signal superposition: Encode the common message and the covert message and superpose them according to the preset transmission power;
[0014] Step 5, information transmission: Alcie sends the common information and the secret message with the superposed power of transmission power P 1,i and P 2,i ; And where P max represents the maximum transmission power of Alice;
[0015] Step 6, information decoding: Carl regards the covert message in the received signal as interference and directly decodes it; Bob needs to use the serial interference cancellation SIC (Successive Interference Cancellation) technology to first cancel the common signal and then decode the secret message.
[0016] Furthermore, in step 1, Alice effectively trains the agent according to the energy-efficient covert communication algorithm framework based on DDPG, so that the agent can give a transmission strategy according to the CSI. Specifically, it includes: Step 1-1, the energy-efficient covert communication algorithm framework based on DDPG defines five basic elements - state, action, transition probability, reward, and discount factor; Step 1-2, update the algorithm framework, network design, and network parameters.
[0017] Furthermore, step 1-1 is specifically as follows:
[0018] The state set S is defined as the set of measurement vectors K i =[κ 1,i ,κ 2,i , which characterizes the observed downlink NOMA link system environment conditions in N time slots; the state κ k,i is composed of the CSI γ k,i =|h k,i| 2 and the amount of data z that Alice still needs to transmit to the k-th user at this time slot k,i constitute;
[0019] The action set B represents the action vector M executed by Alice within N time slots i =[c 1,i , c 2,i ; c k,i represents the amount of information transmitted by Alice to the k-th user and is also the transmission strategy in the i-th time slot;
[0020] The transition probability set T consists of the transition probability vector H for each time slot i =[ρ 1,i , ρ 2,i ; The transition probability ρ k,i represents the probability that when Alice communicates with the k-th user in the i-th time slot, the state κ k,i transitions to κ k,i+1 ; If the difference between the remaining amount of information z k,i transmitted by Alice to the k-th user and the remaining amount of information z k,i+1 in the (i + 1)-th time slot is equal to the transmission strategy c k,i in the i-th time slot, then ρ k,i = 1; otherwise, ρ k,i = 0;
[0021] The reward set R consists of the reward function values r i received by Alice after making a decision in each time slot; The reward function value r i is calculated using the designed reward function, which takes into account energy efficiency, covert performance, and the constraints of transmission power; Specifically, behaviors that meet the constraint conditions will receive positive rewards, while violations of the constraint conditions will be punished; The reward function in the i-th time slot is expressed as
[0022]
[0023] where, represents Willie's minimum error detection probability, P 1,i and P 2,i represent the appropriate transmission power allocated by Alice, represents the average amount of covert message transmission in the i-th time slot, represents the amount of data transmitted per joule in the i-th time slot, χ i represents the number of time slots when Alice and Bob have not communicated before the current time slot, represents the number of time slots when Alice does not communicate, A jThe coefficient representing the magnitude of the control value and the sign of the control item, where j = 1, 2, …, 6; the reward function coefficient is expressed as
[0024] Table 3 Reward Function Coefficient Design
[0025]
[0026] The discount factor is a key hyperparameter in the DDPG algorithm, which is used to adjust the influence degree of future rewards on the current decision-making; the value range of the discount factor is (0, 1], which directly affects the degree of importance that the agent attaches to future rewards.
[0027] Furthermore, Step 1-2 is specifically as follows:
[0028] The high-energy-efficient covert communication algorithm framework based on DDPG consists of a total of 4 neural networks, namely the actor neural network, the critic neural network, the actor target neural network corresponding to the actor neural network, and the critic target neural network corresponding to the critic neural network; in addition, during the training process of the agent, the algorithm uses a memory pool to store past experiences; this helps to break the correlation between data points, enhance the generalization ability of the algorithm, and stabilize the learning process.
[0029] In the high-energy-efficient covert communication algorithm framework based on DDPG, the two target neural networks have the same structure as their corresponding actor neural network and critic neural network; therefore, only the actor and critic neural networks will be introduced; at the i-th time slot, Alice observes a set of states K i , and these states are input into the actor neural network; through three fully connected layers, after passing through the Tanh activation function and the scaling layer, and then the network outputs a set of actions M i = π(K i |ω Actor ), where ω Actor represents the parameters of the actor neural network; the input of the critic neural network consists of two parts: one part is the state K i at the i-th time slot, and the other part is the action value output by the actor neural network; these two parts are processed through two fully connected layers and one fully connected layer respectively, and then connected and fed into the ReLU layer; after passing through another fully connected layer, the network outputs the action-state value function Q π (K i , M i |ω Actor , ω critic ), where ω critic represents the parameters of the critic neural network;
[0030] In the algorithm, the actor neural network updates its parameters based on the sampled policy gradient, expressed as
[0031]
[0032] Among them, Q(K j , M j | ω Actor , ω critc ) represents the state - action value function obtained by inputting the j - th group of samples drawn from the memory pool into the critic neural network. M j = π(K j | ω Actor ) represents the sampled action value;
[0033] In the algorithm, the critic neural network is updated based on the loss value function, and the loss value function is expressed as
[0034]
[0035] Expressed as
[0036]
[0037] Among them, is the target state - action value function obtained by inputting the j - th group of samples sampled from the memory pool into the critic target neural network;
[0038] Both the actor target neural network and the critic target neural network adopt a soft update mechanism to maintain the stability of network updates; specifically expressed as
[0039]
[0040] Among them, μ is the soft update factor, usually a small value close to zero, and the soft update factor controls the update rate of the target network to the main network;
[0041] Table 4 Training agent parameter design
[0042]
[0043] The training parameters are shown in Table 4.
[0044] Furthermore, in step 3, Alice performs dynamic power allocation and resource scheduling according to the CSI by the agent pre - trained by the DDPG algorithm, and completes the transmission of public messages and secret messages in N time slots, specifically including:
[0045] Power allocation and resource scheduling Alice selects the number of messages c 1,i and c 2,i sent to Carl and Bob according to the CSI of the current time slot, so as to allocate appropriate transmission powers P 1,i and P 2,i; The channel rate of each time slot when Alice sends information to Carl and Bob is expressed as
[0046]
[0047] where, |h 1,i | 2 and |h 2,i | 2 respectively represent the CSI of Alice to Carl and Bob in the i-th time slot, and respectively represent the variances of the Gaussian white noise of the channels between Alice and Carl and between Alice and Bob; The CSI of the i-th time slot between Alice and Willie is expressed as |h 3,i | 2 , and the variance of the Gaussian white noise is expressed as
[0048] In each time slot, all channels experience independent and quasi-static Rayleigh fading, that is, g k,i ~CN(0,λ 2 ), so there is
[0049]
[0050] Considering the distances between Alcie and Carl, Bob and Willie are d 1 , d 2 and d 3 , so |h k,i | 2 obeys an exponential distribution with parameter ; where, η is the path loss exponent, k∈{1,2,3}; Bob needs to use the successive interference cancellation (SIC) technique to eliminate the influence of the common information in order to correctly decode the private message. Therefore, it is necessary to ensure that |h 2,i | 2 ≥|h 1,i | 2 , that is, d 1 ≥(a / 1-a) 1 / η d 2 , where a represents the probability of |h 2,i | 2 ≥|h 1,i | 2 , a→1.
[0051] Furthermore, in step 4, Alice superimposes the powers allocated to Carl and Bob, hiding the secret message with the common information transmitted by Alcie while improving the utilization efficiency of communication resources.
[0052] Further, in step 5, the information transmission Alice transmits the common information and the secret message with the superimposed power of transmit power P 1,i and P 2,i ; Considering that Alice's power cannot be infinitely large, the condition needs to be satisfied, where P max represents Alice's maximum transmit power; At the same time, the covert constraint for Alice to send the secret message also needs to be considered, that is, it is required that Willie can obtain the transmit power P 1,i and P 2,i and use the optimal detection threshold for the minimum error detection probability to be greater than 1 - ε, where ε represents the covert communication tolerance value and ε is an arbitrarily small real number.
[0053] Further, in step 6, the information decoding Carl treats the covert message in the received signal as interference and directly decodes it; Bob needs to use serial interference cancellation (SIC) to first cancel the common signal and then decode the secret message; Willie uses power for detection and analysis, expressed as where D 0 and D 1 respectively represent that Willie detects that Alice has not sent the secret message and has sent the secret message after detection, and represents the average of the signals received by Willie, u represents the channel use and u → ∞; Willie's optimal detection threshold is expressed as
[0054]
[0055] The minimum detection probability is expressed as
[0056]
[0057] Compared with the prior art, the significant progress of the present invention lies in that: when the high - energy - efficient covert wireless communication method based on the DDPG algorithm of the present invention is specifically operated, Alice will select and allocate power according to the transmission strategy selected by the trained agent, and hide the secret message to be transmitted through the transmitted public information to increase the uncertainty of the signal received by Willie. The core idea of the algorithm for training the agent in the present invention is the water - filling algorithm idea, that is, selecting a suitable transmission strategy according to the CSI to ensure the improvement of the overall EE. This makes the method proposed in the present invention not only able to ensure the covert performance of the system but also take into account the improvement of the system EE. Compared with the traditional Greedy fixed - ratio and always transmitting messages with the maximum transmission power transmission scheme, the present invention can dynamically adjust the transmission strategy according to the CSI, which is more suitable for the transmission of secret messages in the case of limited energy.
[0058] To more clearly illustrate the functional characteristics and structural parameters of the present invention, the following further explains in conjunction with the drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0060] Figure 1 is a schematic diagram of the communication system model of a high - energy - efficient covert wireless communication method based on the DDPG algorithm designed by the present invention;
[0061] Figure 2 is a flowchart of training the agent of the high - energy - efficient covert communication algorithm based on DDPG;
[0062] Figure 3 is a comparison schematic diagram of the average total transmission power between the algorithm proposed in the present invention and the traditional greedy algorithm;
[0063] Figure 4 is a comparison schematic diagram of the balanced power between the algorithm proposed in the present invention and the traditional greedy algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0065] As Figure 1The downlink NOMA high-energy-efficiency covert communication system shown includes a base station (Alice), a common user (Carl), a covert user (Bob), and a warden (Willie). In the system, Carl is designated as the far-end user and Bob as the near-end user. Each node in the system uses a single antenna and operates in the time-division duplex mode. In this model, Alice needs to transmit a fixed total amount of data to Carl and Bob respectively within N time slots, and Willie uses power detection to determine whether Alice is sending private information to Bob. To improve the system's energy efficiency and hide the secret message transmitted by Alice, based on the idea of the water-filling method, a DDPG algorithm is used to train an agent that can automatically adjust the transmission strategy according to CSI to control the amount of information sent by Alcie to Carl and Bob each time. The implementation process of the transmission method of the present invention is as follows:
[0066] 1) Design a high-energy-efficiency covert communication algorithm framework based on DDPG and train the agent. The high-energy-efficiency covert communication algorithm framework based on DDPG proposed in the present invention is specifically as Figure 2 shown. This algorithm needs to define five basic elements - state, action, transition probability, reward, and discount factor.
[0067] The state set S is defined as the set of measurement vectors K i =[κ 1,i ,κ 2,i , representing the environmental conditions of the downlink NOMA link system observed in N time slots. The state κ k,i consists of the CSI γ k,i =|h k,i | 2 between Alice and the k-th user in the i-th time slot and the amount of data z k,i that Alice still needs to transmit to the k-th user at this time slot.
[0068] The action set B represents all action vectors M i =[c 1,i ,c 2,i executed by Alice within N time slots. c k,i represents the amount of information transmitted by Alice to the k-th user and represents her transmission strategy in the i-th time slot.
[0069] The transition probability set T consists of the transition probability vectors H i =[ρ 1,i ,ρ 2,i for each time slot. The transition probability ρ k,i represents that when Alice communicates with the k-th user in the i-th time slot, the state κ k,i transfers to κ k,i+1The probability. If the remaining information z transmitted by Alice to the k-th user k,i and the remaining information z in the (i + 1)-th time slot k,i+1 The difference is equal to the transmission strategy c in the i-th time slot k,i , then ρ k,i = 1. Otherwise, ρ k,i = 0.
[0070] The reward set R consists of the reward function values r received by Alice after making decisions in each time slot i . The reward function value r i is calculated using the reward function designed by the present invention, which mainly considers energy efficiency, covert performance, and the constraint of transmission power. Specifically, behaviors that satisfy the constraint conditions will receive positive rewards, while violations of the constraint conditions will be punished. The reward function in the i-th time slot can be expressed as
[0071]
[0072] where ξ i * represents Willie's minimum error detection probability, P 1,i and P 2,i represent the appropriate transmission power allocated by Alice, represents the average amount of covert message transmission in the i-th time slot, represents the amount of data transmitted per joule in the i-th time slot, χ i represents the number of time slots when Alice and Bob have not communicated before the current time slot, represents the number of time slots when Alice does not communicate, A j represents the coefficient that controls the magnitude of the value and the sign of the control term, where j = 1, 2,..., 6. The reward function coefficients designed by the present invention can be expressed as
[0073] Table 5 Reward Function Coefficient Design
[0074]
[0075] The discount factor is a key hyperparameter in the DDPG algorithm, which is used to adjust the influence degree of future rewards on the current decision. Its value range is (0, 1], and it directly affects the agent's emphasis on future rewards.
[0076] In order to effectively train the agent using the high-energy-efficiency covert communication algorithm framework based on DDPG, three key areas must be concerned: the algorithm framework, network design, and update of network parameters.
[0077] The high - energy - efficient covert communication algorithm framework based on DDPG consists of a total of 4 neural networks. These four neural networks are the actor neural network, the critic neural network, the actor target neural network corresponding to the actor neural network, and the critic target neural network corresponding to the critic neural network. In addition, during the training process of the agent, the algorithm uses a memory pool to store past experiences. This helps to break the correlation between data points, enhance the generalization ability of the algorithm, and stabilize the learning process.
[0078] In the high - energy - efficient covert communication algorithm framework based on DDPG, the two target neural networks have the same structure as their corresponding actor neural network and critic neural network. Therefore, only the actor and critic neural networks will be introduced. At the i - th time slot, Alice observes a set of states K i , and these states are input into the actor neural network. Through three fully - connected layers, after passing through the Tanh activation function and the scaling layer, the network then outputs a set of actions M i = π(K i |ω Actor ), where ω Actor represents the parameters of the actor neural network. The input of the critic neural network mainly consists of two parts: one part is the state K i at the i - th time slot, and the other part is the action value output by the actor neural network. These two parts are processed through two fully - connected layers and one fully - connected layer respectively, and then connected and fed into the ReLU layer. After passing through another fully - connected layer, the network outputs the action - state value function Q π (K i , M i |ω Actor , ω critic ), where ω critic represents the parameters of the critic neural network.
[0079] In this algorithm, the actor neural network updates its parameters based on the sampled policy gradient, which can be expressed as
[0080]
[0081] where Q(K j , M j |ω Actor , ω critc ) represents the state - action value function obtained by inputting the j - th group of samples drawn from the memory pool into the critic neural network, and M j = π(K j |ω Actor ) represents the sampled action value.
[0082] In this algorithm, the critic neural network is updated based on the loss value function, and the loss value function can be expressed as
[0083]
[0084] can be expressed as
[0085]
[0086] wherein is the target state-action value function obtained by inputting the j-th group of samples sampled from the memory pool into the critic target neural network.
[0087] Both the actor target neural network and the critic target neural network adopt a soft update mechanism to maintain the stability of network updates. Specifically expressed as
[0088]
[0089] where μ is the soft update factor, usually a small value close to zero, which controls the update rate of the target network towards the main network.
[0090] The training parameters are shown in the following table
[0091] Table 6 Training Agent Parameter Design
[0092]
[0093] 2) CSI acquisition: The user periodically sends pilot signals to the base station so that the base station can estimate the uplink CSI. The base station uses the reciprocity principle to convert the uplink CSI into downlink CSI.
[0094] 3) Power allocation and resource scheduling: Alice selects the number of information c 1,i and c 2,i sent to Carl and Bob according to the CSI of the current time slot, so as to allocate appropriate transmit powers P 1,i and P 2,i . The channel rate of each time slot when Alice sends information to Carl and Bob can be expressed as
[0095]
[0096] where |h 1,i | 2 and |h 2,i | 2 respectively represent the CSI between Alice and Carl and Bob in the i-th time slot, and respectively represent the Gaussian white noise variances of the channels between Alice and Carl and Bob. The CSI between Alice and Willie in the i-th time slot can be expressed as |h3,i | 2 , the Gaussian white noise variance can be expressed as
[0097] It is worth mentioning that in each time slot, all channels experience independent, quasi-static Rayleigh fading, i.e., g k,i ~CN(0,λ 2 ), so there is
[0098]
[0099] Considering that the distance between Alcie and Carl, Bob and Willie is d 1 ,d 2 and d 3 , so |h k,i | 2 The obedience parameters are where η is the path loss exponent, k∈{1,2,3}. Bob needs to use the successive interference cancellation (SIC) technique to eliminate the impact of public information in order to correctly decode the private message, so it is necessary to ensure that |h 2,i | 2 ≥|h 1,i | 2 , that is, d 1 ≥(a / 1-a) 1 / η d 2 , where a represents |h 2,i | 2 ≥|h 1,i | 2 The probability that a→1.
[0100] 4) Information Coding and Signal Superposition The public message and the covert message are encoded and superimposed according to the previously set transmission power.
[0101] 5) Information transmission Alcie with transmission power P 1,i and P 2,i The superimposed power sends public information and secret messages. Considering that Alice's power cannot be infinite, the condition must be met here Where P max represents Alice’s maximum transmission power. At the same time, we also need to consider the concealment constraint of Alice sending secret messages, that is, it requires Willie to obtain the transmission power P 1,i and P 2,i Use the optimal detection threshold in the case The minimum error detection probability To be greater than 1-ε, where ε represents the covert communication tolerance value, which is an arbitrarily small real number.
[0102] 6) Information decoding Carl regards the hidden message in the received signal as interference and directly decodes it. Bob needs to use Successive Interference Cancellation (SIC) to first cancel the common signal and then decode the secret message. Willie uses power for detection and analysis, which can be expressed as
[0103]
[0104] where D 0 and D 1 respectively represent that Willie detects that Alice has not sent the secret message and has sent the secret message, and represents the average of the signals received by Willie, u represents the channel usage and u→∞. The optimal detection threshold of Willie can be expressed as
[0105]
[0106] Thus, the minimum detection probability can be expressed as:
[0107]
[0108] The simulation diagram of the average total power varying with SNR in the method of the present invention is as shown in Figure 3 . Among them, the bandwidth is set to the unit bandwidth 1 (Hz), the number of time slots N = 20, Φ 1 = 150 (bits / Hz), Φ 2 = 3 (bits / Hz). To ensure that Bob can correctly decode using the SIC technology, the distance from Alice to Carl is set to d 1 = d 2 + 700 (m), where d 2 (m) is the distance from Alice to Bob. In addition, the distance from Alice to Willie is set to d 3 = d 2 (m), the path loss exponent is set to η = 3.2, the variance of the channel fading coefficient is set to λ 2 = 1, the maximum transmission power is set to P max = 0.001 (W), and the requirement for covert communication is set to ε = 0.1. It should be noted that considering the influence of noise, here SNR is used to measure it, which can be expressed as
[0109]
[0110] where From Figure 3It can be seen that regardless of the distance, under the same SNR condition, the method of the present invention can greatly improve the EE efficiency of the system compared with the traditional Greedy algorithm.
[0111] The schematic diagram of the balanced power (the ratio of the average total power to the probability of successful transmission and undetected) of the method of the present invention and the traditional Greedy algorithm changing with SNR is as Figure 4 shown, where the bandwidth is set to 1 (Hz), the number of time slots N = 20, Φ 1 = 150 (bits / Hz), Φ 2 = 3 (bits / Hz). To ensure that Bob can correctly decode using SIC technology, the distance from Alice to Carl is set to d 1 = 800 (m), and the distance from Alice to Bob is d 2 = 100 (m). In addition, the distance from Alice to Willie is set to d 3 = 100 (m), the path loss exponent is set to η = 3.2, the variance of the channel fading coefficient is set to λ 2 = 1, the maximum transmission power is set to P max = 0.001 (W), and the requirement for covert communication is set to ε = 0.1. It should be noted that considering the influence of noise, SNR is used to measure it here, specifically as shown in (49), where, From Figure 4 it can be seen that within a certain range of SNR, the performance of the method of the present invention in terms of EE, Quality of Service (QoS), and covertness is better than that of the traditional Greedy algorithm.
[0112] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0113] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A high energy efficiency covert wireless communication method based on DDPG algorithm, characterized in that: It includes base station Alice, public user Carl, hidden user Bob and detection user Willie. Carl is designated as the remote user and Bob as the near user in the system. Each node in the system uses a single antenna and works in time division duplex mode. In the model, Alice needs to transmit a fixed total amount of data to Carl and Bob respectively in N time slots, and Willie uses power to detect whether Alice is sending private information to Bob. The DDPG algorithm is used to train an agent that automatically adjusts the transmission strategy according to the channel state information CSI to control the amount of information Alcie sends to Carl and Bob each time.
2. According to claim 1, a high energy efficiency covert wireless communication method based on DDPG algorithm is characterized in that: The specific steps include: Step 1: Design a high-efficiency covert communication algorithm framework based on DDPG and train the agent; first design the reward function, and further design the network parameters of the actor neural network, critic neural network, actor target neural network, and critic target neural network; after a preset round of training, obtain the agent; Step 2: CSI acquisition: The user periodically sends a pilot signal to the base station so that the base station can estimate the uplink CSI. The base station uses the reciprocity principle to convert the uplink CSI into the downlink CSI. Step 3: Power allocation and resource scheduling; Alice inputs the CSI, the amount of data Alice needs to transmit, and the number of time slots that have passed into the agent based on the CSI of the current time slot. The agent then outputs the amount of information c that Alice needs to send to Carl and Bob in the current time slot. 1,i and c 2,i , thereby allocating appropriate transmission power P 1,i and P 2,i ; Step 4: Information encoding and signal superposition: Encode the public message and the covert message and superimpose them according to the preset transmission power; Step 5: Information transmission: Alcie transmits with a power of P 1,i and P 2,i The superimposed power transmits the public information and the secret message; and Where P max represents Alice's maximum transmit power; Step 6: Information decoding: Carl treats the hidden message in the received signal as interference and decodes it directly; Bob needs to use serial interference cancellation (SIC) technology to first eliminate the public signal and then decode the secret message.
3. The high energy efficiency covert wireless communication method based on DDPG algorithm according to claim 2, characterized in that: In step 1, Alice effectively trains the intelligent agent according to the high-efficiency covert communication algorithm framework based on DDPG, so that the intelligent agent can give a transmission strategy based on CSI, which specifically includes: step 1-1, the high-efficiency covert communication algorithm framework based on DDPG defines five basic elements - state, behavior, transition probability, reward and discount factor; step 1-2, updates the algorithm framework, network design and network parameters.
4. The high energy efficiency covert wireless communication method based on DDPG algorithm according to claim 3, characterized in that: Step 1-1 is as follows: The state set S is defined as the measurement vector K i =[κ 1,i ,κ 2,i ], representing the downlink NOMA link system environmental conditions observed in N time slots; state κ k,i The CSIγ between Alice and the kth user in the i-th time slot k,i =|h k,i | 2 and the amount of data z that Alice still needs to transmit to the kth user in this time slot k,i composition; Action set B represents the action vector M executed by Alice in N time slots i =[c 1,i ,c 2,i ];c k,i represents the amount of information transmitted by Alice to the kth user, which is also the transmission strategy in the i-th time slot; The transition probability set T consists of the transition probability vector H of each time slot i =[ρ 1,i ,ρ 2,i ] composition; transition probability ρ k,i It means that when Alice communicates with the kth user in the i-th time slot, the state κ k,i Transfer to κ k,i+1 The probability that Alice transmits the remaining information z to the kth user k,i and the residual information z of the (i+1)th time slot k,i+1 The difference is equal to the transmission strategy c of the i-th time slot k,i , then ρ k,i =1; otherwise ,r k,i =0; The reward set R is composed of the reward function value r received by Alice after making a decision in each time slot i Composition; reward function value r i The reward function is designed to calculate the energy efficiency, concealment performance and transmission power constraints. Specifically, behaviors that meet the constraints will be actively rewarded, while behaviors that violate the constraints will be punished. The reward function of the i-th time slot is expressed as in, represents Willie’s minimum false detection probability, P 1,i and P 2,i represents the appropriate transmit power allocated by Alice, represents the average amount of covert message transmission in the i-th time slot, represents the amount of data transmitted per joule in the i-th time slot, χ i represents the number of time slots before the current time slot when Alice and Bob had no communication, A represents the number of time slots when Alice does not communicate. j The coefficient representing the size of the control value and the positive and negative nature of the control item, where j = 1, 2, ..., 6; the reward function coefficient is expressed as Table 1 Reward function coefficient design The discount factor is a key hyperparameter in the DDPG algorithm, which is used to adjust the impact of future rewards on current decisions. The discount factor ranges from (0,1] and directly affects the importance the agent places on future rewards.
5. The high energy efficiency covert wireless communication method based on DDPG algorithm according to claim 3, characterized in that: Steps 1-2 are as follows: The energy-efficient covert communication algorithm framework based on DDPG contains four neural networks, namely the actor neural network, the critic neural network, the actor target neural network corresponding to the actor neural network, and the critic target neural network corresponding to the critic neural network. In addition, during the training process of the agent, the algorithm uses a memory pool to store past experience. In the energy-efficient covert communication algorithm framework based on DDPG, the two target neural networks have the same structure as their corresponding actor neural network and critic neural network; at the i-th time slot, Alice observes a set of states K i ,These states are fed into the actor neural network; Through three fully connected layers, after the Tanh activation function and the scaling layer, the network outputs a set of actions M i =π(K i |ω Actor ), where ω Actor represents the parameters of the actor neural network; the input of the critic neural network consists of two parts: one is the state K of the i-th time slot i , and the other part is the action value output by the actor neural network; these two parts are processed by two fully connected layers and one fully connected layer respectively, and then connected and fed to the ReLU layer; after passing through another fully connected layer, the network outputs the action state value function Q π (K i ,M i |ω Actor ,ω critic ), where ω critic represents the parameters of the critic neural network; In the algorithm, the actor neural network updates its parameters based on the sampled policy gradient, expressed as Among them, Q(K j ,M j |ω Actor ,ω critc ) represents the state action value function obtained by inputting the j-th group of samples extracted from the memory pool into the critic neural network, M j =π(K j |ω Actor ) represents the sampled action value; In the algorithm, the critic neural network is updated based on the loss function, which is expressed as Expressed as in, is the target state action value function obtained by inputting the j-th group of samples sampled from the memory pool into the critic target neural network; Both the actor target neural network and the critic target neural network use a soft update mechanism to maintain the stability of network updates; specifically, Among them, μ is the soft update factor, which is usually a small value close to zero. The soft update factor controls the rate at which the target network updates to the main network; Table 2. Design of training agent parameters The training parameters are shown in Table 2.
6. The high energy efficiency covert wireless communication method based on DDPG algorithm according to claim 2, characterized in that: In step 3, Alice uses the agent pre-trained by the DDPG algorithm to dynamically allocate power and schedule resources based on the CSI, and completes the transmission of public and secret messages in N time slots, including: Power Allocation and Resource Scheduling Alice selects the amount of information c to send to Carl and Bob based on the CSI of the current time slot. 1,i and c 2,i , thereby allocating appropriate transmission power P 1,i and P 2,i ; The channel rate for each time slot that Alice sends information to Carl and Bob is expressed as Among them, |h 1,i | 2 and |h 2,i | 2 They represent the CSI of Alice, Carl and Bob in the i-th time slot respectively, and They represent the Gaussian white noise variance of the channels between Alice and Carl and Bob respectively; the CSI of the i-th time slot between Alice and Willie is expressed as |h 3,i | 2 , the Gaussian white noise variance is expressed as In each time slot, all channels experience independent, quasi-static Rayleigh fading, i.e., g k,i ~CN(0,λ 2 ), so there is Considering that the distances between Alcie and Carl, Bob and Willie are d1, d2 and d3, then |h k,i | 2 The obedience parameters are exponential distribution; where η is the path loss exponent, k∈{1,2,3}; Bob needs to use the continuous interference cancellation SIC technique to eliminate the influence of public information in order to correctly decode the private message, so it is necessary to ensure |h 2,i | 2 ≥|h 1,i | 2 , that is, d1≥(a / 1-a) 1η d2, where a represents |h 2,i | 2 ≥|h 1,i | 2 The probability that a→1.
7. The high energy efficiency covert wireless communication method based on DDPG algorithm according to claim 2, characterized in that: In step 4, Alice adds the power allocated to Carl and Bob, improving the efficiency of communication resource utilization while using the public information transmitted by Alice to hide the secret message.
8. The high energy efficiency covert wireless communication method based on DDPG algorithm according to claim 2, characterized in that: In step 5, the information is transmitted Alcie with a transmission power P 1,i and P 2,i The superimposed power sends public information and secret messages; considering that Alice's power cannot be infinite, the condition needs to be met Where P max represents Alice’s maximum transmission power; at the same time, we also need to consider the concealment constraint of Alice sending secret messages, that is, requiring Willie to obtain the transmission power P 1,i and P 2,i Use the optimal detection threshold in the case The minimum error detection probability To be greater than 1-ε, where ε represents the covert communication tolerance value and ε is an arbitrarily small real number.
9. The high energy efficiency covert wireless communication method based on DDPG algorithm according to claim 2, characterized in that: In step 6, information decoder Carl treats the hidden message in the received signal as interference and decodes it directly; Bob needs to use serial interference cancellation SIC to eliminate the public signal first and then decode the secret message; Willie uses power for detection and analysis, which is expressed as Among them, D0 and D1 respectively represent that Willie believes Alice did not send a secret message and sent a secret message after detection, and represents Willie's average of the received signal, u represents the channel usage and u→∞; Willie's optimal detection threshold Expressed as The minimum detection probability is expressed as
Citation Information
Patent Citations
Intelligent adaptive power control method for large-scale OFDM (Orthogonal Frequency Division Multiplexing) system
CN114980293A
RSMA-based power distribution method, communication system, equipment and medium
CN118400804A
Hidden transmission method based on cooperative non-orthogonal multiple access communication system
CN118474770A
Information age optimization method and system of unmanned aerial vehicle assisted uplink hidden transmission system
CN119603709A
Full-duplex non-orthogonal multiple access-based transmit power control device employing deep reinforcement learning
US20240422693A1
Cited By
Three-dimensional deployment and power distribution joint optimization method for flight base station of unmanned aerial vehicle
CN113206701A
A three-dimensional deployment and power allocation joint optimization method of a UAV
CN113206701B