Terahertz communication security beam forming method and communication system
By using deep reinforcement learning techniques to construct state vectors and output beamforming vectors, the problems of high computational complexity and poor robustness in terahertz communication are solved. This achieves low-complexity, high-security-rate adaptive beam optimization, thereby improving the physical layer security performance of terahertz communication.
Patent Information
- Application Number
- CN202511774278.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-02-10
AI Technical Summary
In existing terahertz communication, secure beamforming methods have high computational complexity, poor robustness, and weak adaptability, making them unable to effectively cope with dynamically changing channel environments and attacks from intelligent eavesdroppers.
By employing deep reinforcement learning (DRL) technology, a state vector is constructed by observing the state of the intelligent agent system. The beamforming vector is output by the policy network and normalized in the power constraint layer. The network parameters are then updated in combination with the reward function to achieve adaptive beam optimization and adapt to the dynamic channel environment.
Without relying on precise channel information, we have achieved low-complexity, high-security-rate, and robust adaptive beam optimization, which improves the physical layer security performance of terahertz communication.
Smart Images

Figure CN121508595A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of communication, in particular to a terahertz communication security beamforming method and a communication system. BACKGROUND
[0002] With the rapid evolution of mobile communication technology, the sixth generation mobile communication system (6G) is generally considered to achieve the fusion communication target of ultra-high spectral efficiency, ultra-large bandwidth and extremely low latency. In this context, terahertz (THz) communication, with the high-frequency carrier characteristics of 0.1-10 THz, can provide a transmission rate of hundreds of Gbps, and is considered as one of the core supporting technologies of wireless links in the 6G era. Its high-frequency characteristics not only can significantly improve the spatial resolution and data throughput, but also can support short-distance ultra-high-speed wireless connections, such as indoor optical fiber replacement, inter-chip communication, satellite link enhancement, and other new application scenarios.
[0003] However, THz signals have physical characteristics of strong directivity and high propagation loss: due to the extremely short wavelength (usually less than 1 mm), the signal is extremely sensitive to factors such as occlusion, scattering, humidity, etc.; at the same time, the THz beam energy is concentrated, which makes the communication link extremely easy to be located and intercepted, forming a serious physical layer security risk. In the prior art, methods based on convex optimization (such as semi-definite relaxation SDR) or linear preprocessing (such as zero forcing ZF) are often used for security beamforming. However, these methods have the following inherent defects:
[0004] 1. High computational complexity: SDR and other methods have exponential growth in computational complexity when the antenna scale is large (such as large-scale MIMO), which is difficult to meet the real-time communication requirements.
[0005] 2. Poor robustness: It is severely dependent on accurate channel state information (CSI), and when there is channel estimation error, the security performance decreases sharply.
[0006] 3. Weak adaptability: It is a static optimization design and cannot effectively cope with dynamically changing channel environments and attacks by intelligent eavesdroppers.
[0007] Deep reinforcement learning (DRL) provides a new way to solve the above problems. However, there is no mature solution that can effectively apply DRL, especially algorithms suitable for continuous action space, to the security beamforming of terahertz large-scale MIMO systems, and systematically solve the robustness problem under imperfect CSI. SUMMARY
[0008] In view of the above defects and deficiencies existing in the prior art, the present application aims to provide a terahertz communication security beamforming method and a communication system, which aims to realize low-complexity, high-security-rate and strong-robust adaptive beam optimization without relying on accurate channel information, and effectively guarantee the physical layer security of terahertz communication.
[0009] The object of the present application can be achieved by the following technical solutions:
[0010] A terahertz communication security beamforming method applied to an intelligent agent, wherein the intelligent agent interacts with a communication environment including a terahertz base station, a legitimate user and a potential eavesdropper; the method comprises:
[0011] A state vector is constructed by observing a system state through the intelligent agent, wherein the system state comprises channel state information of the legitimate user and the potential eavesdropper;
[0012] A beamforming vector is output by a policy network Actor in the intelligent agent according to the state vector;
[0013] The beamforming vector is normalized by a power constraint layer to obtain a final beamforming vector, and signal transmission is performed;
[0014] A reward value is calculated according to the result after signal transmission, and the network parameters of the deep reinforcement learning intelligent agent are updated based on the reward value.
[0015] A further improvement of the present application is that the state vector is expressed as:
[0016]
[0017] Wherein: is a real part of channel estimation of the legitimate user, is an imaginary part of channel estimation of the legitimate user; is a real part of channel estimation of the potential eavesdropper; is an imaginary part of channel estimation of the potential eavesdropper; is a performance index of the legitimate user in the last time slot; is a performance index of the potential eavesdropper in the last time slot; is a secrecy rate in the last time slot; is system noise power.
[0018] A further improvement of the present application is that the unnormalized beamforming vector output by the policy network Actor is , wherein is the number of antennas; and the expression for normalizing the beamforming vector by the power constraint layer is:
[0019]
[0020] in: For the final beamforming vector; constant ensure And it is differentiable; It is the system's maximum transmission power; Is the agent in the first The dynamic transmit power budget available for each time slot.
[0021] A further improvement of the present invention lies in that, in the final beamforming vector Superimposed zero-mean OU noise And use the superimposed beamforming vector To transmit signals.
[0022] A further improvement of the present invention lies in the calculation of the reward value. The reward function used includes a robust term and a power constraint term. The expression for the reward function is as follows:
[0023]
[0024] Basic security rate The expression is:
[0025]
[0026] in: It is a constant; It uses beamforming vectors The signal-to-noise ratio received by legitimate users; It uses beamforming vectors The signal-to-noise ratio received by a potential eavesdropper;
[0027] Robust item The calculation expression is:
[0028]
[0029] in: ;
[0030] The expression is:
[0031]
[0032] in: This represents the total number of perturbation samples; It is an instantaneous security rate function and satisfies and ; and They are the first Channel estimation errors for legitimate users and potential eavesdroppers under a perturbation sampling;
[0033] The expression is:
[0034]
[0035] in: To be according to worst in ascending order quantile index It is the first The instantaneous security rate corresponding to each perturbation sample;
[0036] Power constraints The calculation expression is:
[0037]
[0038] in: This is the power penalty factor, used to control the penalty for exceeding the allowable power budget in transmit power, ensuring that beamforming meets hardware power limits and communication specification requirements.
[0039] A further improvement of the present invention is that the agent includes a value network, the input of which includes a state vector. With the beamforming vector as the action The concatenated vector is output as a scalar for evaluating the long-term returns of the strategy. value.
[0040] A further improvement of the present invention is that the agent includes the policy network. Corresponding target network and value networks Corresponding target network ;in , , , These are the corresponding network parameters; in updating network parameters, the target network... as well as Perform a soft update.
[0041] A further improvement of the present invention is that, in each time slot, the state vector of the current time slot is... Beamforming vector Calculated reward value The state vector of the next time slot Store in playback buffer During the network parameter update process:
[0042] Updating the value network Loss function used The mean square error between the predicted and target values is expressed as follows:
[0043]
[0044] Where B is the batch size. From the playback buffer Random sampling in the middle; For the first The first-order approximation of the target Q value for each sampled time slot; for any time slot The first-order approximation of the objective Q is expressed as:
[0045]
[0046] in: The target Q value is a first-order approximation; This is the reward value for that time slot; For time slots The state vector; It is a discount factor used to control the importance of long-term rewards;
[0047] Policy Network With deterministic gradient updates and optimization using the Adam optimizer, the expression is:
[0048]
[0049] in: This indicates the direction of performance improvement for the policy network under the current parameters, and is used to update the Actor network; For the deterministic policy gradient term, it is used to measure the degree to which the action output by the Actor network in the current state affects the Q value;
[0050] Using the updated policy network parameters and parameters of the value network For the target network respectively as well as The expression for a soft update is:
[0051]
[0052]
[0053] in: It is a constant.
[0054] A further improvement of this invention is that gradient pruning or weight decay is used during the network parameter update process.
[0055] The present invention also provides a terahertz communication system, comprising: a terahertz base station and an intelligent agent; the intelligent agent provides a beamforming vector to the terahertz base station through the above-described terahertz communication secure beamforming method.
[0056] Compared with the prior art, the present invention has the following significant advantages:
[0057] 1. High performance: Simulation results show that, under the same transmit power and antenna size, the method proposed in this invention can improve the system security rate by about 13% to 30% compared with the traditional ZF and SDR schemes.
[0058] 2. Strong robustness: When channel estimation errors exist, the performance of traditional algorithms drops by more than 35%, while the method of this invention drops by only about 13%, demonstrating excellent robustness.
[0059] 3. Low complexity and real-time performance: After training, online decision-making only requires one forward propagation of the neural network, and the delay of a single beam update is less than 1.5ms. The complexity is much lower than convex optimization methods such as SDR, making it suitable for terahertz high-speed real-time communication scenarios.
[0060] 4. Adaptive capability: It can learn and adapt to dynamically changing communication environments without relying on an accurate channel model. Attached Figure Description
[0061] Figure 1. Schematic diagram of the application scenario of the system described in the embodiment of the present invention;
[0062] Figure 2 shows the simulation results, illustrating the security rate performance of the method of the present invention and the comparative method under different transmission powers;
[0063] Figure 3 shows the simulation results, illustrating the security rate performance of the method of the present invention and the comparative method under different numbers of antennas;
[0064] Figure 4 shows the simulation results, illustrating the confidentiality rate performance of the proposed method and the comparative method under channel estimation error.
[0065] Figure 5 shows the ablation experiment, illustrating the convergence speed and stability of the method of this invention and two algorithm variants. Detailed Implementation
[0066] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0067] like Figure 1This invention provides a terahertz communication secure beamforming method based on reinforcement learning technology. With the goal of improving physical layer security performance, a joint beamforming optimization problem is established. The optimization objective is to maximize the system's security rate while satisfying transmit power constraints and QoS requirements.
[0068]
[0069]
[0070] To address the high-dimensional non-convexity, channel imperfections, and dynamic environments inherent in terahertz physical layer secure beamforming, this invention employs the Deep Deterministic Policy Gradient (DDPG) algorithm to learn beamforming vectors in a continuous action space. This method does not require an exact channel model and can achieve online adaptive optimization even under imperfect CSI conditions.
[0071]
[0072] like Figure 1 As shown, the method of the present invention is applicable to intelligent agents, which are connected to terahertz base stations (equipped with...) The agent interacts with the communication environment of the root antenna, the legitimate user (Bob), and the potential eavesdropper (Eve); the agent is set in the beamforming controller.
[0073] The terahertz communication secure beamforming method in this embodiment includes:
[0074] (S1) Construct a state vector by observing the system state through the intelligent agent. The system state includes the channel state information of the legitimate user and the potential eavesdropper; the state vector... The expression is:
[0075]
[0076] in: It is the real part of the channel estimate for legitimate users. It is the imaginary part of the channel estimate for legitimate users; It is the real part of the channel estimate by the potential eavesdropper; It is the imaginary part of the channel estimate by a potential eavesdropper; It is the performance metric of the legitimate user in the previous time slot; It is the performance indicator of a potential eavesdropper in the previous time slot; It is the security rate of the previous time slot, which serves as an important feedback feature of the agent's long-term reward trend; It is the system noise power, which is used as a state input to improve the robustness of the agent under different noise conditions.
[0077] (S2) Through the policy network (Actor) in the agent, based on the state vector Output beamforming vector ; Provides unnormalized beamforming vectors.
[0078] (S3) The power constraint layer normalizes the beamforming vector to obtain the final beamforming vector and transmits the signal. Its expression is:
[0079]
[0080] in: For the final beamforming vector; constant ensure And it is differentiable; It is the system's maximum transmission power; It is the dynamic transmit power budget that the agent can use in the nth time slot.
[0081] To encourage exploration of the continuous action space, in some embodiments, in the final beamforming vector Superimposed zero-mean OU noise And use the superimposed beamforming vector To transmit signals.
[0082] (S4) Calculate the reward value based on the result after signal transmission. Based on the reward value Update the network parameters of the deep reinforcement learning agent.
[0083] The policy network (Actor) in the agent employs a three-layer fully connected structure with ReLU activation, and the output layer has linear activation. The agent also includes a value network (Critic). Its input includes a state vector With the beamforming vector as the action The concatenated vector is output as a scalar for evaluating the long-term returns of the strategy. Value. The Critic network is also a three-layer fully connected network.
[0084] Furthermore, the agent includes the policy network. Corresponding target network and value networks Corresponding target network ;in , , , These are the corresponding network parameters; to stabilize training, the target network is updated during network parameter updates. as well as Perform a soft update.
[0085] In each time slot, the state vector of the current time slot is... Beamforming vector Calculated reward value The state vector of the next time slot Store in playback buffer .
[0086] During the network parameter update process: Based on the Bellman equation, the first-order approximate target Q value is generated using the target network:
[0087]
[0088] in: It is a discount factor;
[0089] Updating the value network Loss function used The mean square error between the predicted and target values is expressed as follows:
[0090]
[0091] Where B is the batch size. From the playback buffer Random sampling in the middle; For the first The first-order approximation of the target Q value for each sampled time slot; for any time slot The first-order approximation objective Q is calculated according to equation (3). By minimizing Network parameters Gradually approaching the Bellman fixed point, making It converges to the optimal value function.
[0092] Policy Network With deterministic gradient updates and optimization using the Adam optimizer, the expression is:
[0093]
[0094] in: This indicates the direction of performance improvement for the policy network under the current parameters, and is used to update the Actor network; For the deterministic policy gradient term, it is used to measure the degree of influence of the Actor network's output action (i.e., beamforming vector) on the Q value in the current state;
[0095] Using the updated policy network parameters and parameters of the value network For the target network respectively as well as The expression for a soft update is:
[0096]
[0097]
[0098] in: The constant is used. Through soft updates and experience replay mechanisms, the algorithm can achieve stable convergence in a non-convex continuous action space, thereby learning the optimal strategy for high-dimensional beamforming. To reduce overestimation bias, dual Critic and target noise smoothing can be employed.
[0099] In this embodiment, the reward value is calculated. The reward function used includes a base security rate, a robustness term, and a power constraint term;
[0100] Basic security rate The expression is:
[0101]
[0102] in: It is a constant; It uses beamforming vectors The signal-to-noise ratio received by legitimate users; It uses beamforming vectors The signal-to-noise ratio received by a potential eavesdropper;
[0103] statistical error or Next, sampling Group perturbation, defined as:
[0104]
[0105] in: This represents the total number of perturbation samples; It is an instantaneous security rate function (equivalent to formula (6)) and satisfies and ; and They are the first Channel estimation errors for legitimate users and potential eavesdroppers under perturbation sampling.
[0106]
[0107] in To be according to worst in ascending order quantile index It is the first The instantaneous security rate corresponding to each perturbation sample. compromise Finally, by superimposing power constraint terms:
[0108]
[0109] in: This is the power penalty factor, used to control the penalty for exceeding the allowable power budget in transmit power, ensuring that beamforming meets hardware power limits and communication specification requirements.
[0110] The final reward function is obtained:
[0111]
[0112] In this embodiment, to improve stability, an empirical replay buffer is used (i). Random sampling breaks the correlation; (ii) soft updates of the target network suppress target drift; (iii) gradient pruning and weight decay alleviate gradient explosion and overfitting.
[0113] In summary, the proposed deep reinforcement learning-based secure beamforming method can be viewed as a two-layer optimization framework with parallel operation of two sub-modules: "policy learning" and "value assessment." Its core idea lies in achieving adaptive decision-making for beamforming through an Actor-Critic collaborative mechanism. The Actor network generates continuous beam weights to improve the signal quality for legitimate users, while the Critic network evaluates the current policy's merits based on immediate rewards and the target Q-value, guiding gradient updates. Both networks iterate cyclically with the support of experience replay and the target network, enabling the agent to gradually approach the optimal secure policy under imperfect CSI and dynamic channel conditions.
[0114] like Figures 2-5 As shown. Figure 2 This paper compares the security rate of the proposed method with that of traditional beamforming methods under different transmit power conditions. As the transmit power increases, the security rate of all methods increases accordingly. However, the beamforming strategy based on deep reinforcement learning in this invention shows a significant advantage across the entire power range, with a particularly noticeable increase in the medium-to-high power region. This is because the agent can adaptively adjust the beam direction according to channel changes under high SNR conditions, further enhancing the gain for legitimate users while effectively suppressing the received power of potential eavesdroppers, thus achieving superior physical layer security performance. Figure 3This paper demonstrates the security rate performance of various methods as the size of the transmit antenna array increases. It shows that the more transmit antennas there are, the higher the available spatial degrees of freedom for the system, and all methods achieve improved security rates. However, the method of this invention shows the most significant gain, and the gain gap further widens as the array element size increases. In contrast, traditional optimization methods such as SDR experience a rapid increase in computational complexity with large-scale arrays, while the deep reinforcement learning-based strategy of this invention maintains stable complexity as the array dimension increases. By learning finer-grained spatial features, it achieves more efficient narrow beamforming and interference suppression, demonstrating good scalability and engineering adaptability. Figure 4 This paper demonstrates the trend of the security rate of various methods as a function of different channel estimation errors (CSI errors). The results show that traditional methods such as ZF and SDR are extremely sensitive to CSI errors, with performance dropping by more than 30%–40% when the error reaches 0.15, exhibiting significant instability. In contrast, the method of this invention shows a smaller decrease in security rate as the error increases, maintaining approximately 90% of its performance even at an error of 0.15, demonstrating significantly better robustness than other methods. This is because the agent integrates error environment characteristics and historical reward trends into its state vector, enabling it to automatically learn optimization strategies to adapt to imperfect CSI conditions, thus making it more feasible for deployment in real-world THz scenarios. Figure 5 The paper presents a comparison of the convergence speed and stability of three algorithm configurations in ablation experiments. Experimental results show that the complete DDPG method proposed in this invention converges fastest and exhibits the best stability, with a steadily rising reward curve. Removing the robust term makes the algorithm more sensitive to channel disturbances, significantly slowing down the convergence speed. Removing the power constraint term leads to unstable beam output, oscillations, and even degradation of the system reward. Therefore, the robust security term and power constraint term introduced in this invention are both crucial for improving system security performance and learning stability, and neither can be omitted.
[0115] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A terahertz communication secure beamforming method, applied to an intelligent agent, wherein the intelligent agent interacts with a communication environment including a terahertz base station, a legitimate user, and potential eavesdroppers; characterized in that, This method includes: A state vector is constructed by observing the system state through the intelligent agent, and the system state includes the channel state information of the legitimate user and the potential eavesdropper. The beamforming vector is output based on the state vector through the policy network Actor in the intelligent agent. The power constraint layer normalizes the beamforming vector to obtain the final beamforming vector, and then transmits the signal. The reward value is calculated based on the result of signal transmission, and the network parameters of the deep reinforcement learning agent are updated based on the reward value.
2. The terahertz communication secure beamforming method according to claim 1, characterized in that, The state vector The expression is: ; in: It is the real part of the channel estimate for legitimate users. It is the imaginary part of the channel estimate for legitimate users; It is the real part of the channel estimate by the potential eavesdropper; It is the imaginary part of the channel estimate by a potential eavesdropper; It is the performance metric of the legitimate user in the previous time slot; It is the performance indicator of a potential eavesdropper in the previous time slot; It is the security rate of the previous time slot; It is the system noise power.
3. The terahertz communication secure beamforming method according to claim 2, characterized in that, The unnormalized beamforming vector output by the policy network Actor ,in The number of antennas; the expression for normalizing the beamforming vector by the power constraint layer is: ; in: For the final beamforming vector; constant ensure And it is differentiable; It is the system's maximum transmission power; Is the agent in the first The dynamic transmit power budget available for each time slot.
4. The terahertz communication secure beamforming method according to claim 3, characterized in that, In the final beamforming vector Superimposed zero-mean OU noise And use the superimposed beamforming vector To transmit signals.
5. A terahertz communication secure beamforming method according to claim 3, characterized in that, Calculate reward value The reward function used includes a robust term and a power constraint term. The expression for the reward function is as follows: ; Basic security rate The expression is: ; in: It is a constant; It uses beamforming vectors The signal-to-noise ratio received by legitimate users; It uses beamforming vectors The signal-to-noise ratio received by a potential eavesdropper; Robust item The calculation expression is: ; in: ; The expression is: ; in: This represents the total number of perturbation samples; It is an instantaneous security rate function and satisfies and ; and They are the first Channel estimation errors for legitimate users and potential eavesdroppers under a perturbation sampling; The expression is: ; in: To be according to worst in ascending order quantile index It is the first The instantaneous security rate corresponding to each perturbation sample; Power constraint terms The calculation expression is: ; in: This is the power penalty factor, used to control the penalty for exceeding the allowable power budget in transmit power, ensuring that beamforming meets hardware power limits and communication specification requirements.
6. The terahertz communication secure beamforming method according to claim 5, characterized in that: The agent includes a value network, whose inputs include a state vector. With the beamforming vector as the action The concatenated vector is output as a scalar for evaluating the long-term returns of the strategy. value.
7. A terahertz communication secure beamforming method according to claim 6, characterized in that, The agent includes the policy network. Corresponding target network and value networks Corresponding target network ;in , , , These are the corresponding network parameters; in updating network parameters, the target network... as well as Perform a soft update.
8. A terahertz communication secure beamforming method according to claim 7, characterized in that, In each time slot, the state vector of the current time slot will be... Beamforming vector Calculated reward value The state vector of the next time slot Store in playback buffer During the network parameter update process: Updating the value network Loss function used The mean square error between the predicted and target values is expressed as follows: ; Where B is the batch size. From the playback buffer Random sampling in the middle; For the first The first-order approximation of the target Q value for each sampled time slot; for any time slot The first-order approximation of the objective Q is expressed as: ; in: The target Q value is a first-order approximation; This is the reward value for that time slot; For time slots The state vector; It is a discount factor used to control the importance of long-term rewards; Policy Network With deterministic gradient updates and optimization using the Adam optimizer, the expression is: ; in: This indicates the direction of performance improvement for the policy network under the current parameters, and is used to update the Actor network; For the deterministic policy gradient term, it is used to measure the degree to which the action output by the Actor network in the current state affects the Q value; Using the updated policy network parameters and parameters of the value network For the target network respectively as well as The expression for a soft update is: ; ; in: It is a constant.
9. A terahertz communication secure beamforming method according to claim 8, characterized in that, Gradient pruning or weight decay is used during network parameter updates.
10. A terahertz communication system, characterized in that... include: Terahertz base station and intelligent agent; the intelligent agent provides beamforming vectors to the terahertz base station using the terahertz communication secure beamforming method according to any one of claims 1 to 9.