MIMO hybrid beamforming method based on multi-behavior deep reinforcement learning
By employing a MIMO hybrid beamforming method based on multi-behavior deep reinforcement learning, the matrix variables of the transmitting and receiving ends are directly optimized, solving the problems of high complexity and slow convergence speed of traditional methods, and achieving high spectral efficiency and stable communication quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 广安理工学院筹建处
- Filing Date
- 2026-04-09
- Publication Date
- 2026-07-07
Smart Images

Figure CN122348760A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication, and in particular to a MIMO hybrid beamforming method based on multi-behavior deep reinforcement learning. Background Technology
[0002] Beamforming methods can compensate for path loss in millimeter-wave MIMO technology, thereby improving spectral efficiency. Traditional all-digital MIMO beamforming requires a dedicated radio frequency (RF) chain for each antenna element. The large number of RF chains required for numerous antenna elements makes the hardware cost and power consumption prohibitively high, making this beamforming method too expensive for millimeter-wave systems. Hybrid analog and digital beamforming (HBF) architecture separates the entire beamformer into a low-dimensional baseband digital pre-encoder matrix and a high-dimensional analog beamforming vector implemented using phase shifters, significantly reducing the number of RF chains while ensuring sufficient beamforming gain. In massive MIMO systems, precoding / combining techniques can simplify hardware complexity, reduce bit error rate, and improve system spectral efficiency. An effective and widely used approach is to treat HBF design as a matrix factorization problem and minimize the Euclidean distance between the hybrid beamformer and the all-digital beamformer.
[0003] However, most HBF (Hybrid Radiation Facility) designs primarily separate the original problem into two sub-problems: the transmitter and the receiver. Within each sub-problem, the analog RF precoder / combiner and the digital baseband precoder / combiner employ an alternating optimization approach. This involves fixing one matrix variable and using an iterative optimization algorithm to find the optimal solution for the other matrix, then switching the optimization objectives. Traditional model-based optimization methods require nested loops, resulting in high complexity, slow convergence, and significant time consumption. Currently, deep reinforcement learning strategies are effective and efficient in HBF design. Furthermore, there is a lack of research on simultaneously optimizing multiple matrix variables at the HBF transmitter and receiver using deep reinforcement learning. Summary of the Invention
[0004] To address the aforementioned shortcomings in existing technologies, this invention provides a MIMO hybrid beamforming method based on multi-behavior deep reinforcement learning, which solves the problems of high complexity, slow convergence speed, and high cost of traditional beamforming methods.
[0005] To achieve the aforementioned objectives, the technical solution adopted by this invention is: a MIMO hybrid beamforming method based on multi-behavior deep reinforcement learning, comprising: The millimeter-wave environmental channel is acquired through the channel feature information acquisition module; the environmental channel includes information on the channel feature matrix, the transmitter module matrix, and the receiver module matrix. A deep reinforcement learning network model is pre-trained using the channel feature matrix, the transmitter module matrix of the previous time step, and the receiver module matrix of the previous time step as state inputs, and outputs the transmitter module matrix and the receiver module matrix of the current time step. MIMO hybrid beamforming is performed based on the current transmitter module matrix and the current receiver module matrix.
[0006] Furthermore, training a deep reinforcement learning network model includes the following steps: The millimeter-wave environmental channel is acquired through the channel feature information acquisition module; the environmental channel includes information on the channel feature matrix, the transmitter module matrix, and the receiver module matrix. The deep reinforcement learning network model takes the channel feature matrix, the transmitter module matrix of the previous time step, and the receiver module matrix of the previous time step as state inputs and outputs the transmitter module matrix and the receiver module matrix of the current time step. Calculate the spectral efficiency at the current time based on the transmitter module matrix and the receiver module matrix at the current time. The reward value is determined based on the change in spectral efficiency at the current time relative to the spectral efficiency at the previous time; wherein, if the spectral efficiency increases, the current reward value increases by 1; if the spectral efficiency remains unchanged, the current reward value is zero; if the spectral efficiency decreases, the current reward value decreases by 1. The reward value is fed back to the deep reinforcement learning network model, and the parameters of the deep reinforcement learning network model are updated with the goal of maximizing the cumulative reward, thereby optimizing the hybrid beamforming parameters and bringing the system spectral efficiency close to the optimal value.
[0007] Furthermore, the deep reinforcement learning network model adopts a multi-behavior output model.
[0008] Furthermore, the transmitter module matrix includes a digital baseband coding matrix and an analog radio frequency coding matrix.
[0009] Furthermore, the receiver module matrix includes a digital baseband decoding matrix and an analog radio frequency decoding matrix.
[0010] Furthermore, the analog RF encoding matrix and the analog RF decoding matrix satisfy the constant modulus constraint condition.
[0011] The present invention also provides a computer device, characterized in that it includes a memory and a processor, the memory storing a computer program, and when the computer program is executed by the processor, causing the processor to perform the steps of a MIMO hybrid beamforming method based on multi-behavior deep reinforcement learning.
[0012] The present invention also provides a computer-readable storage medium, characterized in that it stores a computer program, which, when executed by a processor, causes the processor to perform the steps of a MIMO hybrid beamforming method based on multi-behavior deep reinforcement learning.
[0013] The beneficial effects of this invention are as follows: 1. By directly optimizing the four matrices of HBF with the spectral efficiency (SE) performance index as the optimization target, the antenna array becomes more directive and the energy is more focused, so that the antenna beam is pointed in a specific direction, and the energy of the antenna is concentrated towards a specific user, thereby making the signal received by the user end more concentrated and the communication quality more stable and reliable.
[0014] 2. By using a multi-behavior output model, joint optimization of the transmitting and receiving modules is achieved, thereby maximizing spectral efficiency. This approach eliminates the need for nested loops, has low complexity, requires less computation, saves computational resources, and improves the efficiency of hybrid beamforming. Attached Figure Description
[0015] Figure 1 A flowchart of a MIMO hybrid beamforming method based on multi-behavior deep reinforcement learning is provided for an embodiment. Figure 2 A schematic diagram of a point-to-point millimeter-wave MIMO system with hybrid beamforming provided for an embodiment; Figure 3 This is a schematic diagram of a deep reinforcement learning network model. Detailed Implementation
[0016] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0017] like Figure 1 As shown, in one embodiment of the present invention, a MIMO hybrid beamforming method based on multi-behavior deep reinforcement learning includes the following steps: S1. The millimeter-wave environmental channel is acquired through the channel feature information acquisition module; the environmental channel includes information on the channel feature matrix, the transmitting module matrix, and the receiving module matrix.
[0018] In this embodiment, a point-to-point (single-user, single-link) multiple-input multiple-output (MIMO) communication system operating in the millimeter-wave band, employing a hybrid beamforming (HBF) architecture, is described. Figure 2 As shown. The transmitter in the system has RF radio frequency, One receives RF radio frequency. One transmitting antenna, One receiving antenna, of which . The number of input data streams. The channel matrix between the receiver and transmitter. It is based on the extended Saleh-Valenzudel cluster channel model.
[0019] The channel matrix is as follows:
[0020] In the formula, It is the number of scattering clusters, each cluster containing A streak of scattered light. It is the first In the scattering cluster, the th The gain of the associated paths of a path. It is the normalized receiver array response. It is the normalized transmitter array response. and They represent the first The arrival angle and departure angle of the path.
[0021] The transmitter module matrix consists of a digital baseband encoding matrix. and analog radio frequency coding matrix composition.
[0022] The receiver module matrix consists of a digital baseband decoding matrix. and analog RF decoding matrix Composition. The analog RF encoding matrix and the analog RF decoding matrix satisfy the constant modulus constraint condition.
[0023] S2. A deep reinforcement learning network model is pre-trained with the channel feature matrix, the transmitter module matrix of the previous time step, and the receiver module matrix of the previous time step as state inputs, and outputs the transmitter module matrix and the receiver module matrix of the current time step. S3. Perform MIMO hybrid beamforming based on the current transmitter module matrix and the current receiver module matrix.
[0024] The signal received at the beamforming end can be represented as follows:
[0025] In the formula, This indicates the final signal obtained after the signal at the receiving end has undergone beamforming processing. Indicates average received power. For a noise vector that follows a Gaussian distribution, This indicates the conjugate transpose. In order to send a signal, , It expresses expectation.
[0026] The spectral efficiency SE that the final signal can achieve is:
[0027] In the formula, This represents the noise covariance.
[0028] The objective of this invention is to simultaneously optimize the digital beamforming matrix and analog beamforming matrix of the transmitter and receiver under certain constraints, such as constant mode constraints and transmitter power normalization constraints, thereby maximizing the spectral efficiency (SE).
[0029] like Figure 3 As shown, this invention uses a deep reinforcement learning network model to generate corresponding multiple actions based on the input state. First, the original action is divided into n sub-actions. Then, n neural networks are used to estimate the value of each sub-action. The sub-actions are... Finally, these sub-actions are merged into the original action. All networks output values for each word action: .
[0030] Figure 3 middle, For state and before The union of actions, where, . Figure 3 The dashed box in the diagram represents a single entity, namely an intelligent agent. The intelligent agent obtains state from the environment. After the intelligent agent makes decisions, it outputs multiple actions. The resulting behavior is then applied to the environment, and the environment calculates the reward generated in response to the behavior. And jump to the next state. This process of trial and error training eventually produces the behavior best suited to the environment. In this embodiment, n is 4, which represents the number of behaviors. It includes four matrix variables for hybrid beamforming.
[0031] State in this invention ,Behavior and rewards They are respectively: state: ,in, Indicates channel information, This represents the radio frequency analog precoding matrix of the previous time step. This represents the digital precoding matrix from the previous time step; This represents the RF analog combination matrix from the previous moment. This represents the matrix of combinations of numbers at the previous time step.
[0032] Behavior: ,in, This represents the radio frequency analog precoding matrix at the current moment. The digital precoding matrix represents the current time step; This represents the RF analog combination matrix from the previous moment. A matrix representing the combination of numbers at the current moment.
[0033] The reward is calculated as follows: an increase in spectral efficiency (SE) increases the reward by 1, a decrease in spectral efficiency (SE) decreases the reward by 1, and no change in spectral efficiency (SE) results in no change in the reward. The goal is to find the maximum reward through training a deep reinforcement learning neural network, thereby maximizing spectral efficiency (SE).
[0034] The training process of a deep reinforcement learning network model is as follows: A1. The millimeter-wave environmental channel is acquired through the channel feature information acquisition module; the environmental channel includes information on the channel feature matrix, the transmitting module matrix, and the receiving module matrix; A2. Input the deep reinforcement learning network model with the channel feature matrix, the transmitter module matrix of the previous time step, and the receiver module matrix of the previous time step as the state input, and output the transmitter module matrix and the receiver module matrix of the current time step. A3. Calculate the spectral efficiency at the current time based on the current transmitter module matrix and the current receiver module matrix; A4. Determine the reward value based on the change in spectral efficiency at the current time relative to the spectral efficiency at the previous time; wherein, if the spectral efficiency increases, the reward value increases by 1; if the spectral efficiency remains unchanged, the reward value is zero; if the spectral efficiency decreases, the reward value decreases by 1. A5. Feed the reward value back to the deep reinforcement learning network model, and update the parameters of the deep reinforcement learning network model with the goal of maximizing the cumulative reward, thereby optimizing the hybrid beamforming parameters and making the system spectral efficiency approach the optimal value.
[0035] This invention addresses the problems of high signal attenuation, short transmission distance, and poor transmission capability that occur when using high-frequency millimeter-wave / terahertz electromagnetic waves to transmit signals in 5G and above wireless communication technologies. It proposes using hybrid beamforming technology at both the base station and mobile terminal to improve the directivity of the antenna array and focus energy more effectively. This allows the antenna beam to be pointed in a specific direction, concentrating the antenna's energy towards a specific user, resulting in a more concentrated signal received by the user and more stable and reliable communication quality. Simultaneously, the user terminal can receive beams from a designated direction through beamforming technology, shielding against interfering beams. By employing beamforming technology on both the transmitting and receiving sides, signal gain is improved, frequency efficiency is increased, and communication quality is enhanced.
Claims
1. A MIMO hybrid beamforming method based on multi-behavior deep reinforcement learning, characterized in that, include: The millimeter-wave environmental channel is acquired through the channel feature information acquisition module; the environmental channel includes information on the channel feature matrix, the transmitter module matrix, and the receiver module matrix. A deep reinforcement learning network model is pre-trained using the channel feature matrix, the transmitter module matrix of the previous time step, and the receiver module matrix of the previous time step as state inputs, and outputs the transmitter module matrix and the receiver module matrix of the current time step. MIMO hybrid beamforming is performed based on the current transmitter module matrix and the current receiver module matrix.
2. The method according to claim 1, characterized in that, Training a deep reinforcement learning network model includes the following steps: The millimeter-wave environmental channel is acquired through the channel feature information acquisition module; the environmental channel includes information on the channel feature matrix, the transmitter module matrix, and the receiver module matrix. The deep reinforcement learning network model takes the channel feature matrix, the transmitter module matrix of the previous time step, and the receiver module matrix of the previous time step as state inputs and outputs the transmitter module matrix and the receiver module matrix of the current time step. Calculate the spectral efficiency at the current time based on the transmitter module matrix and the receiver module matrix at the current time. The reward value is determined based on the change in spectral efficiency at the current time relative to the spectral efficiency at the previous time; wherein, if the spectral efficiency increases, the current reward value increases by 1; if the spectral efficiency remains unchanged, the current reward value is zero; if the spectral efficiency decreases, the current reward value decreases by 1. The reward value is fed back to the deep reinforcement learning network model, and the parameters of the deep reinforcement learning network model are updated with the goal of maximizing the cumulative reward, thereby optimizing the hybrid beamforming parameters and bringing the system spectral efficiency close to the optimal value.
3. The method according to claim 2, characterized in that, The deep reinforcement learning network model adopts a multi-behavior output model.
4. The method according to claim 1, characterized in that, The transmitter module matrix includes a digital baseband coding matrix and an analog radio frequency coding matrix.
5. The method according to claim 4, characterized in that, The receiver module matrix includes a digital baseband decoding matrix and an analog radio frequency decoding matrix.
6. The method according to claim 5, characterized in that, The analog RF encoding matrix and the analog RF decoding matrix satisfy the constant modulus constraint.
7. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to perform the steps of the method according to any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, The device stores a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described in any one of claims 1 to 6.