Beam forming and power distribution joint optimization method and system based on deep reinforcement learning
By employing a joint optimization method of beamforming and power allocation based on deep reinforcement learning in the terahertz UM-MIMO-ISCAP system, combined with a hybrid optimization algorithm of DNN and PSO, the problems of high system energy consumption and high computational complexity are solved, achieving efficient sensing, communication and energy transmission, and improving system energy efficiency and algorithm convergence.
Patent Information
- Application Number
- CN202511493484.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies in terahertz UM-MIMO-ISCAP systems suffer from high energy consumption, high computational complexity, slow convergence speed, and difficulty in meeting millisecond-level latency and extremely high reliability requirements. In particular, in large-scale multi-agent environments, the model scalability, training energy consumption, and online decision-making stability are insufficient, and there is a lack of a unified optimization framework to handle the deep coupling of communication, sensing, and energy transmission.
A joint optimization method for beamforming and power allocation based on deep reinforcement learning is adopted. By constructing a hybrid near-field UM-MIMO-ISCAP system model, DNN is used for adaptive pilot design and channel estimation. Combined with the global search capability of PSO, a second-order time difference method and a binary tree experience pool structure are introduced to optimize the MADDQN algorithm to improve system energy efficiency.
It significantly improves sensing resolution and communication capacity, while achieving energy transmission efficiency, reducing system energy consumption, improving algorithm convergence and computational efficiency, avoiding local optima, and enhancing the overall energy efficiency of the system.
Smart Images

Figure CN121001114A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of communication, and particularly relates to a beamforming and power allocation joint optimization method and system based on deep reinforcement learning. BACKGROUND
[0002] In order to compensate for the path loss of the high frequency band, the terahertz UM-MIMO-ISCAP system often uses a large-scale array of thousands of units, and is equipped with a large number of phase shifters, true time delay units, switching networks and power amplifiers. With the expansion of the array size, the direct current power supply, heat dissipation and radio frequency link loss of the base station end increase exponentially, resulting in unacceptable system overall energy consumption, volume and cost; at the same time, the power consumption of the adjustable true time delay network required for beam squint compensation is particularly prominent, further reducing the overall energy efficiency.
[0003] Existing energy efficiency improvement schemes mostly rely on iterative algorithms such as convex approximation, alternating optimization and semidefinite relaxation. In the face of high-dimensional mixed beamforming, power / spectrum allocation and multi-objective constraints of UM-MIMO-ISCAP, such algorithms often have high computational complexity, slow convergence speed, and are easy to fall into local optimum; in the rapidly changing terahertz channel environment, frequent CSI estimation and feedback further slow down the real-time performance, making it difficult to meet the millisecond-level latency and extremely high reliability requirements.
[0004] Most researches focus on communication performance or single ISAC scenario, and lack a unified optimization framework for the deep coupling of communication, sensing and energy transmission. The position error of energy collection users, clutter scattering and multi-user interference will introduce significant uncertainty, which will not only reduce the sensing accuracy, but also weaken the efficiency of wireless energy transmission, and further slow down the overall energy efficiency. In addition, although existing machine learning and reinforcement learning methods can reduce the computational burden, they still face the bottlenecks of model scalability, training energy consumption and online decision stability in large-scale multi-agent environments. SUMMARY
[0005] In view of the problems existing in the prior art, the application provides a beamforming and power allocation joint optimization method based on deep reinforcement learning.
[0006] The application is implemented in the following manner: a beamforming and power allocation joint optimization method based on deep reinforcement learning comprises the following steps: Step 1, constructing a UM-MIMO-ISCAP system model based on mixed far and near fields; Step 2, the echo signal is adaptively pilot-designed and channel-estimated by a DNN composed of a dimension reduction network and a reconstruction network, and the CSI in the frequency-angle domain is extracted by an end-to-end data-driven method; then, the estimated CSI is three-dimensionally tensor-decomposed to extract the accurate angle of arrival and angle of departure, which are two CSI parameters; and the two parameters are taken as the input of the subsequent MADDQN to dynamically adjust the beamforming matrix and power allocation parameter, so as to maximize the system energy efficiency; Step 3, the global search capability of PSO is excavated and used for rapid generation of the initial beamforming matrix and power allocation vector; a second-order time difference method and a binary tree experience pool structure are introduced for optimization of the MADDQN algorithm; the second-order time difference method is used to reconstruct the loss function in the MADDQN model training process, and the binary tree structure is used for storing experience replay.
[0007] Further, the system model comprises: A. a communication model; B. a perception model; C. an energy transmission model; D. problem modeling.
[0008] Further, the communication model comprises: It is assumed that the UM-MIMO-based ISCAP base station adopts a full-connection hybrid beamforming architecture to serve EH and ID receivers with RF chains, wherein ; it is assumed that the base station sets a UM-MIMO antenna array based on a uniform linear array (ULA), adopts a hybrid plane wave and spherical wave modeling method, uses a plane wave modeling method within a subarray, and uses a spherical wave model between subarrays to improve modeling accuracy; let denote the transmission energy signal of the th EH receiver with a power of , and let denote the transmission information-bearing signal of the th ID receiver with a power of , then the signal transmitted by the base station can be represented as: (1) wherein contains unit power data streams for IDs, is a hybrid field channel matrix, and represent an analog beamforming matrix and a digital beamforming matrix, respectively; it is assumed that the signal streams are independent of each other, that is , is the K-dimensional identity matrix; the transmit power can be written as: (2) The channel matrix in the hybrid modeling approach can be expressed as: (3) where, denotes the propagation path, denotes the line-of-sight path, denotes the non-line-of-sight path; is the amplitude of the path gain; denotes the th path, denotes the distance from the th base station terminal array to the th user terminal array; denotes the array response vector of the transmit sub-array, denotes the array response vector of the receive sub-array, which can be expressed as: (4) (5).
[0009] Further, the perception model: considers the near-field to have a multi-point target single-static perception, and uses the perception echo signal to detect the target; it is assumed that and are the distance and angle from the detection target to the origin, respectively, denotes the echo signal received at the base station on the th time slot, and can be defined as: (6) where, denotes the reflection coefficient of the th user, is the Gaussian white noise at the base station receiver, is the channel matrix of the base station-user-end-base station link, where, and represent the receive steering vector and the transmit steering vector of the base station, respectively, and since the expressions of the two are similar, we let be the AoA of the th time slot, and is given as: (7).
[0010] Further, the energy transfer model: assumes that the base station and the There is a line-of-sight (LoS) path between each ID receiver and For a non-line-of-sight (NLoS) path, since the UM-MIMO-ISCAP channel of this invention is in the easily blocked terahertz band, the NLoS component in the high-frequency band can be ignored; for each ID receiver in the far field of the base station, according to the plane wavefront propagation model
[43] , the channel characteristics can be expressed as: (8) in, It is the first Complex channel gain of each ID receiver, Let represent the far-field channel rotation vector, where It is the spatial angle at the base station. Indicates the distance from the center of the base station to the... AoD of each ID receiver; make This represents the beam scheduling indication set from the UM-MIMO base station, where and They represent the first The first EH receiver and the first The binary scheduling variables of each ID receiver; specifically, if the EH receiver is scheduled through the base station, then ,otherwise , And it is the same; therefore, the first The received signal of a far-field ID receiver can be represented as: (9) in, The additive white Gaussian noise received at the ID receiver has a mean of 0 and a power of Therefore, in the first The signal-to-interference-plus-noise ratio (SINR) at each ID receiver can be expressed as: (10) in, It is the first Power is allocated to each ID receiver. It is the first Power is allocated to each EH receiver ; Assuming at the base station and the first There is a Loss path between the EH receivers and a NLoS path, the channel from the base station to the th EH receiver can be estimated as ; therefore, the near-field channel from the base station to the th EH receiver can be simply modeled as: (11) where represents the near-field channel steering vector; denotes the distance between the base station center and the th EH receiver, and denotes the spatial angle at the base station, is the AoD from the base station center to the th EH receiver; for wireless energy transmission, since the broadcast nature of the wireless channel, each EH receiver can harvest wireless energy from both energy and information signals; therefore, the present invention neglects the noise power at the near-field EH receiver and assumes a linear EH model, then the energy received by the th EH receiver is: (12) where denotes the energy reception efficiency.
[0011] Further, the problem is modeled as: To facilitate implementation, it is assumed that the base station will direct a beam towards the EH receiver to maximize the energy efficiency of the system, i.e. and ; then, the achievable rate of each ID receiver can be represented as equation (14): (13) Equation (12) can be written as: (14) where ; the energy efficiency of the th user is: (15) Let denote the predefined power weight of the th EH receiver, where the larger , the more preferred is to deliver energy to the EH receiver ; therefore, the weighted power sum delivered to all EH receivers can be represented as: Therefore, the total energy efficiency of the system can be written as ; Under the sum-rate constraints of all ID receivers and the constraint of the total transmit power of the base station, the hybrid beamforming matrix and the power allocation are jointly optimized to maximize the energy efficiency of the system, so the optimization problem can be represented as the following formula: (17a) (17b) (17c) (17d) (17e) Wherein, (17b) represents the sum-rate constraint of all ID receivers; (17c) represents the binary beam scheduling indicator of each EH receiver and ID receiver; (17d) is the maximum transmit power constraint from the base station; However, the optimization problem (P1) is a mixed integer optimization problem caused by the binary beam scheduling optimization variable of (17c) and the continuous power allocation variable of (17d), so the present application will introduce a continuous variable to eliminate the binary optimization variable, that is ; Therefore, , further, formula (13) can be re-expressed as: (18) Since the ID receiver is susceptible to the interference of the EH receiver when receiving information when allocating power, the EH receiver cannot efficiently collect energy, therefore, it is necessary to study the correlation between the EH receiver and the ID receiver channel; first, the correlation between two near-field rotation vectors is defined: (19) Therefore, based on the above analysis, formula (14) and formula (17) can be re-expressed as a function of EH and ID correlation, as shown in the following formula: (20) (21) In order to facilitate analysis, the present application defines a correlation matrix: (22) Wherein, represents the EH / ID receiver and the EH / ID receiver the correlation between the channel steering vectors of the ID receivers; next, without considering the energy collection of the ID receivers, the problem (P1) is further re-expressed in a more compact form, the present application will set some diagonal elements to 0, and the new correlation matrix can be expressed as: (23) Let represent the vector containing all the power allocation optimization variables; then, the problem (P1) can be redefined as: (24a) (24b) (24c) (24d) wherein, and .
[0012] Another object of the present application is to provide a deep reinforcement learning-based beamforming and power allocation joint optimization system comprising: a construction module configured to construct a hybrid far and near field-based UM-MIMO-ISCAP system model; a CSI estimation module configured to perform adaptive pilot design and channel estimation on echo signals by using a DNN composed of a dimension reduction network and a reconstruction network, extract CSI in the frequency-angle domain through an end-to-end data-driven method, then perform three-dimensional tensor decomposition on the estimated CSI to extract accurate angle of arrival and angle of departure, two CSI parameters, and use the two parameters as inputs of a subsequent MADDQN state space to dynamically adjust a beamforming matrix and a power allocation parameter, and maximize system energy efficiency; an optimization module configured to mine the global search capability of PSO and use the global search capability to rapidly generate an initial beamforming matrix and a power allocation vector, introduce a second-order time difference method and a binary tree experience pool structure into the MADDQN algorithm for optimization, use the second-order time difference method to reconstruct a loss function in the MADDQN model training process, and use the binary tree structure to store experience replay.
[0013] Another object of the present application is to provide a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the computer program is executed by the processor to enable the processor to perform the steps of the deep reinforcement learning-based beamforming and power allocation joint optimization method.
[0014] Another object of the present application is to provide a computer readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method for joint optimization of beamforming and power allocation based on deep reinforcement learning.
[0015] Another object of the present application is to provide an information data processing terminal for implementing the system for joint optimization of beamforming and power allocation based on deep reinforcement learning.
[0016] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present application are analyzed from the following aspects: The deep integration of terahertz UM-MIMO and ISCAP (Integrated Sensing, Communication and Powering) can significantly improve the sensing resolution and communication capacity while achieving energy transmission efficiency. However, the expansion of the functions of the terahertz UM-MIMO-ISCAP system and the increase in the number of antennas will lead to an increase in system energy consumption, which in turn will lead to a decrease in energy efficiency. Therefore, the present application proposes a joint optimization algorithm for hybrid beamforming and power allocation based on DNN and PSO-MADDQN. First, an energy collection (EH) receiver is deployed in the near field of the base station, and an information decoder (ID) receiver is placed in the far field, while using UM-MIMO based on hybrid plane wave and spherical wave to effectively suppress the near-field effect. Second, due to imperfect channel state information (CSI), the joint optimization problem will fall into a local optimum, so the present application uses a deep neural network (DNN) to adaptively design the pilot and channel estimation of the echo signal, and performs three-dimensional tensor decomposition on the obtained CSI to extract accurate angle parameters. In order to further solve the problem of high parameter space dimension and difficult exploration in the joint optimization process of hybrid beamforming and power allocation in the terahertz UM-MIMO-ISCAP system, the present application proposes a MADDQN algorithm improved by particle swarm optimization (PSO), which uses the adaptive PSO algorithm as a global optimization tool, provides high-quality initial solutions for MADDQN through iterative search of the particle swarm, and optimizes the beamforming matrix and power allocation vector to improve the system energy efficiency. In addition, in order to further improve the convergence and computational efficiency of the algorithm, in the training process of the MADDQN model, the second-order time difference optimization loss function is used to avoid the problem of overestimation in the traditional method, and the binary tree structure is used to store the experience replay to improve the utilization efficiency of historical information and reduce the computational overhead. The simulation results show that the joint optimization algorithm based on DNN and PSO-MADDQN proposed in the present application realizes more accurate hybrid beamforming and power allocation, and improves the overall energy efficiency of the system.
[0017] In the hybrid field terahertz ultra-massive MIMO-ISCAP system, the plane wave cannot accurately describe the channel characteristics, and the matrix dimension of the spherical wave channel model increases significantly, and the calculation complexity increases greatly. In order to balance the accuracy and overhead of channel modeling, the UM-MIMO-ISCAP system model based on hybrid near and far field is constructed. Specifically, the EH receiver is deployed in the near field of the base station to realize high-power energy transmission, and the ID receiver is placed in the far field of the base station to make the signal of the far field ID receiver can effectively power the near field EH receiver within a certain angle range, and the UM-MIMO adopts the way of modeling of hybrid plane wave and spherical wave, effectively solves the problem that the spherical wave channel model causes too many parameters and the calculation complexity increases greatly due to the dramatic increase of channel matrix dimension.
[0018] In a multi-user scenario, if the CSI is not accurate, the system cannot accurately allocate power, resulting in an increase in system energy efficiency, and when solving the energy efficiency maximization problem, the algorithm is easy to fall into a local optimal solution. Therefore, the present application proposes a CSI estimation scheme based on DNN. This method uses the DNN composed of dimension reduction network and reconstruction network to perform adaptive pilot design and channel estimation on the echo signal, and extracts the CSI in the frequency-angle domain through the end-to-end data-driven method. Then, the estimated CSI is three-dimensional tensor decomposition, and the accurate Angle of Arrival (AoA) and Angle of Departure (AoD) two CSI parameters are extracted. And take it as the input of the subsequent MADDQN state space to dynamically adjust the beamforming matrix and power allocation parameters, so as to realize the maximization of system energy efficiency.
[0019] Because the joint optimization of beamforming and power allocation in the UM-MIMO-ISCAP system involves a large number of beamforming matrices and power allocation vectors, a large-dimensional parameter space is generated, making the exploration of MADDQN more difficult. In order to solve these problems and further improve the energy efficiency of the system, the present application proposes a joint optimization algorithm based on PSO-MADDQN. This algorithm exploits the global search ability of PSO and uses it to quickly generate initial beamforming matrices and power allocation vectors. In addition, in order to further improve the convergence and computational efficiency of the algorithm, the present application also introduces the second-order time difference method and the binary tree experience pool structure into the optimization MADDQN algorithm. This method uses the second-order time difference method to reconstruct the loss function during the training process of the MADDQN model, and at the same time uses the binary tree structure to store the experience replay, improves the utilization efficiency of historical information, avoids the problem of overestimation in the traditional method, improves the convergence of the algorithm and reduces the computational overhead. The present application proposes a hybrid beamforming and power allocation joint optimization algorithm based on DNN and PSO-MADDQN. Specifically, first, the hybrid spherical wave and plane wave modeling method is used for channel modeling of the terahertz UM-MIMO-ISCAP system, and the energy harvesting (Energy Harvest, EH) receiver is placed in the near field and the information decoding (Information Decode, ID) receiver is placed in the far field, balancing the accuracy and overhead of channel modeling. Secondly, in view of the situation that imperfect CSI leads to poor solution of the joint optimization problem, a DNN-based CSI estimation algorithm is proposed to estimate and extract AoD / AoA. Finally, in order to solve the problem that MADDQN has limited exploration ability in high-dimensional space and may fall into a local optimal solution, a PSO-MADDQN algorithm is proposed to improve the global convergence and search efficiency of the algorithm. In addition, in order to further improve the convergence of the algorithm, the second-order time difference method and the binary tree structure are used to improve the utilization rate of historical information and reduce the computational overhead.
[0020] The present application aims at the problem that the expansion of the function and the increase of the number of antennas of the terahertz ultra-massive MIMO-ISCAP system will lead to the increase of the energy consumption of the system, and further lead to the decrease of the energy efficiency. A hybrid beamforming and power allocation joint optimization algorithm based on DNN and PSO-MADDQN is proposed to improve the energy efficiency of the system. First, in order to balance the accuracy and overhead of channel modeling in the hybrid field of the terahertz UM-MIMO-ISCAP system, the present application constructs a UM-MIMO-ISCAP system model based on hybrid near and far field. The EH receiver is deployed in the near field of the base station to realize high-power energy transmission, and the ID receiver is placed in the far field of the base station, and the signal of the far field ID receiver can effectively power the near field EH receiver within a certain angle area through energy leakage, effectively solving the problem of excessive parameters and greatly increased computational complexity caused by the dramatic increase of the dimension of the channel matrix in the spherical wave channel model. In the multi-user scenario, if the CSI is not accurate, the system cannot accurately allocate power, leading to the increase of the system energy efficiency, therefore, the present application estimates the CSI of the perception echo by using the DNN composed of the dimension reduction network and the reconstruction network, and performs three-dimensional tensor decomposition on the estimated CSI to extract the two CSI parameters of AoA and AoD. Since the joint optimization of hybrid beamforming and power allocation in the UM-MIMO-ISCAP system will involve a large number of hybrid beamforming matrices and power allocation vectors, resulting in a large parameter space with high dimension, making the exploration of MADDQN more difficult, in order to further improve the energy efficiency of the system, the present application proposes a joint optimization algorithm based on PSO-MADDQN. This algorithm excavates the global search ability of PSO and uses it to quickly generate the initial hybrid beamforming matrix and power allocation vector. Then, the present application performs fine parameter optimization operation through MADDQN, at this time, each agent will independently optimize the hybrid beamforming and power allocation of EH or ID, which significantly improves the overall energy efficiency of the system while reducing the training complexity of the MADDQN network. In addition, in order to further improve the convergence and calculation efficiency of the algorithm, the present application uses the second-order time difference method to reconstruct the loss function in the training process of the MADDQN model, and uses the binary tree structure to store the experience replay, improves the utilization efficiency of historical information, avoids the problem of overestimation in the traditional method, improves the convergence of the algorithm and reduces the calculation overhead. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 is the flow chart of the joint optimization method of beamforming and power allocation based on deep reinforcement learning provided by the embodiment of the present application.
[0022] Figure 2 is the structure block diagram of the joint optimization system of beamforming and power allocation based on deep reinforcement learning provided by the embodiment of the present application.
[0023] Figure 3 is a mixed field-based ISCAP model diagram provided by an embodiment of the application.
[0024] Figure 4 is an improved DNN channel state information extraction structure diagram provided by an embodiment of the application.
[0025] Figure 5 is a DDQN algorithm structure diagram improved based on adaptive PSO provided by an embodiment of the application.
[0026] Figure 6 is a total energy efficiency comparison diagram of the proposed algorithm changing with the number of episodes under different discount rates provided by an embodiment of the application.
[0027] Figure 7 is a loss value change diagram of the proposed algorithm with the number of iterations under different learning rates provided by an embodiment of the application.
[0028] Figure 8 is a MES change diagram of different CSI extraction algorithms with the signal-to-noise ratio provided by an embodiment of the application.
[0029] Figure 9 is a comparison diagram of the energy efficiency of different algorithms changing with the number of iterations provided by an embodiment of the application.
[0030] Figure 10 is a comparison diagram of the energy efficiency of different algorithms changing with the number of EH users provided by an embodiment of the application.
[0031] Figure 11 is a comparison diagram of the energy efficiency of different algorithms changing with the maximum transmit power provided by an embodiment of the application.
[0032] Figure 12 is a comparison diagram of the energy efficiency of different algorithms changing with the achievable rate (a) the number of antennas is 64; (b) the number of antennas = 1024 provided by an embodiment of the application. Figure 13 is a comparison diagram of the energy efficiency of different algorithms changing with the noise amplitude provided by an embodiment of the application.
[0033] Figure 14 is a comparison diagram of the energy efficiency of different algorithms changing with the number of antennas provided by an embodiment of the application. DETAILED DESCRIPTION
[0034] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0035] As Figure 1As shown, the beamforming and power allocation joint optimization method based on deep reinforcement learning provided by this embodiment of the invention includes the following steps: S101, Construct a UM-MIMO-ISCAP system model based on hybrid near and far fields; S102 utilizes a DNN composed of a dimensionality reduction network and a reconstruction network to perform adaptive pilot design and channel estimation on the echo signal. The frequency-angle signal identification (CSI) is extracted using an end-to-end data-driven method. Then, the estimated CSI is decomposed into three-dimensional tensors to extract the accurate angle of arrival (AQ) and angle of deviation (OG). These parameters are then used as inputs to the subsequent MADDQN to dynamically adjust the beamforming matrix and power allocation parameters, thereby maximizing the system's energy efficiency. S103 explores the global search capability of PSO and uses it to rapidly generate the initial beamforming matrix and power allocation vector; it introduces the second-order time difference method and binary tree experience pool structure to optimize the MADDQN algorithm; it uses the second-order time difference method to reconstruct the loss function during the training process of the MADDQN model, and uses the binary tree structure to store experience playback.
[0036] like Figure 2 As shown, an embodiment of the present invention provides a beamforming and power allocation joint optimization system based on deep reinforcement learning, comprising: The building block is used to construct a UM-MIMO-ISCAP system model based on hybrid near and far fields; The CSI estimation module is used to perform adaptive pilot design and channel estimation of echo signals using a DNN composed of a dimensionality reduction network and a reconstruction network. It extracts the CSI in the frequency-angle domain through an end-to-end data-driven method. Then, it performs three-dimensional tensor decomposition on the estimated CSI to extract the accurate angle of arrival and angle of deviation CSI parameters. These parameters are then used as inputs to the subsequent MADDQN to dynamically adjust the beamforming matrix and power allocation parameters to maximize system energy efficiency. An optimization module is used to explore the global search capability of PSO and apply it to the rapid generation of the initial beamforming matrix and power allocation vector; a second-order time difference method and a binary tree experience pool structure are introduced to optimize the MADDQN algorithm; the second-order time difference method is used to reconstruct the loss function during the training of the MADDQN model, while the binary tree structure is used to store experience playback.
[0037] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to perform the steps of the beamforming and power allocation joint optimization method based on deep reinforcement learning.
[0038] Another object of the present application is to provide a computer readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method for joint optimization of beamforming and power allocation based on deep reinforcement learning.
[0039] Another object of the present application is to provide an information data processing terminal for implementing the system for joint optimization of beamforming and power allocation based on deep reinforcement learning.
[0040] The present application is embodied as follows: 1. The present application proposes a hybrid beamforming and power allocation joint optimization algorithm based on DNN and PSO-MADDQN. Specifically, first, the hybrid spherical wave and plane wave modeling method is used for channel modeling of the terahertz UM-MIMO-ISCAP system, and the energy harvesting (Energy Harvest, EH) receiver is placed in the near field, and the information decoding (Information Decode, ID) receiver is placed in the far field, balancing the accuracy and overhead of channel modeling. Second, in view of the situation that imperfect CSI leads to poor solution of the joint optimization problem, a CSI estimation algorithm based on DNN is proposed to estimate and extract AoD / AoA and obtain CSI. Finally, in order to solve the problem that MADDQN has limited exploration ability in high-dimensional space and may fall into a local optimal solution, a PSO-MADDQN algorithm is proposed to improve the global convergence and search efficiency of the algorithm. In addition, in order to further improve the convergence of the algorithm, the second-order time difference method and binary tree structure are used to improve the utilization rate of historical information and reduce the computational overhead.
[0041] II. System model Considering the near-field effect of the terahertz UM-MIMO-ISCAP system, in order to balance the accuracy and overhead of channel modeling, a UM-MIMO-ISCAP model based on hybrid near and far field is constructed, as shown in Figure 3 The single antenna EH receiver is deployed in the near field of the base station, and the spherical wave model is used to accurately capture the near-field signal propagation characteristics to achieve high-power energy transmission, and the ID receiver is placed in the far field of the base station, and the plane wave is used to process signal transmission. The signal of the far-field ID receiver can effectively supply energy to the near-field EH receiver within a certain angle range by using energy leakage.
[0042] A. Communication model It is assumed that the UM-MIMO-based ISCAP base station uses a fully connected hybrid beamforming architecture to serve users with One EH and one ID receiver, among which Assume the base station is equipped with a UM-MIMO antenna array based on a Uniform Linear Array (ULA), employing a hybrid plane wave and spherical wave modeling approach. Plane wave modeling is used within a subarray, while spherical wave models are used between subarrays to improve modeling accuracy. Let... Indicates each power is The The transmitted energy signal of each EH receiver indicates a power of The Each ID receiver transmits information carrying a signal. Therefore, using hybrid beamforming technology, the signal transmitted by the base station can be represented as: (1) in, Includes targeting ID A unit power data stream, For the mixed field channel matrix, and These represent the analog beamforming matrix and the digital beamforming matrix, respectively. It is assumed that the signal flows are independent of each other. , It is a K-dimensional identity matrix. The transmit power can be written as: (2) The channel matrix under the hybrid modeling approach can be represented as: (3) in, Indicates the propagation path, Indicates the line-of-sight path. Indicates a non-line-of-sight path; The amplitude of the path gain; Indicates the first The first path The base station terminal array to the first The distance between user terminal arrays; This represents the array response vector of the transmitting subarray. The array response vectors of the receiving subarrays can be represented as: (4) (5) B. Perceptual Model This invention considers near-field multi-point target single static sensing, utilizing the sensed echo signal to detect the target. Assume... and respectively the distance and angle of the detection target to the origin, represents the echo signal received at the base station on the time slot, may be defined as: (6) where, represents the reflection coefficient of the th user, is the Gaussian white noise at the base station receiver, is the channel matrix of the base station-user-end-base station link, where, and represent the receive steering vector and transmit steering vector of the base station, respectively, and since the two expressions are similar, we let is the AoA of the th time slot, here we only give is: (7) C. Energy transmission model Assume that there is a line of sight (LoS) path and non-line of sight (NLoS) path between the base station and the th ID receiver, since the UM-MIMO-ISCAP channel of the present application is in the terahertz band which is prone to obstruction, the NLoS component in the high frequency band can be ignored. For each ID receiver in the far field of the base station, according to the plane wave front propagation model
[43] , the channel characteristics can be represented as: (8) where, is the complex channel gain of the th ID receiver, represents the far-field channel rotation vector, where, is the spatial angle at the base station, represents the AoD from the center of the base station to the th ID receiver.
[0043] Let represent the set of beam scheduling indications from the UM-MIMO base station, where and represent the binary scheduling variables of the th EH receiver and the th ID receiver, respectively. Specifically, if the EH receiver is scheduled by the base station , otherwise , and it is the same. Therefore, the The received signal of a far-field ID receiver can be represented as: (9) in, The additive white Gaussian noise received at the ID receiver has a mean of 0 and a power of Therefore, in the first The signal-to-interference-plus-noise ratio (SINR) at each ID receiver can be expressed as: (10) in, It is the first Power is allocated to each ID receiver. It is the first Power is allocated to each EH receiver .
[0044] Assuming at the base station and the first There is a Loss path between the EH receivers and The NLoS path then extends from the base station to the... The channel of each EH receiver can be estimated as follows: Therefore, from the base station to the... The near-field channel of an EH receiver can be simply modeled as: (11) in, This represents the near-field channel steering vector. Indicates the base station center and the first The distance between the EH receivers, and Indicates the spatial angle at the base station. It is from the base station center to the first The AoD of the EH receiver. For wireless power transfer, due to the broadcast characteristics of the wireless channel, each EH receiver can obtain wireless power from both the power and information signals. Therefore, this invention ignores the noise power at the near-field EH receiver and assumes the selection of a linear EH model, then the... The energy received by each EH receiver is: (12) in, This indicates the energy receiving efficiency.
[0045] D. Problem Modeling For ease of implementation, assume that the base station will direct a beam toward the EH receiver to maximize the system's energy efficiency, i.e. as well as . Then, the achievable rate of each ID receiver can be expressed as equation (14): (13) Then, equation (12) can be written as: (14) where . Then, the energy efficiency of the th user is: (15) Let denote the predefined power weight of the th EH receiver, where the larger , the more the th EH receiver prefers to have energy delivered to it compared to other EH receivers. Therefore, the weighted power sum delivered to all EH receivers can be expressed as: (16) Then, the total energy efficiency of the system can be written as .
[0046] The research goal of the present invention is to jointly optimize the hybrid beamforming matrix and power allocation to maximize the energy efficiency of the system under the sum rate constraints of all ID receivers and the constraint of the total transmit power of the base station. Therefore, the optimization problem can be expressed as the following equations: (17a) (17b) (17c) (17d) (17e) where (17b) represents the sum rate constraints of all ID receivers; (17c) represents the binary beam scheduling indicators of each EH receiver and ID receiver; and (17d) is the maximum transmit power constraint from the base station.
[0047] However, the optimization problem (P1) is a mixed integer optimization problem caused by the binary beam scheduling optimization variables of (17c) and the continuous power allocation variables of (17d). Therefore, the present invention introduces continuous variables to eliminate the binary optimization variables, i.e. . Therefore, , further, equation (13) can be re-expressed as: (18) Since the ID receiver is vulnerable to the EH receiver when receiving information when allocating power, the EH receiver cannot efficiently collect energy, so it is necessary to study the correlation between the EH receiver and the ID receiver channel. First, the correlation between the two near-field rotation vectors is defined: (19) Therefore, based on the above analysis, formula (14) and formula (17) can be re-expressed as a function of EH and ID correlation, as shown in the following equation: (20) (21) For ease of analysis, the present application defines a correlation matrix: (22) Where, represents the correlation between the channel rotation vectors of the EH / ID receiver and the EH / ID receiver . Next, without considering the energy collection of the ID receiver, the problem (P1) is further re-expressed in a more compact form, the present application will set some diagonal elements to 0, the new correlation matrix can be expressed as: (23) Let represent the vector containing all the power allocation optimization variables. Then, problem (P1) can be redefined as: (24a) (24b) (24c) (24d) Where, and .
[0048] III. Hybrid beamforming and power allocation joint optimization algorithm based on DNN and PSO-MADDQN The deep integration of terahertz UM-MIMO antenna array and ISCAP can not only compensate for path loss, but also significantly improve the perception resolution, communication capacity and energy transmission efficiency. However, with the expansion of system functions and the increase of antenna number, the system may not receive enough energy, thereby causing a series of problems. On the one hand, due to the increase of energy consumption caused by the expansion of the functions of the terahertz UM-MIMO-ISCAP system, the system energy efficiency is reduced. On the other hand, due to the significant near-field effect in the terahertz communication system, the traditional optimization method is difficult to effectively handle the difference between the far-field and near-field propagation mechanisms, further increasing the non-convexity and solving difficulty of the problem, thereby causing the algorithm convergence speed to slow down and the system performance to decrease significantly. Therefore, the present application will design a hybrid beamforming and power allocation joint optimization algorithm based on DNN and PSO-MADDQN.
[0049] A. CSI estimation algorithm based on improved DNN In a multi-user scenario, if the CSI is not accurate, the system cannot accurately allocate power, resulting in an increase in system energy efficiency, and when solving the energy efficiency maximization problem, the algorithm is prone to fall into a local optimal solution, therefore the present application proposes a CSI estimation scheme based on DNN. The scheme uses a DNN network to perform adaptive pilot design and channel estimation on the echo signal, and effectively estimates the CSI in the frequency-angle domain through an end-to-end data-driven method. Then, the present application performs three-dimensional tensor decomposition on the CSI obtained by DNN estimation to extract accurate angle parameters.
[0050] Perception echo signal As shown in formula (6), due to the compressibility of the UM-MIMO channel in the frequency-angle domain, the frequency-space domain channel can be converted into the frequency-angle domain channel , that is , is the transform domain matrix, then formula (6) can be rewritten as: (25) wherein, .
[0051] In the adaptive pilot design stage, a dimension reduction network is used to simulate the compression process of high-dimensional signals. First, the real and imaginary parts of the channel are separated and stacked
[44] . Then, the channel is multiplied by the combination matrix on the left and the pilot signal on the right, thereby obtaining the low-dimensional measurement value, as shown in formula (26): (26) At this time, for each measurement value output, the real and imaginary parts are the weighted linear combination of the real and imaginary parts corresponding to all input channel values.
[0052] In the channel estimation stage, the reconstruction network used to reconstruct the channel from the low-dimensional measurements consists of two cascaded parts: the first part uses a fully connected (FC) layer without bias term and nonlinear activation function to represent the FC layer simulating the fitting operation of the greedy compressed sensing (CS) algorithm, and to calculate the inner product of the measurement matrix and the actual measurement value. In the second part, due to the compressibility of the frequency-angle domain channel in UM-MIMO, the present application can be roughly estimated and reshaped. By refining the rough estimate, the reconstruction of the channel is realized, and a more accurate channel estimation is obtained. In the DNN architecture, a refinement unit is included, and each unit consists of a convolutional layer, a batch normalization layer and a ReLU activation function. The present application uses a data-driven DNN to perform channel estimation, and selects the mean square error (MSE) to quantify the difference between the input and the output, that is: (27) wherein, denotes the number of samples in each batch in the training set, and denote the real sample and the estimation of the real sample of the th batch, respectively. is the fitting function between the input , the comprehensive weight of the DNN and the output . The comprehensive weight is composed of the weight in the dimension reduction network, the FC network weight in the reconstruction network and the convolutional layer weight , denotes the Frobenius norm.
[0053] In the end-to-end DNN architecture, the comprehensive weight of the DNN will be trained by formula (28), and the Adam algorithm is used to optimize it.
[0054] (28) wherein, , represent the adaptive functions of the dimension reduction network and the reconstruction network, respectively. In order to extract the estimated channel parameters, the estimated channel is expressed as a tensor model. By standard multi-component decomposition of the tensor, the corresponding AoAs and AoDs are obtained.
[0055] For the channel matrix , all time slots can be connected to obtain Each subcarrier is then built as a third-order tensor which can be expressed as: (29) where, is a vector related to the time slot, which contains the complex gain connected to the th path. The third mode of is related to the time slot, and by superimposing multiple time slots, each is decoupled into three independent dimensions, and then the factor matrix of each mode can be expressed as: (30) However, under the basic condition, the standard multi-component decomposition is unique, that is, there is a correlation between the estimated factor matrix and the actual factor matrix, as shown in equation (31): (31) where, represents an uncertain permutation matrix, denotes the estimation error related to the three estimated factor matrices, is a non-singular diagonal matrix, which satisfies After tensor decomposition, the extracted AoA and AoD can be expressed as: (32) (33) In order to solve the problem of joint optimization of hybrid beamforming and power allocation caused by imperfect CSI, the DNN composed of dimension reduction network and reconstruction network based on automatic encoder is adopted to realize the synchronous implementation of adaptive pilot design and CSI estimation. In addition, in order to extract CSI, the channel to be estimated is decomposed into a third-order linear tensor model containing three factor matrices, and the structure characteristics of the factor matrix are used to decompose the standard multi-component tensor to realize the extraction of AoA and AoD, and the structure of the algorithm is as shown in Figure 4 Algorithm 1 introduces the specific process of the CSI estimation algorithm based on the improved DNN in this section.
[0056] B. Hybrid beamforming and power allocation joint optimization algorithm based on PSO-MADDQN Due to the high dimensionality of the parameter space, the number of hybrid beamforming matrices and power allocation vectors is huge, which limits the exploration ability of MADDQN in high-dimensional space. PSO has the advantages of fast convergence and can avoid local optimum, so in the present application, PSO algorithm will be used to generate the initial hybrid beamforming matrix and power allocation vector. Therefore, the present application will further propose a joint optimization algorithm combining PSO and improved MADDQN to maximize the global search ability of PSO algorithm and the flexibility and adaptability of MADDQN in dynamic environment.
[0057] State of the agent Including CSI , the current hybrid beamforming matrix And the power allocation vector , can be expressed as , the action of the agent Adjust the hybrid beamforming matrix and the power allocation vector, which can be expressed as . Assuming the position of the particle is , wherein represents the hybrid beamforming matrix, represents the power allocation vector. The velocity of the particle is initialized as a random value. Then, the fitness function of the particle is defined as the system energy efficiency, that is, the weighted total power collected by all EH receivers. The fitness function can be expressed as: (34) , wherein represents the energy collected by the th EH receiver, represents the predefined power weight.
[0058] In each iteration of PSO, the particle will update its position and velocity according to its own experience and group experience respectively. The update formula of the velocity and position of the PSO particle can be expressed as: (35) (36) , wherein and respectively represent the velocity and position of the particle in the th iteration; is the inertia weight factor; represents the individual optimal position of the particle in the th iteration; represents the global optimal position of the group in the th iteration; and are learning factors, and are two random numbers within the range of
[0059] To further improve the search efficiency, the present application uses the adaptive PSO algorithm to dynamically adjust the inertia weight and learning factor based on the current search state in the process of updating the particle speed and position, which enhances the exploration ability of the algorithm. Then, the learning factor and the inertia weight factor can be re-expressed as: (37) (38) (39) wherein, and respectively represent the upper and lower bounds of the inertia weight factor. In order to ensure the stability of the particle convergence, the inertia weight factor should be avoided to be set too high or too low. and are the upper and lower limits of the individual learning factor; , represent the upper and lower limits of the group learning factor, respectively. is a nonlinear factor, wherein, and respectively represent the existing iteration number and the maximum iteration number.
[0060] In order to meet the power constraint and the rate constraint, the penalty function method is used to introduce the constraint condition into the fitness function. For the power constraint, a penalty term can be introduced: (40) wherein, α is the penalty coefficient. The final fitness function can be expressed as: (41) The reward function of the agent is defined as the improvement of the system energy efficiency. For the EH receiver, the reward function can represent the increment of the energy obtained by the EH receiver, as shown in equation (42): (42) Then, for the ID receiver, the reward function can be expressed as the sum rate increment of the ID receiver: (43) MADDQN learns optimal hybrid beamforming and power allocation strategies through interaction with the environment. At each time step, the agent selects an action based on the current state, and obtains a new state and reward after executing the action. However, the action selection and feedback among multiple agents are extremely complex, leading to significant errors in Q-value estimation and reducing the convergence speed of MADDQN. Therefore, this invention utilizes a second-order time difference method to reconstruct the loss function to improve the algorithm's convergence. In the MADDQN structure with second-order time difference, random sampling is performed through empirical replay, and training is conducted using the second-order time difference function to realize the weights. The update. The calculation method for the second-order time difference is shown in formula (44): (44) in, Indicates the state Take action When, the value function is at the th Round using weight parameters The estimated Q value; It is in state Take action When, the value function is at the th Round using weight parameters The estimated Q value; It is in state When taking action The value function in the first place Round using weight parameters The estimated Q value. The parameter, ranging from 0 to 1, is used to distinguish the difference between the two weighted estimation functions. Used to distinguish the difference between two weighted estimated functions
[45] .
[0061] Next, These are considered as the weight parameters of the first layer target network (MADDQN-1) of MADDQN. and These represent the weight parameters of the value network and the target network in the second layer of the MADDQN network (MADDQN-2), respectively. During each stage of model training, the parameters of the value network in MADDQN_2 are updated simultaneously with the values in MADDQN_1. , ,in, This represents the weight parameter of the next state in MADDQN_2, each The next step is to pass the parameters of the value network in MADDQN_1 to the target network, i.e. At the same time, the parameters of the value network in DDQN_2 are passed to the target network, that is . Therefore, the improved loss function can be expressed as: (45) where, is the parameter of the model, represents all possible states , actions , rewards and next states expectations; represents the Q value of the second estimation function calculated using the target network; represents the Q value of the first estimation function calculated using the target network; represents the Q value of the second estimation function calculated using the current network. The loss function describes the difference between the two estimated value functions and the variance of the parameter values of the two networks. Solving the inverse of the loss function in formula (45) and the performance, the weight gradient is as shown in formula (46): (46) The weight gradient is used to update the parameters of the value network, keeping the weight parameters of the current network unchanged. This helps the Q value network to better approximate the true value and prevents the instability of the algorithm convergence.
[0062] Although the traditional experience pool structure can effectively store historical information, its retrieval efficiency is low when dealing with large-scale data, which may cause a long time to recover historical information, thereby affecting the real-time performance and response speed of the system. Therefore, the binary tree structure is used to replace the traditional experience pool structure to store the results obtained by the above-mentioned second-order time difference method, so as to reduce the time consumed in recovering historical information. In order to simply express the binary tree formula, the present application will the second round and the first round of the moment estimation function are denoted as , that is . .
[0063] In the present application, the value of each binary tree node will be measured by , and its calculation method can be expressed as formula (47): (47) where, is a very small constant, , represents the capacity of the binary tree.
[0064] After the fusion of the second-order time-difference method and the binary tree storage structure with the MADDQN method, the principle of the priority sampling of this method considers that the priority of the leaf node increases with the increase of its value, at which time the probability of the selection of this leaf node will also increase. However, since the priority selection may cause the uneven distribution of the results obtained by the algorithm, leading to outliers and premature convergence. In order to improve the robustness in the selection process, this section further introduces a new sampling method, which estimates some distribution properties by regulating the probability distribution to reduce the variance. The calculation method of the importance sampling weight is as follows: (48) wherein represents the importance sampling weight, is the current time, represents the priority at the current sampling time, is the probability of the sampling of the leaf node. This method well improves the robustness of the sampling, while reducing outliers.
[0065] After that, the neural network parameters will be updated using backpropagation and formula (49): (49) wherein represents the updated network parameters after iterations, is the current parameter, represents the importance sampling weight, and indicates the time error. This step is used to adjust the parameters of the neural network to improve the performance of the algorithm.
[0066] After obtaining the updated neural network parameters, the strategy can be estimated using the approximation of the action value function. If the next iteration does not reach the final state, the action of the next iteration can be determined according to , and the algorithm will continue to execute the loop until the final state is achieved.
[0067] In the process of PSO and DDQN, the continuous optimal solution and the discrete optimal solution centered at the current position are calculated using PSO and DDQN, and then the weighted fusion is performed using the following formula: (50) wherein and represent the positions of the agent at the current and next steps, respectively. and denotes the weight of the corresponding candidate solution, and is calculated as follows: (51) (52) wherein and are the particle positions at the next time calculated by PSO and DDQN based on the existing particle positions, denotes the evaluation function.
[0068] The present application takes the PSO algorithm as a global optimization tool, globally searches for the optimal hybrid beamforming and power allocation scheme by iteratively searching for the particle swarm and using the fitness function to measure the energy efficiency of the system, thereby providing stronger global exploration capability in the optimization process. Secondly, the second-order time difference method for computing load optimization is introduced in MADDQN to avoid the problem of overestimation in the traditional method. Finally, the experience replay is stored in a binary tree structure, which improves the utilization efficiency of historical information and reduces the computational overhead. The structure of the algorithm is shown in Figure 5 , and the specific process of the algorithm is shown in Algorithm 2.
[0069] IV. Simulation analysis The present application proposes a hybrid beamforming and power allocation joint optimization algorithm based on DNN and MADDQN for the energy efficiency maximization problem in the terahertz UM-MIMO-ISCAP system. In order to explore the influence of the algorithm on the energy efficiency in the terahertz UM-MIMO-ISCAP system, energy efficiency comparison experiments of different algorithms, and experiments on the influence of iteration number, signal-to-noise ratio and number of transmit antennas on system energy efficiency performance are designed. In the simulation experiment of the present application, based on the hardware platform of Intel Core i7-10700 CPU @ 2.90GHz (16 CPUs), Radeon RX 550X (4GB video memory), 16GB DDR4 memory, Windows 10 64-bit operating system is adopted, and relies on PyCharm 2022.1.2 (Python3.10) software platform to carry out simulation experiment. The specific parameter settings of the system simulation are shown in Table 1, the hyperparameter settings of DNN are shown in Table 2, and the hyperparameter settings of MADDQN are shown in Table 3.
[0070] Table 1 System simulation parameters Table 2 DNN parameter design Table 3 PSO-MADDQN network parameter design The discount rate is one of the important hyperparameters in MADDQN, so the present application sets the discount rate to 0.7, 0.8, 0.9 and 0.99 in the simulation experiment to verify the change trend of the total energy efficiency of the proposed algorithm with the number of iterations, as shown in FIG. 6. When the discount rate is 0.7, the overall energy efficiency of the system is the lowest, and there is still a downward trend when the number of iterations reaches 100. In addition, when the discount rate is 0.99, the total user energy efficiency is the highest, and gradually converges with the increase of the number of iterations. There are two reasons for this result. On the one hand, the action of the agent will affect the future instantaneous reward, and then affect the cumulative discounted reward, and the gradually increasing discount rate will further make the agent pay more attention to the influence of the next moment reward on the current moment reward, so the total energy efficiency will become more favorable with the increase of the discount rate. From another point of view, the increase of the discount rate will increase the proportion of the agent's prediction of the next moment reward, which promotes the stability of the total energy efficiency of the system with the number of iterations. Therefore, in the simulation of the present application, the discount rate is set to 0.99, and under this condition, the proposed algorithm can achieve the optimal system total energy efficiency and gradually stabilize. Figure 6
[0071] In order to explore the change of the loss value of the algorithm proposed in the paper with the number of iterations, experiments are designed to focus on the influence of different learning rates on the convergence speed and stability of the algorithm, as shown in FIG. 7. The learning rate parameters are set to 0.05, 0.005 and 0.0005, respectively. The change trend of the loss value of the proposed algorithm in the training process with the number of iterations is shown in FIG. 7. When the learning rate is 0.05, the curve fluctuates violently in the early training period, and the average loss peak value even reaches 5 or more, and in the later period, it oscillates between 0 and 3, indicating that a large learning rate will lead to obvious instability. When the learning rate is 0.005, it is relatively balanced, and the average loss fluctuates between 0 and 2, and there is still a certain amplitude of jitter in the training process, but the overall trend gradually decreases. When the learning rate is 0.0005, the loss curve is at a relatively low level and fluctuates between 0 and 1, which indicates that the training process at this time is relatively smooth and more stable. A smaller learning rate can reduce the parameter update step and avoid large fluctuations and deviations. For the terahertz UM-MIMO-ISCAP system used in the present application, smooth training is more stable for the algorithm to converge to the optimal solution, reduces a large number of network parameter updates, and thus avoids local oscillation. Figure 7
[0072] Figure 8 It can be seen that the MSE of all algorithms decreases with the increase of SNR, because higher SNR means clearer signal reception, so the accuracy of CSI estimation will be improved and the MSE will be reduced. Compared with the higher error and smaller decline of CS and LS, the DNN-based CSI extraction algorithm still maintains the lowest error, because the algorithm realizes the automatic learning of CSI in the noise environment through end-to-end training. This simulation result shows that the DNN method can more accurately reconstruct CSI when facing different SNR channel estimation problems, and has higher adaptability and robustness. In addition, the DNN-based algorithm converts the channel into the frequency-angle domain and uses the sparsity of the channel to compress the redundant information, reducing the interference of noise. This shows that DNN can continuously provide more accurate CSI estimation when dealing with channel estimation problems in high SNR environment. The reason why the advantages of LS and CS, two traditional methods, gradually disappear under high SNR conditions is that the increase in the number of antennas in the terahertz UM-MIMO system leads to an increase in the dimension of the channel matrix and the non-Gaussian nature of near-field interference reduces sparsity and robustness. In summary, in a complex channel environment, the advantage of the CSI estimation algorithm based on the DNN algorithm lies in its strong signal feature extraction capability and robustness, especially in channel estimation, which can learn the underlying signal rules through complex models to achieve low CSI estimation error.
[0073] As shown in Figure 9 , it can be seen that the energy efficiency of all algorithms tends to be stable with the increase of iteration times. Compared with AO, SDR, Q-learning, Greedy, Random and other algorithms, the PSO-MADDQN algorithm proposed in this chapter performs best and can quickly converge and maintain high energy efficiency in a short iteration period. The performance of AO, SDR, Q-learning, Greedy, Random and other algorithms is relatively close, and the energy efficiency improvement is relatively small, and in a longer training process, their energy efficiency has almost no further improvement after reaching a certain level. This is because these algorithms lack global search and adaptive ability in optimization selection, and Q-learning and greedy algorithm converge slowly due to high-dimensional action space and sparse reward problems. This chapter combines the initial solution generation ability of PSO with the dynamic optimization ability of MADDQN to avoid local optimal problems. In addition, through the second-order time difference and binary tree experience pool, the convergence speed of DDQN is significantly improved.
[0074] When the learning rate is 0.005, as shown in Figure 9 (a), the curve rises rapidly at the beginning of the iteration of the algorithm, indicating that the algorithm can quickly approach the optimal solution at the beginning, and the energy efficiency of the PSO-MADDQN algorithm can reach 11.28 bps / Hz / W.Figure 9 As shown in (b), the energy efficiency of each algorithm is generally higher when the learning rate is 0.0005 than when the learning rate is 0.005, and the energy efficiency of PSO-MADDQN can reach 12.97 bps / Hz / W, which indicates that the training process is relatively smooth and more stable at this time, and the convergence is more rapid. This shows that a smaller learning rate makes the updating process of the algorithm more stable, avoiding excessive jumping, while improving the accuracy and convergence speed. As the learning rate decreases, the training process becomes more stable, avoiding excessive oscillation around the optimal solution, thereby achieving more detailed optimization.
[0075] Table 4 Comparison of running time of each iteration (minutes) of different algorithms Table 4 compares the average calculation time required for each iteration of different algorithms, and it can be clearly seen that the PSO-MADDQN algorithm proposed in this chapter has the shortest running time, while the alternating optimization and SDR have relatively long calculation time for each iteration, because the time consumption of alternating optimization and SDR will present exponential growth with the dimension of variables when dealing with large-scale matrix operations. The calculation process of Greedy algorithm and Random strategy is relatively simple, so the calculation time is the shortest. However, these algorithms lack efficient power allocation and hybrid beamforming strategies, resulting in poor energy efficiency performance. On the other hand, although the IMADDQN algorithm proposed in the third chapter has relatively short running time, it still lags behind the PSO-MADDQN algorithm proposed in this chapter, because the algorithm uses second-order time difference method to reconstruct the loss function in the training process of MADDQN model, and uses binary tree structure to store experience replay, improving the utilization efficiency of historical information, thereby reducing the calculation overhead and running time. The hybrid beamforming and power allocation joint optimization algorithm based on PAO-MADDQN proposed in this chapter has longer calculation time than Random and Greedy algorithms due to the calculation load of MADDQN joint optimization. However, through efficient joint optimization of power allocation and hybrid beamforming, it can significantly improve the total energy efficiency of the system while ensuring shorter calculation time, which shows that the algorithm proposed in this chapter can efficiently handle complex optimization tasks.
[0076] In order to consider the total energy efficiency performance of the UM-MIMO-ISCAP system under different numbers of EH receivers. Set the number of EH users of the system to {2, 3, 4, 5, 6}, and in addition to the three EHs set at the beginning, the newly added EH receivers are uniformly distributed in the radius range of The spatial angle is 20 dB, the number of runs is 300, and the number of iterations is 50. The system can reach a rate of 5.0 bps / Hz. As the number of EH receivers increases, the energy efficiency of all algorithms will increase, as shown in Figure 10 Compared with other algorithms, the PSO-MADDQN algorithm proposed in this chapter performs best and is significantly better than the algorithm proposed in Chapter 3. This is because AO relies on step-by-step iteration, and SDR needs to make a convex approximation to a non-convex problem, which will cause the dimension to increase when the number of EH receivers increases, resulting in a deviation from the global optimum. At the same time, due to the dimension disaster in high-dimensional space of Q-learning and the greedy algorithm rule only relying on local optimal strategy, it cannot effectively coordinate the resource competition among multiple users, resulting in lower energy efficiency than our proposed algorithm. Compared with the algorithm proposed in Chapter 3, the PSO-MADDQN algorithm proposed in this chapter uses PSO to generate hybrid beamforming and power allocation candidate solutions in the initial stage, and dynamically adjusts the strategy according to the CSI by MADDQN, so that the algorithm can efficiently explore the high-dimensional space when the number of EH users increases. In addition, when the number of users increases, the combination of binary tree structure and limited sampling speeds up the convergence and avoids the performance degradation of traditional Q-learning due to sparse rewards. Therefore, the PSO-MADDQN algorithm can perform better adaptability and efficiency and energy efficiency when dealing with the joint optimization problem of hybrid beamforming and power allocation in a multi-user environment.
[0077] As shown in Figure 11 , the energy efficiency of all algorithms increases significantly with the increase of the transmit power at the base station. Alternating optimization and SDR also perform relatively well at higher transmit power, but their energy efficiency is still lower than that of PSO-MADDQN. The Q-learning algorithm performs well at low power, but as the transmit power increases, the problems of dimension disaster and low exploration efficiency limit its energy efficiency improvement. The greedy algorithm and the random algorithm rely too much on local optimal solutions and random strategies, and their performance improves slowly. Their energy efficiency lags behind other algorithms as the transmit power increases. The PSO-MADDQN algorithm proposed in this chapter has significantly improved performance compared to the algorithm proposed in Chapter 3. This is because PSO-MADDQN uses an adaptive penalty function to avoid power overflow. In addition, the algorithm uses second-order time difference and importance sampling weight to prioritize learning key features and maintain the stability of the strategy at high transmit power, thereby improving the energy efficiency of the system. In summary, the experimental results of the simulation prove that the PSO-MADDQN algorithm can still maintain high energy efficiency at high base station transmit power, and it is superior to other traditional algorithms in terms of energy efficiency at high power, demonstrating strong robustness and adaptability.
[0078] Figure 12The simulation results show the comparison of the total energy efficiency of different algorithms under the change of achievable sum rate and sum rate constraint. With the increase of achievable sum rate and sum rate, the energy efficiency of all algorithms decreases, because the higher the sum rate requirement is, the more energy is allocated to ID receivers, which leads to the decrease of energy efficiency. From the performance of each algorithm, the alternating optimization and SDR algorithms have higher energy efficiency at low sum rate constraint, but their energy efficiency performance gradually decreases with the increase of sum rate. In contrast, the performance of the Q-learning algorithm is relatively poor, because when the achievable sum rate and sum rate requirement are high, Q-learning will face the problem of dimension disaster and inefficient exploration, which leads to the rapid decline of its energy efficiency. The greedy algorithm performs poorly at high sum rate constraint because it only relies on local optimal decision, ignoring the effective coordination among multiple users in the resource allocation process. The PSO-MADDQN algorithm proposed in this chapter performs best in the entire sum rate range, effectively balancing the sum rate requirement and energy efficiency, and the decline of energy efficiency is relatively small with the increase of sum rate requirement. The performance of the proposed PSO-MADDQN algorithm is improved, and the decline of energy efficiency is slower, because the algorithm uses the second-order time difference method to compare the action estimation error of the previous two rounds to realize the dynamic adjustment of network parameters. In addition, PSO-MADDQN prioritizes learning high-error samples through binary tree experience pool to accelerate the convergence of key strategies, thereby reducing the conflict between energy efficiency and sum rate and slowing down the decline of energy efficiency. At the same time, Figure 12 The simulation comparison of (a) and (b) shows that when the number of antennas is 1024, the total energy efficiency of the system is higher than that when the number of antennas is 64, and the decline of energy efficiency is smaller than that when the number of antennas is 64. This is because when the number of antennas is 64, the hybrid beamforming technology that gives priority to precision and rate improvement is adopted, which weakens the interference suppression ability and causes the rapid decline of energy efficiency in the case of high sum rate. On the contrary, when the number of antennas increases to 1024, the increase of the number of antennas expands the spatial degrees of freedom, enabling the system to more efficiently allocate power in the case of high sum rate, thereby reducing energy loss and delaying the decline of energy efficiency.
[0079] To study the impact of noise on system energy efficiency, the total energy efficiency of different algorithms is compared with the change of noise. As shown in Fig. 12, the energy efficiency of all algorithms decreases with the increase of noise. The energy efficiency of the alternating optimization and SDR algorithms is higher at low noise, but their energy efficiency performance gradually decreases with the increase of noise. In contrast, the performance of the Q-learning algorithm is relatively poor, because when the noise is high, Q-learning will face the problem of dimension disaster and inefficient exploration, which leads to the rapid decline of its energy efficiency. The greedy algorithm performs poorly at high noise because it only relies on local optimal decision, ignoring the effective coordination among multiple users in the resource allocation process. The PSO-MADDQN algorithm proposed in this chapter performs best in the entire noise range, effectively balancing the noise requirement and energy efficiency, and the decline of energy efficiency is relatively small with the increase of noise requirement. The performance of the proposed PSO-MADDQN algorithm is improved, and the decline of energy efficiency is slower, because the algorithm uses the second-order time difference method to compare the action estimation error of the previous two rounds to realize the dynamic adjustment of network parameters. In addition, PSO-MADDQN prioritizes learning high-error samples through binary tree experience pool to accelerate the convergence of key strategies, thereby reducing the conflict between energy efficiency and sum rate and slowing down the decline of energy efficiency. At the same time, Figure 13As shown, the energy efficiency of all algorithms shows a clear downward trend with the increase of noise. This is because as the signal strength increases, the transmission efficiency of the system gradually approaches its theoretical limit, and the sensitivity to signal noise decreases, making the energy efficiency tend to be stable. Specifically, the PSO-MADDQN algorithm proposed in this chapter can maintain a relatively high energy efficiency, especially in the condition of low noise amplitude. The energy efficiency of the PSO-MADDQN algorithm is relatively high, and as the noise increases, the energy efficiency decreases relatively small. This is because the algorithm proposed in this chapter learns power enhancement actions by importance sampling and adaptive penalty function when the noise changes, thereby keeping a small energy fluctuation. In contrast, the alternating optimization algorithm has a higher energy efficiency in low noise conditions, but its energy efficiency decreases significantly as the noise amplitude increases. This is because the alternating optimization algorithm relies on step-by-step iteration, and each time only one variable can be optimized, which will cause its local optimal solution to deviate from the global optimal solution, resulting in a decrease in energy efficiency performance. In summary, the PSO-MADDQN-based algorithm can maintain good energy efficiency performance under different noise conditions, especially in low noise, showing strong adaptability and robustness, while other algorithms face different degrees of performance decline in high noise amplitude environment.
[0080] As Figure 14As shown, the energy efficiency of all algorithms is increasing with the increase of the number of antennas, which shows that the increase of the number of antennas can significantly improve the energy efficiency of the system, especially in a multi-user system. When the number of antennas increases from 64 to 1024, the total energy efficiency of the PSO-MADDQN algorithm steadily increases from 8.5 bps / Hz / W to 12.2 bps / Hz / W, which is significantly better than other algorithms. The alternating optimization, SDR and other algorithms have limited performance improvement in the super large antenna scenario. This is because in the super large antenna scenario, the coupling relationship between the hybrid beamforming matrix and the power allocation becomes very complex, so the iteration of the alternating optimization converges very slowly. For SDR, its convex approximation method will accumulate errors with the increase of the number of antennas, so that the final solution deviates from the global optimum, and therefore the energy efficiency improvement effect is not ideal. Although Q-learning and greedy algorithm perform well when the number of antennas is small, their energy efficiency increases gradually decrease with the increase of the number of antennas. This is because Q-learning has low exploration efficiency in high-dimensional space, and the greedy algorithm only relies on local optimal solution and cannot effectively coordinate the resource competition among multiple users, thereby affecting the further improvement of energy efficiency. The algorithm proposed in this chapter uses the second-order time difference method to dynamically compare the action estimation errors of the previous and next rounds, optimizes the super large antenna hybrid beamforming weight in priority, and combines the binary tree experience pool to speed up the convergence of the jamming suppression action, avoiding the divergence problem of DDQN in high-dimensional space in the third chapter. In summary, with the increase of the number of antennas, the PSO-MADDQN algorithm shows significant advantages in the multi-antenna system through its global search and dynamic adjustment strategy, and can continuously improve the energy efficiency in a large-scale antenna environment.
[0081] It should be noted that the embodiments of the present application can be realized by hardware, software or a combination of software and hardware. The hardware part can be realized by special logic; the software part can be stored in a memory and executed by a suitable instruction execution system, such as a microprocessor or a specially designed hardware. Those skilled in the art can understand that the above-mentioned devices and methods can be realized by computer executable instructions and / or included in processor control code, such as provided on a carrier medium, such as a magnetic disk, CD or DVD-ROM, a programmable memory, such as a read-only memory (firmware), or a data carrier, such as an optical or electronic signal carrier. The devices of the present application and their modules can be realized by hardware circuits, such as very large scale integrated circuits or gate arrays, semiconductors, such as logic chips, transistors, etc., or programmable hardware devices, such as field programmable gate arrays, programmable logic devices, etc., by software executed by various types of processors, or by a combination of the above-mentioned hardware circuit and software, such as firmware.
[0082] The above merely illustrates the specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any modification, equivalent replacement and improvement within the technical range disclosed by the present application and within the spirit and principle of the present application should be covered within the protection scope of the present application.
Claims
1. A method for joint optimization of beamforming and power allocation based on deep reinforcement learning, characterized in that, The method comprises the following steps: a. Constructing a unified modeling far and near field super large scale multiple input multiple output communication sensing and energy transmission integrated system; b. Using a deep neural network composed of a dimension reduction network and a reconstruction network to perform pilot design and channel estimation on the received echo signal, and obtaining channel state information in the frequency-angle domain; c. Performing three-dimensional tensor decomposition on the channel state information to extract the angle of arrival and the angle of departure parameters; d. Inputting the parameters into a multi-agent double deep Q network to dynamically update the beamforming matrix and the power allocation vector through reinforcement learning; e. Generating an initial beamforming matrix and power allocation vector using a particle swarm optimization algorithm; f. Introducing a second-order time difference strategy and a binary tree experience pool mechanism based on reward error ranking during the training process of the multi-agent double deep Q network to obtain the converged joint optimization result.
2. The method of claim 1, wherein, The particle swarm optimization algorithm uses global optimal particles and individual optimal particles for iterative updating in the speed updating process, and limits the range of particle search space in the position updating stage to control the optimization boundary.
3. The method of claim 1, wherein, The binary tree experience pool prioritizes the reward error of historical training samples for priority scheduling, and preferentially samples samples in the high reward fluctuation interval for training.
4. The method of claim 1, wherein, The second-order time difference strategy calculates the target Q value in the forward view form, and embeds the multi-step reward sequence into the reinforcement learning loss function to enhance the convergence performance.
5. A system for joint optimization of beamforming and power allocation, the system comprising: It comprises: a. A modeling unit for constructing a far and near field unified modeling super large scale MIMO communication sensing and energy transmission system; b. A channel estimation unit for extracting frequency-angle domain channel state information using a deep neural network; c. A parameter extraction unit for performing three-dimensional tensor decomposition on the channel state information and extracting the angle of arrival and the angle of departure parameters; d. A reinforcement learning unit for generating a beamforming matrix and a power allocation vector using a multi-agent double deep Q network; e. A heuristic initialization unit for generating an initial solution of beamforming and power allocation based on a particle swarm optimization algorithm; f. A learning scheduling module for introducing a binary tree experience pool and a second-order time difference mechanism to improve training efficiency and stability.
6. The system of claim 5, wherein, The reinforcement learning unit includes a policy update module, a gradient backpropagation module, and a parameter synchronization update module of a double network architecture, and supports distributed training deployment.
7. A base station device, characterized by comprising: The optimization system of claim 5 is integrated, and the final beamforming matrix is executed through a radio frequency front-end circuit and an antenna array to complete data communication, sensing tasks, and directional energy transmission.
8. A method for angle parameter extraction based on three-dimensional tensor decomposition, characterized in that, It comprises: Receiving frequency-angle domain channel state information, constructing a three-dimensional channel tensor, and using CP or Tucker decomposition algorithm to extract channel angle of arrival and angle of departure for subsequent feature input of reinforcement learning network.
9. A method for experience sample scheduling for a reinforcement learning training process, the method comprising: By constructing a binary tree data structure based on the size of the reward error, dynamic priority is assigned to the experience samples, and experience data is sampled according to the priority for gradient update during the training phase.
10. An optimization method combining particle swarm optimization and deep reinforcement learning, characterized in that, First, an initial beamforming and power allocation solution is generated using a particle swarm optimization algorithm, and then the initial solution is used as the starting point of exploration for the reinforcement learning network to obtain the optimal solution through policy iteration.