Hybrid precoding method based on low-dimensional discrete soft actor commentator

Through a hybrid precoding method based on low-dimensional discrete soft actor critics, combining channel estimation and prehaul capacity limitation under the C-RAN architecture, a hybrid precoder is designed to solve the problems of constant mode constraints and large variable space in the simulation domain, and efficient channel utilization and system capacity improvement are achieved.

CN120049925APending Publication Date: 2025-05-27UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510274852.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing hybrid precoding technology is difficult to effectively solve the problems of constant mode constraints and large variable space in the analog domain. At the same time, there is the coupling problem between the digital domain and the analog domain, and the deep reinforcement learning algorithm is poor in general use.

Method used

Using a hybrid precoding method based on low-dimensional discrete soft actor critics, a hybrid precoder is designed to explore and maximize the trade-offs of optimization goals through low-dimensional discrete soft actor critic D-SAC reinforcement learning algorithm, combining channel estimation and prefab capacity limitations under the C-RAN architecture.

Benefits of technology

It effectively solves the problems of constant mode constraints and large variable space in the simulation domain, reduces interference between users, increases the total capacity of the system, and improves the generality of the algorithm through reinforcement learning algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120049925A_ABST
    Figure CN120049925A_ABST
Patent Text Reader

Abstract

The invention discloses a hybrid precoding method based on a low-dimensional discrete soft actor commentator, which comprises the following steps of: obtaining an analog precoder and a service user selected by each remote radio head (RRH) according to channel state information (CSI) of each user and each RRH and a D-SAC reinforcement learning algorithm of the low-dimensional discrete soft actor commentator; pilot channel estimation is carried out according to the analog precoder selected by each RRH, and an equivalent estimation channel from each RRH to a corresponding service user is obtained; and obtaining a digital precoder and a forward compression noise covariance matrix in a central baseband processing unit (BBU) through a fractional programming algorithm on the basis of a user of each RRH selection service and an equivalent estimation channel, thereby completing the design of the hybrid precoder. According to the method, channel estimation, forward capacity limitation under a C-RAN architecture and tradeoff of exploration and maximization optimization targets are considered at the same time, so that the problems of constant modulus constraint and large variable space of a simulation domain are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technology, and in particular to a hybrid precoding method based on low-dimensional discrete soft actor critic. Background Art

[0002] For the practical implementation of millimeter-wave massive Multiple-input Multiple-output (mmWave mMIMO) systems, energy-efficient hybrid precoding architectures significantly reduce the number of radio frequency (RF) chains to provide sufficient channel gain while maintaining affordable hardware cost and energy consumption. Therefore, the design of hybrid precoders is a classic problem in communication systems. The main challenge in designing hybrid precoders is that analog precoding is implemented using constant modulus phase shifters, which brings non-convex constant modulus constraints to hybrid precoding. At the same time, future communication systems may adopt a cloud radio access network (C-RAN) architecture. In C-RAN, baseband signal processing is migrated to a centralized baseband unit (BBU) in the cloud, and RF processing is performed in remote radio heads (RRHs). Therefore, when designing hybrid precoders, it is necessary to consider the coordinated interference management of different RRHs and the capacity limitations of the fronthaul link from the centralized baseband unit to multiple RRHs.

[0003] In response to the above-mentioned problems of hybrid precoding, the prior art has proposed relevant algorithm solutions, including an iterative algorithm based on the weighted minimum mean square error (WMMSE) method, but it needs to iteratively solve two convex optimization problems, which is highly complex. At the same time, the performance is degraded due to the quasi-convex treatment of the constant modulus constraint. Another solution is based on the beam scanning framework, which proposes to first determine the analog beamformer from the codebook, and then perform digital beamforming and compression design based on semi-definite relaxation (SDR) and convex approximation. It still directly splits the coupling problem into two sub-problems, resulting in a suboptimal solution. There is also a solution that performs precoding based on the deep reinforcement learning (DRL) algorithm under fixed CSI, and develops different distributed architectures at the same time. However, due to the fixed CSI, the algorithm has poor versatility and cannot be actually deployed.

[0004] In summary, although some work has attempted to design a hybrid precoder, there is still a lot of room for optimization in terms of how to solve the constant modulus constraint brought by the phase shifter, solve the coupling between the digital domain and the analog domain, and improve the versatility of the DRL algorithm. Summary of the invention

[0005] The purpose of the present invention is to provide a hybrid precoding method based on a low-dimensional discrete soft actor critic, which simultaneously considers the channel estimation, the fronthaul capacity limitation under the C-RAN architecture, the trade-off between exploration and maximization optimization objectives, so as to solve the problems of constant modulus constraints and large variable space in the analog domain, so that the analog domain variables can take into account both exploration and optimization.

[0006] The objective of the present invention is achieved through the following technical solutions:

[0007] A hybrid precoding method based on low-dimensional discrete soft actor-critic, the method comprising:

[0008] Step 1: According to the channel state information CSI between each user and each remote radio head RRH, a simulated precoder selected by each RRH and the user served are obtained according to the low-dimensional discrete soft actor critic D-SAC reinforcement learning algorithm;

[0009] Step 2: perform pilot channel estimation according to the analog precoder selected by each RRH to obtain an equivalent estimated channel from each RRH to the corresponding service user;

[0010] Step 3: Based on the users selected by each RRH and the equivalent estimated channel, the digital precoder and the fronthaul compression noise covariance matrix are obtained in the central baseband processing unit BBU through the fractional programming algorithm to complete the design of the hybrid precoder.

[0011] An electronic device comprises a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the method.

[0012] A computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method.

[0013] It can be seen from the technical solution provided by the present invention that the above method simultaneously considers the trade-offs between channel estimation, fronthaul capacity limitation under the C-RAN architecture, exploration and maximization optimization objectives, so as to solve the problems of constant modulus constraints and large variable space in the analog domain, so that the analog domain variables can take into account both exploration and optimization, thereby reducing interference between users and improving the total capacity of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required in the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0015] Figure 1 Schematic diagram of the hybrid precoding method based on low-dimensional discrete soft actor-critic provided by the embodiment of the present invention;

[0016] Figure 2 Schematic diagram of the low-dimensional D-SAC network structure described in the embodiment of the present invention;

[0017] Figure 3 Schematic diagram of the training process of low-dimensional D-SAC of an agent described in the embodiment of the present invention;

[0018] Figure 4 Schematic diagram of the offline deployment process of the method described in the embodiment of the present invention;

[0019] Figure 5 Schematic diagram of the online deployment process of the method described in the embodiment of the present invention;

[0020] Figure 6 Schematic diagram of the comparison of the total user rate with the change of SNR in the example of the present invention;

[0021] Figure 7 Schematic diagram of the comparison of the total user rate with the change of fronthaul capacity limit in the example of the present invention;

[0022] Figure 8 Schematic diagram of the comparison of the total user rate with the change of channel time slots in the example of the present invention. Detailed implementation manners

[0023] The following combines the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments, which does not constitute a limitation to the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0024] As Figure 1 shown, it is a schematic diagram of the hybrid precoding method based on low-dimensional discrete soft actor-critic provided by the embodiment of the present invention, and the method includes:

[0025] Step 1: According to the channel state information (CSI) between each user and each remote radio head (RRH), based on the low-dimensional discrete soft actor-critic (D-SAC) reinforcement learning algorithm, obtain the analog precoder selected by each RRH and the served users;

[0026] In this step, the low-dimensional discrete soft actor-critic (D-SAC) reinforcement learning algorithm includes a Markov decision process (MDP), a low-dimensional D-SAC network structure, and a training process of low-dimensional D-SAC. Among them: The Markov decision process includes an agent, a state, an action, a reward, and a state transition. The analog precoder selected by each RRH and the served users are equivalent to the beam-user pairs selected by each agent. Specifically:

[0027] Agent: Consider each base station as an agent, and interact with the environment based on the policy π m to observe the current state s t,m of the beam-user pair a t,m , where is the set of base stations; and obtain the reward r t,m .

[0028] State: The system model contains K users and N beam directions. The state of each agent includes two parts: 1) According to the Toeplitz property of the channel autocorrelation matrix, the local channel state information CSI statistical value R m formed by the diagonal elements of the channel autocorrelation matrix of each RRH for each user, which is expressed as: represents the real number field; 2) The currently accumulated selected beam-user pairs If the element g m in the n-th row and k-th column of G nk = 1, it means that the k-th user and the n-th beam are selected;

[0029] Action: The action of each agent is all the beam-user pairs it can select, that is, the action where B is the number of alternative beams, and it is reduced to through the Softmax function of a deep neural network (DNN). Then the coupling relationship of the beam-user pairs is implicit in the reward function;

[0030] Reward: The reward of each agent is the same. When the time step t is less than the maximum time step L, the reward is 0; when t = L, there is a total reward r L , which is expressed as:

[0031]

[0032] where ω k represents the weight of user k; R k represents the communication rate of user k;

[0033] State transition: Each experience of each agent has L time steps, and at each time step, a different beam-user pair a t,m is selected to change the G t+1,m part in the state s m at the next moment.

[0034] In specific implementation, as Figure 2 shown in the schematic diagram of the low-dimensional D-SAC network structure described in the embodiment of the present invention, the low-dimensional D-SAC network structure includes an actor network and a critic network. Each agent has a corresponding actor-critic network. In the m-th agent, the actor network inputs the state s t,m at the current time step t, and then outputs a B×1 beam selection probability vector and a K×1 user selection probability vector through two Softmax functions respectively. According to these probability vectors, the beam direction a t0,m and the served user a t1,m at the current time step t are selected. Then, the beam-user pair a t,m is expressed as:

[0035] a t,m ={a t0,m , a t1,m};

[0036] In the m-th agent, the critic network inputs the state s t,m at the current time step t and the selected beam-user pair a t,m to obtain its Q value, that is, the long-term expected return. Since the action is discrete, the critic network inputs the state s t,m at the current time step t and outputs the Q value of each a t,m , and finally outputs a Q value vector; the critic network predicts the Q values of all actions through two neural networks;

[0037] Generally speaking, the beam selection probability vector output by the m-th actor network at the current time step t is The user selection probability vector is s t is the state vector at the current time step t, represents a one-dimensional vector with a length of B, represents a one-dimensional vector with a length of K; the Q value vector output by the m-th critic network at the current time step t is

[0038] Among them, the beam selection Q-value vector is The user selection Q-value vector is θ i , where i = 1, 2 represents the double Q-network, that is, the Double Deep Q-Network (DDQN) structure.

[0039] In specific implementation, as Figure 3 shown in the schematic diagram of the training process of the low-dimensional D-SAC of an agent described in the embodiment of the present invention, the training process of the low-dimensional D-SAC is specifically as follows:

[0040] The agent continuously interacts with the environment through the actor network, generating multiple (s t , a t , r t , s t+1 ) samples and storing them in the replay buffer; for the convenience of expression, here s t is used to replace s t,m to represent the state, r t is used to replace r t,m to represent the reward, m represents the agent, and the other parameters are the same;

[0041] Then, a small batch of samples are taken from the replay buffer to train the actor network, the critic network, and the entropy temperature coefficient, and the entropy temperature coefficient is used to balance the ratio of network exploration and exploitation; since we have made a low-dimensional improvement to the D-SAC algorithm, the loss functions of its actor network and critic network also need to be improved. The loss function J π (φ) is expressed as:

[0042]

[0043] Where represents the expectation; φ represents the trainable parameters of the actor network; α is the entropy temperature coefficient; D represents the replay memory of all previous state-action-reward-next state experiments;

[0044] The loss function J Q (θ i ) is expressed as:

[0045]

[0046] Where, represents two target Q-networks, which are periodically updated from the original double Q-network θ i , i = 1, 2; represents the Q-network function; r represents the reward; η represents the discount factor; represents the Q-value corresponding to the state s and the action a;

[0047] The loss function \(L(\alpha)\) of the entropy temperature coefficient is expressed as:

[0048]

[0049] where is the target entropy; \(H(\pi(\cdot|s t )) = \log(\pi(a t |s t ))).

[0050] Step 2: Perform pilot channel estimation according to the analog precoders selected by each RRH to obtain the equivalent estimated channel from each RRH to the corresponding served user;

[0051] In this step, after obtaining the analog precoders selected by each RRH and the served user D mk it is possible to perform pilot channel estimation, where is the set of base stations, is the set of users, and the specific process is as follows:

[0052] Each user sends an orthogonal pilot sequence is the complex domain. Here, pilot contamination between users is not considered, that is, \(\tau\geq K\), where \(K\) is the total number of users and \(\tau\) is the pilot length. Then the received signal \(Y m of the \(m\)-th RRH is expressed as:

[0053]

[0054] where is the noise matrix composed of \(\tau\) noise vectors; \(\rho k is the uplink transmit power of the \(k\)-th user; \((\cdot) H denotes taking the conjugate; \(h mk denotes the channel vector from the \(k\)-th user to the \(m\)-th RRH;

[0055] Multiply the right side of Equation (5) by to solve for the equivalent estimated channel from the \(m\)-th RRH to the \(k\)-th user, which is expressed as:

[0056]

[0057] Step 3: Based on the users served by each RRH and the equivalent estimated channels, obtain the digital precoder and the fronthaul compression noise covariance matrix in the central baseband processing unit BBU through the fractional programming algorithm to complete the design of the hybrid precoder.

[0058] In this step, based on the users D served by each RRHmk and the equivalent estimated channel Obtain the digital precoder F at the central baseband processing unit BBU through the fractional programming algorithm BB and the fronthaul compression noise covariance matrix Ω, where The specific process is as follows:

[0059] After introducing the pilot channel estimation, the received signal y of the k-th user k is expressed as:

[0060]

[0061] where, e mk is the channel estimation error vector from the m-th RRH to the k-th user; represents the digital precoding vector; s i represents the user signal; q m represents the compression error; n k represents the noise; e mk is expressed according to Equation (6) as:

[0062]

[0063] The channel estimation error vector satisfies i.e., a complex Gaussian distribution with variance Ξ mk where the estimation error covariance matrix Ξ mk is expressed as:

[0064]

[0065] is the received noise of the RRH during the pilot channel estimation, and the subscript u represents user; ρ k is the uplink transmission power of the k-th user;

[0066] To simplify the expression, the following vectors and matrices are constructed: the equivalent estimated channel vector of the k-th user the channel estimation error vector of the k-th user the aggregated effective estimated channel vector of the l-th user and the k-th user the aggregated channel estimation error vector of the l-th user and the k-th user the fronthaul compression noise covariance matrix Ω = diag(Ω 1 , …, Ω M );

[0067] Finally, the signal-to-interference-plus-noise ratio γ of the k-th user k is expressed as:

[0068]

[0069] Among them, represents the aggregated effective estimated channel vector; represents digital precoding; ζ k (F BB , Ω) is the sum of interference and noise, expressed as:

[0070]

[0071] In the above formula, tr represents taking the trace; represents the noise power, and are composed of the estimated error covariance matrix, so the overall optimization problem is revised to:

[0072]

[0073] For the above optimization problem (12) related to the weighted sum rate, the first constraint c m (F BB , Ω m ) ≤ C m means that the fronthaul capacity is greater than the communication rate, and the second constraint p m (F BB , Ω m ) ≤ P m represents the power constraint;

[0074] It is solved by using the fractional programming algorithm. First, introduce the auxiliary variable The objective function is transformed quadratically and expressed as f o (F BB , Ω, m), and the expression is:

[0075]

[0076] Among them represents taking the real part; the auxiliary variable m can be iteratively optimized. Let F o in f BB (F BB , Ω, m) and Ω be fixed, and the optimal value m' k of the auxiliary variable m is obtained, expressed as:

[0077]

[0078] Thus, the objective function is transformed into f o (F BB , Ω, m'), where m' is the optimal auxiliary variable, which is convex with respect to the variables F BB and Ω, but the fronthaul capacity constraint is still non-convex. Use the inequality (15) to transform the constraint:

[0079] log2 |A|≤log 2 |B|+tr(B -1 A)-N (15)

[0080] Where N represents the dimension; positive definite Hermitian matrix When A=B, take the equal sign;

[0081] At the same time, another auxiliary variable is introduced Finally, the forward data rate is obtained to satisfy the inequality:

[0082]

[0083] The right side of the inequality g m (F BB ,Ω m ,Σ m ) is bounded by:

[0084]

[0085] When another auxiliary variable Σ m Requirements:

[0086]

[0087] Equation (16) takes the equal sign, and we have:

[0088]

[0089] where c m (F BB ,Ω m ) represents the forward transmission capacity; g m (F BB ,Ω m ,Σ′ m ) Regarding the variable F BB , Ω is a convex constraint; therefore, the final optimization problem (12) is transformed into a convex problem, expressed as:

[0090]

[0091] f o is a weighted and rate-dependent optimization objective. The first constraint g m (F BB ,Ω m ,Σ′ m )≤C m Indicates that the forward transmission capacity is greater than the communication rate. The second constraint p m (F BB ,Ω m )≤P m represents the power constraint;

[0092] With the fixed auxiliary variable m and the auxiliary variable Σ m a new F is obtained by solving the convex problem of formula (20). BB And Ω; Iterate the above process until the growth of the objective function is less than the preset value, and output the finally obtained F BB and Ω to complete the design of the hybrid precoder.

[0093] The following further illustrates the embodiments of the invention with specific examples. In this example, consider a mmWave mMIMO C-RAN system, in which M RRHs are uniformly distributed, and each RRH is a uniform linear array (ULA) with N antennas, where the number of RF chains is L, and K single-antenna users are uniformly randomly distributed.

[0094] Further assume that the fronthaul capacity limits of all RRHs are the same and the maximum transmit power is the same In addition, define the SNR during channel estimation as where it is assumed that the uplink transmit power of all users is ρ.

[0095] As Figure 4 shown is the schematic diagram of the offline deployment process of the method described in the embodiment of the present invention. The offline deployment of the proposed invention is completed through the trained actor-critic network. As Figure 5 shown is the schematic diagram of the online deployment process of the method described in the embodiment of the present invention. According to Figure 4 , Figure 5 shown steps, the design of the hybrid precoder is completed. By online training the actor-critic network to adapt to the changing CSI, the main difference from the offline deployment is that a reward function needs to be fed back to the multi-agent to train the actor-critic network.

[0096] Finally, taking the total user rate as the evaluation criterion, the method described in the embodiment of the present invention has better performance than the "WMMSE" and "Heuristic algorithm" proposed in the prior art under fixed CSI and changing CSI, different channel estimation errors γ, and different fronthaul capacity limits. As Figure 6 shown is the schematic diagram of the comparison of the total user rate varying with SNR in the example of the present invention. As Figure 7 shown is the schematic diagram of the comparison of the total user rate varying with the fronthaul capacity limit in the example of the present invention. Figure 6 and Figure 7 show that under the condition of fixed CSI, the method described in the embodiment of the present invention has performance improvement compared with the prior art under different channel estimation errors and different fronthaul capacity limits.

[0097] AsFigure 8 The figure shows a schematic diagram of the comparison of the total user rate according to the example of the present invention as the channel time slot changes. Figure 8 It is demonstrated that under the condition of changing CSI, the method of the present invention converges after 50 channel time slots. At the same time, compared with the prior art, the proposed invention has improved performance under different channel estimation errors.

[0098] It is worth noting that the contents not described in detail in the embodiments of the present invention belong to the prior art known to professional and technical personnel in the field.

[0099] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the method.

[0100] An embodiment of the present invention further provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method.

[0101] In summary, the method described in the embodiment of the present invention has the following advantages:

[0102] 1. Compared with other existing hybrid precoding technologies, the method of the present invention takes into account the channel estimation and the fronthaul capacity limitation under the C-RAN architecture at the same time; at the same time, compared with the similar convex optimization hybrid precoding technology in the background technology, the present invention solves the problems of constant modulus constraint and large variable space in the simulation domain by maximizing the trade-off between action entropy function and maximizing reward, that is, the trade-off between exploration and maximization of optimization objectives, so that the simulation domain variables can take into account both exploration and optimization;

[0103] 2. By setting the reward function, the analog domain and the digital domain have a certain coupling relationship, guiding the reinforcement learning algorithm to consider the impact of analog domain variables on the digital domain while generating them;

[0104] 3. Compared with similar reinforcement learning hybrid precoding techniques in the background art, the method of the present invention reduces the action space by setting a softmax function and modifying the loss function of the actor-critic network, thereby reducing the algorithm complexity;

[0105] 4. The method of the present invention can guide the intelligent agent to learn the changing trend of the channel through channel statistics, the critic network's evaluation of the random environment, and the adaptive change of the entropy temperature coefficient, and feedback the reward function in each time slot, thereby realizing online deployment.

[0106] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art part of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art.

Claims

1. A hybrid precoding method based on low-dimensional discrete soft actor-critic, characterized in that The method comprises: Step 1: According to the channel state information CSI between each user and each remote radio head RRH, a simulated precoder selected by each RRH and the user served are obtained according to the low-dimensional discrete soft actor critic D-SAC reinforcement learning algorithm; Step 2: perform pilot channel estimation according to the analog precoder selected by each RRH to obtain an equivalent estimated channel from each RRH to the corresponding service user; Step 3: Based on the users selected by each RRH and the equivalent estimated channel, the digital precoder and the fronthaul compression noise covariance matrix are obtained in the central baseband processing unit BBU through the fractional programming algorithm to complete the design of the hybrid precoder.

2. The hybrid precoding method based on low-dimensional discrete soft actor-critic according to claim 1, characterized in that: In step 1, the low-dimensional discrete soft actor-critic D-SAC reinforcement learning algorithm includes a Markov decision process MDP, a low-dimensional D-SAC network structure, and a training process of the low-dimensional D-SAC, wherein: The Markov decision process includes agents, states, actions, rewards, and state transitions. The simulated precoder selected by each RRH and the user served are equivalent to the beam-user pair selected by each agent. Specifically: Agent: Each base station is regarded as an agent, based on the strategy π m Interact with the environment and observe the current state s t,m The beam-user pair a t,m ,in is a set of base stations; State: The system model contains K users and N beam directions. The state of each agent It includes two parts: 1) According to the Toeplitz property of the channel autocorrelation matrix, each RRH calculates the local channel state information CSI statistic R composed of the diagonal elements of the channel autocorrelation matrix of each user m , expressed as: represents the real number domain; 2) the currently accumulated selected beam-user pair If G m The element g at the nth row and kth column in nk =1, indicating that the kth user and the nth beam are selected; Action: Each agent's action is the total number of beam-user pairs it can choose, i.e., action Where B is the number of candidate beams, which is reduced to Then the coupling relationship between beam-user pairs is implicit in the reward function; Reward: The reward for each agent is the same. When the time step t is less than the maximum time step L, the reward is 0; when t = L, there is a total reward r L , expressed as: where ω k represents the weight of user k; R k represents the communication rate of user k; State transfer: Each experience of each agent has L time steps, and each time step selects a different beam-user pair a t,m , used to change the state s at the next moment t+1,m G in m part.

3. The hybrid precoding method based on low-dimensional discrete soft actor critic according to claim 2, characterized in that: In step 1, the low-dimensional D-SAC network structure includes an actor network and a critic network. Each agent has a corresponding actor-critic network. In the mth agent, the actor network inputs the state s at the current time step t t,m Then, two Softmax functions are used to output the B×1 beam selection probability vector and the K×1 user selection probability vector respectively, and the beam direction a of the current time step t is selected according to the probability vector. t0,m and users of the service a t1,m , then the beam-user pair a t,m It is expressed as: a t,m ={a t0,m ,a t1,m }; In the mth agent, the critic network is fed with the input state s at the current time step t t,m and the selected beam-user pair a t,m , get its Q value, that is, the long-term expected benefit. Since the action is discrete, the critic network inputs the state s of the current time step t t,m , output each a t,m The Q value of the final output is a Q value vector; In general, the beam selection probability vector output by the mth actor network at the current time step t is The user selection probability vector is s t is the state vector of the current time step t, represents a one-dimensional vector of length B, represents a one-dimensional vector of length K; the Q value vector output by the m-th critic network at the current time step t is Among them, the beam selection Q value vector is The user chooses the Q value vector as θ i ,i=1,2 represents a double-depth Q network.

4. The hybrid precoding method based on low-dimensional discrete soft actor critic according to claim 3, characterized in that: In step 1, the training process of low-dimensional D-SAC is as follows: The agent continuously interacts with the environment through the actor network, generating multiple (s t ,a t ,r t ,s t+1 ) samples are stored in the playback buffer; t Instead of t,m Indicates the state, r t Instead of r t,m represents reward, m represents agent, and the other parameters are similar; Then, small batches of samples are taken from the playback buffer to train the actor network, the critic network and the entropy temperature coefficient, which is used to balance the proportion of network exploration and utilization; The loss function J of the actor network π (φ) is expressed as: in represents expectation; φ represents the trainable parameter of the actor network; α is the entropy temperature coefficient; D represents the replay memory of all previous state-action-reward-next-state experiments; The loss function J of the critic network Q (θ i ) is expressed as: in, Represents two target Q networks, which are composed of the original double Q network θ i ,i=1,2 is updated regularly; Q θi represents the Q network function; r represents the reward; η represents the discount factor; Indicates the Q value corresponding to state s and action a; The loss function L(α) of the entropy temperature coefficient is expressed as: in is the target entropy; H(π(·|s t ))=log(π(a t |s t )).

5. The hybrid precoding method based on low-dimensional discrete soft actor-critic according to claim 1, characterized in that: In step 2, the simulated precoder selected by each RRH is obtained. and the user D of the service mk After that, the pilot channel estimation can be performed, where is a collection of base stations, is a collection of users. The specific process is: Each user sends an orthogonal pilot sequence is a complex domain. Here, pilot contamination between users is not considered, that is, τ ≥ K, K is the total number of users, τ is the pilot length, then the received signal Y of the mth RRH is m It is expressed as: in, is the noise matrix composed of τ noise vectors; ρ k is the uplink transmission power of the kth user; (·) H Indicates conjugation; h mk represents the channel vector between the kth user and the mth RRH; Multiply the right side of formula (5) by The equivalent estimated channel from the mth RRH to the kth user can be solved It is expressed as:

6. The hybrid precoding method based on low-dimensional discrete soft actor-critic according to claim 5, characterized in that: In step 3, the served user D is selected on a per RRH basis. mk and the equivalent estimated channel The digital precoder F is obtained in the central baseband processing unit BBU through the fractional programming algorithm. BB and the forward compression noise covariance matrix Ω, where The specific process is: After the pilot channel estimation is introduced, the received signal y of the kth user k It is expressed as: Among them, e mk is the channel estimation error vector from the mth RRH to the kth user; represents the digital precoding vector; s i represents the user signal; q m Indicates compression error; n k represents noise; e mk According to formula (6), it can be expressed as: The channel estimation error vector satisfies That is, the variance is Ξ mk The complex Gaussian distribution of the estimation error covariance matrix Ξ mk It is expressed as: is the received noise of the RRH during the pilot channel estimation, subscript u represents user; k is the uplink transmission power of the kth user; To simplify the expression, the following vectors and matrices are constructed: The equivalent estimated channel vector of the kth user is The channel estimation error vector of the kth user is The aggregated effective estimated channel vector of the lth user and the kth user is The aggregate channel estimation error vector of the lth user and the kth user is The forward compression noise covariance matrix Ω=diag(Ω1,…,Ω M ); Finally, the signal-to-interference-noise ratio γ of the kth user k It is expressed as: in, represents the aggregated effective estimated channel vector; represents digital precoding; k (F BB ,Ω) is the sum of interference and noise, expressed as: In the above formula, tr means finding the trace; represents the noise power, and It is composed of the estimated error covariance matrix, so the overall optimization problem is corrected as follows: For the above weighted and rate-dependent optimization problem (12), the first constraint c m (F BB ,Ω m )≤C m Indicates that the forward transmission capacity is greater than the communication rate. The second constraint p m (F BB ,Ω m )≤P m represents the power constraint; The fractional programming algorithm is used to solve the problem. First, auxiliary variables are introduced. The objective function is transformed twice and expressed as f o (F BB ,Ω,m), the expression is: The auxiliary variable m can be optimized iteratively, so that f o (F BB ,Ω,m) BB , Ω is fixed, and the optimal value m′ of the auxiliary variable m is obtained k , expressed as: Thus, the objective function is transformed into f o (F BB ,Ω,m′), m′ is the optimal auxiliary variable, relative to the variable F BB , Ω is convex, but the forward capacity constraint is still non-convex. The constraint is transformed using inequality (15): log2|A|≤log2|B|+tr(B -1 A)-N (15) Where N represents the dimension; positive definite Hermitian matrix When A=B, take the equal sign; At the same time, another auxiliary variable is introduced Finally, the forward data rate is obtained to satisfy the inequality: The right side of the inequality g m (F BB ,Ω m ,Σ m ) is bounded by: When another auxiliary variable Σ m Requirements: Equation (16) takes the equal sign, and we have: where c m (F BB ,Ω m ) represents the forward transmission capacity; g m (F BB ,Ω m ,Σ′ m ) Regarding the variable F BB , Ω is a convex constraint; therefore, the final optimization problem (12) is transformed into a convex problem, expressed as: f o is a weighted and rate-dependent optimization objective. The first constraint g m (F BB ,Ω m ,Σ′ m )≤C m Indicates that the forward transmission capacity is greater than the communication rate. The second constraint p m (F BB ,Ω m )≤P m represents the power constraint; With fixed auxiliary variables m and auxiliary variables Σ m In the case of , the new F is obtained by solving the convex problem of formula (20) BB and Ω; iterate the above process until the growth of the objective function is less than the preset value, and output the final F BB and Ω, completing the design of the hybrid precoder.

7. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 6.

8. A computer storage medium, characterized in that: The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1 to 6.