A method and system for massive MIMO user scheduling and beamforming based on MAB
By adopting a two-stage beamforming scheme based on multi-arm gambling machines in the frequency division duplex large-scale MIMO system, the problems of downlink pilot length and channel feedback overhead are solved, and efficient spectrum utilization is achieved.
Patent Information
- Application Number
- CN202411677866.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-11-22
AI Technical Summary
In Frequency Division Duplex (FDD) large-scale MIMO systems, the prior art is difficult to effectively reduce downlink pilot length and channel feedback overhead, especially in multi-user scenarios, resulting in low spectrum efficiency.
A two-stage beamforming (TSB) scheme based on multi-arm gambling machine (MAB) is adopted to convert the design problem of maximizing spectrum efficiency into a combined multi-arm gambling machine problem in two-dimensional space, and solve it using the optimal linear UCB algorithm and the suboptimal linear UCB algorithm to obtain the beamforming matrix, avoiding the estimation of the channel covariance matrix (CCM).
It effectively reduces the downlink training length and channel feedback overhead, improves spectrum efficiency, and significantly improves the spectrum utilization of the system in multiple user scenarios.
Smart Images

Figure CN119182437B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a MAB-based large-scale MIMO user scheduling and beamforming method and system, and belongs to the technical field of wireless communications. Background Art
[0002] The difficulty of obtaining channel state information is crucial for the study of precoding algorithms for large-scale MIMO systems. For the two duplex modes in the current communication field, namely frequency division duplex (FDD) and time division duplex (TDD), their channel state information acquisition methods are significantly different. For the frequency division duplex (FDD) mode, due to the asymmetry of the uplink and downlink channels, the downlink also requires the base station to send a pilot sequence. The terminal devices calculate the state information of their respective channels and feed it back to the base station. The amount of pilot overhead will be proportional to the number of antenna arrays. Obviously, for large-scale MIMO systems, the large number of antennas causes a huge amount of pilot overhead, and pilot pollution also surges. Although dynamic antenna switching technology can reduce the number of pilot links, this technology does not fully utilize the advantages of multi-antenna gain, and there is a signal delay problem on the reverse link, so it is difficult to obtain all the channel information of the system in a timely manner.
[0003] In order to fully utilize the advantages of massive MIMO, it is particularly important to obtain channel state information at the base station (BS). However, channel reciprocity does not hold in frequency division duplex (FDD). In order to obtain instantaneous CSI, a large downlink pilot length (DTL) has to be obtained, which bears a large amount of CSI feedback overhead. In future wireless communications, both TDD and FDD modes are very important, so it is necessary to reduce the downlink pilot length (DTL) and channel feedback in frequency division duplex (FDD) mode. At the same time, when the number of users is greater than what the system can serve, improving spectrum efficiency will also be a rigid demand of the system. In frequency division duplex (FDD) massive MIMO, two-stage beamforming (TSB) using the channel covariance matrix (CCM) can reduce the downlink training length (DTL) and channel feedback. However, the cost of estimating the channel covariance matrix (CCM) is also very high. Therefore, it is necessary to propose a TSB algorithm that does not require the channel covariance matrix (CCM), which can improve the performance of TSB while taking into account the number of users and reducing overhead.
[0004] Adhikary used the spatial correlation and channel sparsity of massive MIMO to study a two-stage precoding scheme based on joint space division multiplexing (JSDM) technology. The key idea is to allocate user devices with different spatial characteristics to different groups through user scheduling, and divide the downlink precoding process into two matrix solution processes. It lays the foundation for the precoding implementation of massive MIMO technology under the frequency division duplex (FDD) standard. When the number of users is greater than the system can serve, some studies have proposed user scheduling schemes to improve spectrum efficiency. These works rely on the CCM available at the base station, but the estimation of the channel covariance matrix (CCM) remains a challenging problem. In addition, in scenarios such as fast user movement, the channel covariance matrix (CCM) changes rapidly, which makes the estimation of CCM more difficult.
[0005] Considering the high cost of estimating CCM, Song designed a TSB scheme based on the multi-armed bandit (MAB), which avoids the estimation of the channel covariance matrix (CCM) and improves the spectrum efficiency. The TSB scheme based on the multi-armed bandit (MAB) transforms the pre-beamforming matrix design problem into a MAB problem, and uses the algorithm to train the pre-beamforming matrix, eliminating the need for the channel covariance matrix (CCM). Summary of the invention
[0006] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a method and system for massive MIMO user scheduling and beamforming based on MAB to solve the problems in the prior art that the multi-user scenario cannot be handled, the downlink training length is too long, and the effective spectrum efficiency is low;
[0007] In order to achieve the above purpose / solve the above technical problems, the present invention is implemented by adopting the following technical solutions:
[0008] A first aspect: A method for massive MIMO user scheduling and beamforming based on MAB, the method comprising:
[0009] Establish the TSB scheme communication model and obtain the design problem of maximizing spectrum efficiency;
[0010] The obtained design problem of maximizing spectrum efficiency is converted into a combinatorial multi-armed bandit problem in two-dimensional space;
[0011] The objective function of the two-dimensional combinatorial multi-armed bandit problem is further divided into two sub-problems. In the first sub-problem, the bandit arms are selected to find the zero value in the power angular spectrum, and in the second sub-problem, the bandit arms are selected to maximize the received energy.
[0012] According to the two sub-problems, the optimal linear UCB algorithm and the suboptimal linear UCB algorithm are used to solve the combined multi-armed bandit problem in two-dimensional space and obtain the beamforming matrix.
[0013] Optionally, the establishing of the TSB scheme communication model to obtain a design problem for maximizing spectrum efficiency includes:
[0014] Using DFT matrix To design the pre-beamforming matrix, , Indicates the unprocessed signal at the sender. Indicated in The selected users in the time window are Indicated in The selected DFT vector in the time window, then the pre-beamforming matrix is expressed as , in The effective channel matrix in the time window is expressed as ;
[0015] set up for The base number, represents a complex set, L is the number of base station array elements, then when the base station sends a signal, the pilot matrix is At that time, The signal received by each user is:
[0016] ;
[0017] in, represents the conjugate transposed matrix, Indicates Users in The signal received at any time, and For the The channel vector of the user Noise in a time window;
[0018] In the applicable channel, the problem is transformed into designing a and , so that the spectrum efficiency is maximized, according to the derivation of Shannon's formula, the energy of the received signal When the maximum value is reached, the spectrum efficiency index will be improved accordingly, because , the objective function of the design problem of the pre-beamforming matrix is:
[0019] ;
[0020] in, Indicated in selected users in a time window, Indicated in The selected DFT vector in the time window, The set of indices representing the selected user, represents the set of indices of the selected DFT vector, For the User in The channel vector in the time window is The index is The DFT vector, The length of downlink training The limiting constant.
[0021] Optionally, converting the obtained design problem of maximizing the spectrum efficiency into a combined multi-armed bandit problem in two-dimensional space includes:
[0022] The two-dimensional combinatorial multi-armed bandit problem has two sets of variables and , and represents the number of elements in each of these two variable sets, is the number of users, is the number of DFT vectors, the combination of each pair of elements of the two variable sets All have an unknown mean, and each mean satisfies independent and identical distribution;
[0023] The problem It is applicable to the problem of a combined multi-armed bandit in a two-dimensional space corresponding to two sets of variables. The set of indices representing the selected user, , represents the set of indices of the selected DFT vector, , the super arm of the multi-armed bandit is The corresponding reward is .
[0024] Optionally, converting the obtained design problem of maximizing the spectrum efficiency into a combined multi-armed bandit problem in two-dimensional space includes:
[0025] The two-dimensional combinatorial multi-armed bandit problem has two sets of variables and , and represents the number of elements in each of these two variable sets, is the number of users, is the number of DFT vectors, the combination of each pair of elements of the two variable sets All have an unknown mean, and each mean satisfies independent and identical distribution;
[0026] The problem It is applicable to the problem of a combined multi-armed bandit in a two-dimensional space corresponding to two sets of variables. The set of indices representing the selected user, , represents the set of indices of the selected DFT vector, , the super arm of the multi-armed bandit is The corresponding reward is .
[0027] Optionally, the objective function of the two-dimensional combined multi-armed bandit problem is further divided into two sub-problems, including:
[0028] The first sub-problem is to standardize the likelihood ratio in the interval from 0 to 1 using a standardized utility function, and the standardized utility function is as follows:
[0029] ;
[0030] in, Indicates The channel vector of the user is The likelihood ratio of the signal energy after the operation of the DFT vector, Then The standardized form of the likelihood ratio The bigger, If the signal is transmitted with the same probability, then Its mapping non-zero;
[0031] The first sub-problem is transformed into how to choose and Can make The sum is the largest, and the formula of the first sub-problem is expressed as follows:
[0032] ;
[0033] ;
[0034] ;
[0035] in, Indicates The channel vector of the user is The normalized likelihood ratio of the signal energy after the operation of the DFT vector, The set of indices representing the selected user, represents the set of indices of the selected DFT vector, is the limiting constant of the downlink training length, In this subproblem, selected users in a time window, In this subproblem, The selected DFT vector in the time window;
[0036] because The likelihood ratio of the group is zero in the remaining set, There are multiple solutions, in each solution In, there is Make have or Make have , it is necessary to remove rows and columns with all zero elements from the solution so that the sum of the likelihood ratios is the maximum value;
[0037] When the solution , select the remaining arms to maximize the received energy, then the second sub-problem is formulated as follows:
[0038] ;
[0039] ;
[0040] ;
[0041] in, Indicated in selected users in a time window, Indicated in The selected DFT vector in the time window, The set of indices representing the selected user, represents the set of indices of the selected DFT vector, For the User in The channel vector in the time window is The index is The DFT vector, is the limiting constant of the downlink training length, Indicated in In the sub-question selected users in a time window, Indicated in In the sub-problem The selected DFT vector in the time window.
[0042] Optionally, according to the two sub-problems, an optimal linear UCB algorithm or a suboptimal linear UCB algorithm is used to solve the combined multi-armed bandit problem in two-dimensional space to obtain a beamforming matrix, including: based on the application scenario of a single-ring channel, an optimal linear UCB algorithm and a suboptimal linear UCB algorithm are used to solve the combined multi-armed bandit problem in two-dimensional space to obtain a beamforming matrix.
[0043] Optionally, based on the application scenario of the single-loop channel, the optimal linear UCB algorithm is used to solve the combined multi-armed bandit problem in two-dimensional space to obtain a beamforming matrix, including:
[0044] In the first subproblem, the utility function As indexed by the DFT vector With user index Gambling Arm The reward of a basic gambling arm in the optimal linear algorithm The UCB value is defined as follows:
[0045] ;
[0046] in, Represents the gambling arm in the optimal algorithm The UCB value, is the number of users, is the number of DFT vectors, Indicates The number of times this gambling arm has been selected in iterations, is the undetermined gambling arm set corresponding to the effective channel vector in the user channel, For the current gambling arm The modulus of the reward value obtained in the iteration; in order to select the super arm that minimizes the UCB value, a vector To indicate whether the user has been selected, use Indicates whether a beam has been selected. Representative beams have been selected, and 0 is the opposite. The problem of selecting the super arm of the slot machine will be expressed as follows:
[0047] ;
[0048] ;
[0049] in, represents the transpose operation, is a 0-1 vector indicating whether each user has been selected. To indicate the A logical value indicating whether the user has been selected. A 0-1 vector indicating whether each DFT vector has been selected. For all channels, For the The channel vector of each user.
[0050] Optionally, based on the application of a single-loop channel, a suboptimal linear UCB algorithm is used to solve the combined multi-armed bandit problem in two-dimensional space to obtain a beamforming matrix, including:
[0051] Consider the first sub-problem, assuming that In the iteration, and is the index set of the selected user and beam, and takes action in the next iteration The rewards will be defined as follows:
[0052] ;
[0053] in, For the The gambling arm Rewards, , is the union of the user set and the DFT vector set in the current gambling arm and the i-1th iteration, To indicate the The channel vector of the user is The normalized likelihood ratio of the signal energy after the operation of the DFT vector;
[0054] because , easy to prove Therefore, in the suboptimal algorithm, In each gamble arm The UCB value is defined as follows:
[0055] ;
[0056] in, Represents the gambling arm in the first sub-problem of the suboptimal algorithm The UCB value, For the The gambling arm The upper limit of the reward value, Indicates The number of times this gambling arm has been selected in iterations, For the current gambling arm The modulus of the reward value obtained in the iteration; maximize the UCB value in the The original problem in this iteration will be listed as follows:
[0057] ;
[0058] ;
[0059] To solve the problem, first calculate the UCB value of each unselected gambling arm , such as In , and select the maximum value from them;
[0060] When considering the second sub-problem, let In the iteration, and The index set of the selected user and vector, which will be acted on in the next iteration The rewards will be defined as follows:
[0061] ;
[0062] in, For the The gambling arm Rewards, , is the union of the user set and the DFT vector set in the current gambling arm and the i-1th iteration, For the User in The channel vector in the time window is The index is DFT vector of
[0063] award It will present a sub-exponential distribution, so the UCB value of each gambling arm will be defined as follows;
[0064] ;
[0065] in, Represents the gambling arm in the second sub-problem of the suboptimal algorithm The UCB value, is the number of users, For the gambling arm, Indicates The number of times this gambling arm has been selected in iterations, For the current gambling arm The modulus of the reward values obtained in the iteration, is the total reward that the current system has obtained in the i-th iteration; The original problem in this iteration will be listed as follows:
[0066] ;
[0067] ;
[0068] Calculate the UCB value of each unselected gambling arm , and select the maximum value from them.
[0069] Optionally, the method further includes: based on a multi-scattering cluster channel, converting the calculation of the downlink training length in the constraint condition into a graph coloring problem, using a suboptimal linear UCB algorithm to solve the color number of the graph, and obtaining a beamforming matrix, wherein converting the calculation of the downlink training length in the constraint condition into a graph coloring problem includes:
[0070] Use the color number of the graph to reduce the computational complexity and satisfy The minimum DTL value is ,in, For users The effective channel matrix is represents transpose, is a 0-1 matrix representing the relationship between the DFT vector and the user pairing, express The chromatic number of the graph of the adjacency matrix is for A solution of , the pilot matrix is designed as ,in, It is arbitrary Unitary matrix of size No. List.
[0071] Optionally, based on a multi-scattering cluster channel, the downlink training length is converted into the color number of a graph, and the color number of the graph is solved using a suboptimal linear UCB algorithm to obtain a beamforming matrix, including:
[0072] Suppose that in the suboptimal linear algorithm In the iteration, and is the index set of the selected users and vectors, The reward is , in Actions in iterations The rewards will be defined as follows:
[0073] ;
[0074] in, Represents the gambling arm in the first subproblem of the MSC channel The UCB value, For the The gambling arm The upper limit of the reward value, Indicates The number of times this gambling arm has been selected in iterations, For the current gambling arm The modulus of the reward values obtained in iterations;
[0075] The problem of maximizing the UCB value of the selected action is The iterations will be listed as follows:
[0076] ;
[0077] ;
[0078] To solve the first sub-problem, first calculate the UCB value of each unselected gambling arm , and then the largest gambling arm of UCB was selected and satisfied , and then use the backtracking algorithm to calculate the color number;
[0079] Solve the second sub-problem, reward It will present a sub-exponential distribution, and the UCB value of each gambling arm will be defined as follows:
[0080] ;
[0081] in, Represents the gambling arm in the second sub-problem of the MSC channel The UCB value, is the number of users, For the gambling arm, Indicates The number of times this gambling arm has been selected in iterations, For the current gambling arm The modulus of the reward values obtained in the iteration, is the total reward that the current system has obtained in the i-th iteration; The original problem in this iteration will be listed as follows:
[0082] ;
[0083] ;
[0084] Calculate under constraints The UCB value of each unselected gambling arm , the number of colors in the constraint will be obtained by the backtracking algorithm, and the gambling arm with the largest UCB value will be selected from it.
[0085] A second aspect: A massive MIMO user scheduling and beamforming system based on MAB, the system comprising:
[0086] Modeling module, used to establish the communication model of TSB scheme and obtain the design problem of maximizing spectrum efficiency;
[0087] A conversion module, used for converting the obtained design problem of maximizing spectrum efficiency into a combinatorial multi-armed bandit problem in two-dimensional space;
[0088] A problem decomposition module is used to further divide the objective function of the two-dimensional combined multi-armed bandit problem into two sub-problems, in which the gambling arms are selected to find the zero value in the power angular spectrum in the first sub-problem, and in the second sub-problem, the gambling arms are selected to maximize the received energy;
[0089] The calculation module is used to solve the combined multi-armed bandit problem in two-dimensional space according to two sub-problems by using the optimal linear UCB algorithm and the suboptimal linear UCB algorithm to obtain a beamforming matrix.
[0090] Compared with the prior art, the present invention has the following beneficial effects:
[0091] The present invention defines a combined multi-armed bandit for a two-dimensional problem, which is used to describe the pre-beamforming matrix and user scheduling design problem in a two-stage beamforming scheme (TSB), and further divides the problem into two sub-problems. First, a bandit arm is selected to find the zero value in the power angular spectrum, and then a bandit arm is selected to maximize the received energy, thereby improving the spectrum efficiency.
[0092] Aiming at the application scenarios of single-ring channels, the present invention proposes an optimal linear UCB algorithm and a suboptimal UCB algorithm to solve the MAB problem; in the multi-scattering cluster (MSC) channel, the present invention proves that the downlink training length (DTL) is equal to the color number of the graph, and proposes a suboptimal linear UCB algorithm with low computational complexity, which improves the effective spectrum efficiency and reduces the feedback length and DTL advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] Figure 1 FIG. 4 is a flow chart of a method for massive MIMO user scheduling and beamforming based on MAB according to the present invention;
[0094] Figure 2 It is a schematic diagram of ESE simulation in a single-ring channel under different numbers of users of different schemes according to an embodiment of the present invention;
[0095] Figure 3 It is a schematic diagram of ESE simulation in MSC channel under different numbers of users in different schemes according to an embodiment of the present invention;
[0096] Figure 4It is a schematic diagram of ESE simulation of different schemes in a single-ring channel under different signal-to-noise ratios according to an embodiment of the present invention;
[0097] Figure 5 It is a schematic diagram of ESE simulation in MSC channel under different signal-to-noise ratios of different schemes according to an embodiment of the present invention;
[0098] Figure 6 Shown is a schematic diagram of a MAB-based massive MIMO user scheduling and beamforming system of the present invention. DETAILED DESCRIPTION
[0099] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below in conjunction with specific implementation methods.
[0100] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and the like are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, features defined as "first", "second", and the like may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.
[0101] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood by specific circumstances.
[0102] Example 1, reference Figure 1-Figure 6 As shown, the present invention considers a single-cell FDD massive MIMO system. The base station is equipped with a The TSB scheme uses a uniform linear array (ULA) with 4 elements. Since the average CSI remains unchanged over a much larger interval than the time slot, it is easier to obtain than the instantaneous CSI. The TSB scheme uses the average CSI to design a pre-beamforming matrix to spatialize the effective channel matrix, thereby reducing the pilot overhead. In the first stage of the TSB scheme, the pre-beamforming matrix is designed. To sparse the effective channel matrix Where , and From the base station to the The base station then sends a pilot signal to the user, and the user terminal performs In the second stage of the scheme, the base station designs a multi-user precoder , to alleviate interference between users and thus improve spectrum utilization.
[0103] use Represents the received signal, where It is The received signal can then be expressed as:
[0104] ;
[0105] in, is the transmitted signal, with Characteristics. is noise, and . Multi-user precoding matrix In the second phase, it is designed To reduce interference, The present invention adopts a channel model based on scatterer clusters, and signals are sent between users through one or more scatterer clusters. Indicates The number of scatterer clusters between users and base stations is evenly distributed in superior, is the maximum number of clusters for all users. User's The angular range of the scatterer clusters is in the interval On, the user CCM for: ;
[0106] in represents conjugation, represents the steering vector, represents the wavelength, Indicates the distance between two antenna elements. is the channel power angular spectrum (PAS), which is modeled as , is a random variable normalized to be uniformly distributed on Using the Karhunen-Loeve formulation, the user The channel vector can be expressed as ,in It is a small-scale fading.
[0107] The following is an introduction to the MAB-based massive MIMO user scheduling and beamforming method proposed in the present invention:
[0108] like Figure 1 As shown, a method for massive MIMO user scheduling and beamforming based on MAB is disclosed, and the method includes:
[0109] Step 1: Establish a TSB scheme communication model to obtain the design problem of maximizing spectrum efficiency;
[0110] Step 2, converting the obtained design problem of maximizing spectrum efficiency into a combinatorial multi-armed bandit problem in two-dimensional space;
[0111] Step 3, further divide the objective function of the two-dimensional combined multi-armed bandit problem into two sub-problems, in the first sub-problem, select the gambling arm to find the zero value in the power angular spectrum, and in the second sub-problem, select the gambling arm to maximize the received energy;
[0112] Step 4: According to the two sub-problems, the optimal linear UCB algorithm and the suboptimal linear UCB algorithm are used to solve the combined multi-armed bandit problem in two-dimensional space to obtain the beamforming matrix.
[0113] In the specific implementation process, step S1 is specifically described, and the TSB scheme communication model is established to obtain the design problem of maximizing spectrum efficiency, including:
[0114] Using DFT matrix To design the pre-beamforming matrix, ,use Indicated in The selected users in the time window are Indicated in The selected DFT vector in the time window, then the pre-beamforming matrix is expressed as , in The effective channel matrix in the time window is expressed as ;
[0115] When the base station sends the pilot matrix , n is The base number, The signal received by each user is:
[0116] ;
[0117] in, represents the conjugate transposed matrix, Indicates Users in The signal received at any time, and For the The channel vector of the user Noise in a time window;
[0118] In applicable channels, Some elements in will be zero, so Represents the index set of non-zero entries, the pilot matrix for any index Should satisfy , the problem then becomes to design a and , so that the spectrum efficiency is maximized, according to the derivation of Shannon's formula, the energy of the received signal When the maximum value is reached, the spectrum efficiency index will be improved accordingly, because , the objective function of the design problem of the pre-beamforming matrix is:
[0119] ;
[0120] in, Indicated in selected users in a time window, Indicated in The selected DFT vector in the time window, For the user set, is the index set of the DFT vector to be selected, For the User in The channel vector in the time window is The index is The DFT vector, is the limiting constant of the downlink training length.
[0121] In the specific implementation process, for step 2, the design problem of maximizing the spectrum efficiency is converted into a two-dimensional combined multi-armed bandit problem, including:
[0122] The two-dimensional combinatorial multi-armed bandit problem has two sets of variables and , the combination of each pair of elements of the two variable sets All have an unknown mean, and each mean satisfies independent and identical distribution;
[0123] The problem Applicable as variable set and The corresponding two-dimensional space combined multi-armed bandit problem, , represents the index set of the selected user, , represents the index set of the selected DFT vector, and the super arm is The corresponding reward is .
[0124] In the specific implementation process, step 3 is specifically described as follows. The objective function of the two-dimensional combined multi-armed bandit problem is further divided into two sub-problems, including:
[0125] The first sub-question is, assuming that when the energy Less than a constant value can be approximated to zero when , otherwise the value is non-zero. Then the energy Whether it is zero in a certain time window can be calculated by the likelihood ratio method. After n selections, the energy The likelihood ratio will be calculated according to the interval integral ratio The formal definition of , the likelihood ratio Represents energy The probability of being zero. The likelihood ratio is standardized in the interval from 0 to 1 using a standardized utility function, and the standardized utility function is as follows:
[0126] ;
[0127] in, Indicates The channel vector of the user is The likelihood ratio of the signal energy after the operation of the DFT vector, Then The standardized form of the likelihood ratio The bigger, If the signal is transmitted with the same probability, then Its mapping non-zero;
[0128] The first sub-problem is transformed into how to choose and Can make The sum is the largest, and the formula of the first sub-problem is expressed as follows:
[0129] ;
[0130] ;
[0131] ;
[0132] in, Indicates The channel vector of the user is The normalized likelihood ratio of the signal energy after the operation of the DFT vector, For the user set, is the index set of the DFT vector to be selected, is the limiting constant of the downlink training length, In this subproblem, selected users in a time window, In this subproblem, The selected DFT vector in the time window;
[0133] because The likelihood ratio of the group is zero in the remaining set, There are multiple solutions, in each solution In, there is Make have or Make have , it is necessary to remove rows and columns with all zero elements from the solution so that the sum of the likelihood ratios is the maximum value;
[0134] When the solution , select the remaining arms to maximize the received energy, then the second sub-problem is formulated as follows:
[0135] ;
[0136] ;
[0137] ;
[0138] in, Indicated in selected users in a time window, Indicated in The selected DFT vector in the time window, For the user set, is the index set of the DFT vector to be selected, For the User in The channel vector in the time window is The index is The DFT vector, is the limiting constant of the downlink training length, Indicated in In the sub-problem selected users in a time window, Indicated in In the sub-problem The selected DFT vector in the time window.
[0139] In the specific implementation process, step 3 is specifically explained. According to the two sub-problems, the optimal linear UCB algorithm or the suboptimal linear UCB algorithm is used to solve the combined multi-armed bandit problem in two-dimensional space to obtain a beamforming matrix, including: based on the application scenario of a single-ring channel, the optimal linear UCB algorithm and the suboptimal linear UCB algorithm are used to solve the combined multi-armed bandit problem in two-dimensional space to obtain a beamforming matrix.
[0140] In the application scenario based on the single-loop channel, the optimal linear UCB algorithm is used to solve the combined multi-armed bandit problem in two-dimensional space to obtain the beamforming matrix, including:
[0141] The problem is solved in a single-ring channel, where the base station transmits the signal through a single scatterer cluster, so the user The channel covariance matrix can be simplified to some extent. Let , then Not When the final result is fixed to zero, it is easy to know is a set of consecutive integers, let time window For location users A collection of unknown energies, energy and The corresponding edge vector (such as and ) has a greater probability of being zero. So the pre-selected set will be converted into ,in , if for a given User channels all are all certain, then is an empty set. Therefore, when only considering the effective channel vector, let for In order to satisfy the elements , pilot matrix Can be designed as
[0142] of the form, in which ,and is a unitary matrix. Then the DTL at this time will be equal to the maximum value.
[0143] In the first subproblem, the utility function As a gambling arm The reward of the optimal linear algorithm is the basis arm The UCB value (the probability of the gambling arm being selected) is defined as follows:
[0144] ;
[0145] in, is the number of users, is the number of DFT vectors, For the gambling arm, Indicates The number of times this gambling arm has been selected in iterations, is the undetermined gambling arm set corresponding to the effective channel vector in the user channel, For the current gambling arm The modulus of the reward value obtained in the iteration; in order to select the super arm that minimizes the UCB value, a vector To indicate whether the user has been selected, use Indicates whether a beam has been selected. Representative beams have been selected, and 0 is the opposite. The problem of selecting the super arm of the slot machine will be expressed as follows:
[0146] ;
[0147] ;in, is a 0-1 vector indicating whether each user has been selected. To indicate the A logical value indicating whether the user has been selected. A 0-1 vector indicating whether each DFT vector has been selected. For all channels, For the The channel vector of each user. In addition, due to the energy Following the chi-square distribution, the calculation of the UCB value can be simplified by eliminating the infinite case in the subsequent steps.
[0148] use Indicates whether the unselected user is selected. Indicates user ( ) is selected. Similarly, Indicates whether the unselected DFT vector is selected. Denotes the DFT vector ( ) is selected. Next, define a matrix A and add and . and Respectively represent User and DFT vectors, we have In order to find the maximum UCB value, the problem becomes a quadratic integer programming problem in the interval from 0 to 1. According to the deformation of the existing mathematical theorem, this quadratic integer programming problem can be transformed into a linear integer programming problem in the interval from 0 to 1 for final solution. In this O-LUCB algorithm, the action converges to the optimal action, and the regret is sublinear with time.
[0149] Based on the application of single-loop channel, the suboptimal linear UCB algorithm is used to solve the combined multi-armed bandit problem in two-dimensional space, and the beamforming matrix is obtained, including:
[0150] Consider the first sub-problem, assuming that In the iteration, and is the index set of the selected user and beam, and takes action in the next iteration The rewards will be defined as follows:
[0151] ;
[0152] in, For the The gambling arm Rewards, , is the union of the user set and the DFT vector set in the current gambling arm and the i-1th iteration, To indicate the The channel vector of the user is The normalized likelihood ratio of the signal energy after the operation of the DFT vector;
[0153] because , easy to prove Therefore, in the suboptimal algorithm, In each gamble arm The UCB value is defined as follows:
[0154] ;
[0155] in, For the The gambling arm The upper limit of the reward value, Indicates The number of times this gambling arm has been selected in iterations, For the current gambling arm The modulus of the reward value obtained in the iteration; maximize the UCB value in the The original problem in this iteration will be listed as follows:
[0156] ;
[0157] ;
[0158] To solve the problem, first calculate the UCB value of each unselected gambling arm , as above In , and select the maximum value from them;
[0159] When considering the second sub-problem, let In the iteration, and The index set of the selected user and vector, which will be acted on in the next iteration The rewards will be defined as follows:
[0160] ;
[0161] in, For the The gambling arm Rewards, , is the union of the user set and the DFT vector set in the current gambling arm and the i-1th iteration, For the User in The channel vector in the time window is The index is DFT vector of
[0162] award It will present a sub-exponential distribution, so the UCB value of each gambling arm will be defined as follows;
[0163] ;
[0164] in, is the number of users, For the gambling arm, Indicates The number of times this gambling arm has been selected in iterations, For the current gambling arm The modulus of the reward values obtained in the iteration, is the total reward that the current system has obtained in the i-th iteration; The original problem in the iteration will be listed as follows;
[0165] ;
[0166] ;
[0167] Calculate the UCB value of each unselected gambling arm , and select the maximum value from them.
[0168] Based on the multi-scattering cluster channel, the calculation of the downlink training length in the constraint condition is converted into a graph coloring problem. The suboptimal linear UCB algorithm is used to solve the color number of the graph to obtain the beamforming matrix. The calculation of the downlink training length in the constraint condition is converted into a graph coloring problem, including:
[0169] Convert the downlink training length to the color number of the graph, including:
[0170] In the MSC scenario, consider the MSC channel where the base station transmits signals to the user through a multi-scatterer cluster. The CCM matrix of has been given above. Assume , As the number of antennas approaches infinity, the index of the DFT vector does not belong to The final result of the system will also be zero. As a matter of course, the user channel There will be some randomly distributed zero elements in . The pilot matrix defines an index matrix , the matrix Row and The elements of the column can be represented as the DFT vector and the logical value of whether the user is matched. At this time, the color number of the graph can be used to reduce the computational complexity and satisfy The minimum DTL value is ,in, For users The effective channel matrix is represents transpose, is a 0-1 matrix representing the relationship between the DFT vector and the user pairing, express The chromatic number of the graph of the adjacency matrix is for A solution of , the pilot matrix is designed as ,in, It is arbitrary Unitary matrix of size No. List.
[0171] Based on the multi-scattering cluster channel, the downlink training length is converted into the color number of the graph, and the suboptimal linear UCB algorithm is used to solve the color number of the graph to obtain the beamforming matrix, including:
[0172] Suppose that in the suboptimal linear algorithm In the iteration, and is the index set of the selected users and vectors, The reward is , in Actions in iterations The rewards will be defined as follows:
[0173] ;
[0174] in, For the The gambling arm The upper limit of the reward value, Indicates The number of times this gambling arm has been selected in iterations, For the current gambling arm The modulus of the reward values obtained in iterations;
[0175] The problem of maximizing the UCB value of the selected action is The iterations will be listed as follows:
[0176] ;
[0177] ;
[0178] To solve the first sub-problem, first calculate the UCB value of each unselected gambling arm , and then the largest gambling arm of UCB was selected and satisfied , and then use the backtracking algorithm to calculate the color number;
[0179] Solve the second sub-problem, reward It will present a sub-exponential distribution, and the UCB value of each gambling arm will be defined as follows:
[0180] ;
[0181] in, is the number of users, For the gambling arm, Indicates The number of times this gambling arm has been selected in iterations, For the current gambling arm The modulus of the reward values obtained in the iteration, is the total reward that the current system has obtained in the i-th iteration; The original problem in this iteration will be listed as follows:
[0182] ;
[0183] ;
[0184] Calculate under constraints The UCB value of each unselected gambling arm Here, the present invention can use a backtracking algorithm to calculate the color number and select the gambling arm with the largest UCB value.
[0185] At this time, the constraint condition is transformed into the color number of the graph, and the calculation of the color number is also very complicated. Here, a S-LUCB algorithm similar to the S-LUCB algorithm in the single-ring channel is proposed.
[0186] As explained above, the S-LUCB algorithm itself selects a set of actions in each iteration Similarly, assuming that in the S-LUCB algorithm In the iteration, and is the index set of the selected users and beams, The reward is , in Actions in iterations The rewards will be defined as follows:
[0187] ;
[0188] The problem of maximizing the UCB value of the selected action is The iterations will be listed as follows:
[0189] ;
[0190] ;
[0191] To solve the problem The UCB value of each unselected gambling arm needs to be calculated , and then the largest gambling arm of UCB can be selected and satisfied Here, a backtracking algorithm is used to calculate the color number.
[0192] Next go to The problem described is similar to the previous one, and the reward It will present a sub-exponential distribution, so the UCB value of each gambling arm will be defined as follows:
[0193] ;
[0194] In the The original problem in this iteration will be listed as follows:
[0195] ;
[0196] ;
[0197] Then we need to calculate the constraints The UCB value of each unselected gambling arm , the color number in the constraint will be obtained by the backtracking algorithm. And the gambling arm with the largest UCB value is selected from it.
[0198] The S-LUCB algorithm makes the actions converge to a suboptimal solution .set up , the upper bound of the regret value of the S-LUCB algorithm is .
[0199] The following is a simulation verification of the method in the embodiment of the present invention in conjunction with the specific implementation mode:
[0200] The performance of the proposed scheme is verified by simulation. Consider a base station equipped with 64 antennas. To provide services to single-antenna users, in the MSC channel model, the azimuth center angle of each cluster is evenly distributed in Then in the proposed scheme, Set to 0.01, The present invention compares the user scheduling and beamforming scheme based on 2D-CMAB design with the two-stage beamforming scheme without user scheduling and the traditional JSDMs (including the complete precoding scheme) using auxiliary modules. The full precoding scheme is inefficient due to the need for large DTL and uplink feedback. For fairness, in the traditional JSDM, the average angular spread of all groups is further set to be equal to NAS, i.e. , so the number of groups is (The value after round operation is taken). In the traditional JSDM, The effective channel dimension of a group is ,in It is The number of users in the group.
[0201] Considering the pilot overhead, the user The effective spectral efficiency (ESE) will be expressed as follows:
[0202] ;
[0203] In the formula, represents the number of symbols per coherence interval, Indicates DTL of each user. A resource block in LTE consists of 168 complex symbols, which is 14 OFDM symbols in two time slots multiplied by 12 subcarriers. Set to 100. In traditional JSDM, each group estimates the effective channel matrix separately, and the minimum DTL of users in the same group is the same. As shown in equation (17), the minimum DTL of all users in the proposed scheme is In order to better demonstrate the effective spectrum efficiency ESE, consider the case where the pilot has the minimum DTL, that is, the DTL is equal to the minimum DTL.
[0204] Figure 2 The ESE of different schemes for different numbers of users in single-loop channel (top) and MSC channel (bottom) are shown. The signal-to-noise ratio is set to 10 dB. τ is set to 30. The proposed 2D-CMAB-based scheme has higher ESE than other schemes. When the number of users is small, the performance of the MAB-based scheme and the proposed scheme are similar. The reasons are as follows. For a small number of users, the interference between users is not large, and the multi-user gain improves the ESE. However, when the number of users is large, the interference between users is high, which reduces the ESE. Figure 3 , Figure 4 and Figure 5 The ESE of single-ring channel (upper) and MSC channel (lower) under different SNRs are compared. The proposed 2D-CMAB-based scheme has higher ESE than other schemes.
[0205] In summary, the present invention overcomes the deficiencies in the prior art, provides an improved solution for two-stage beamforming taking user scheduling into consideration, and solves the problems in the prior art such as the inability to cope with multi-user scenarios, too long downlink training length, and low effective spectrum efficiency. The present invention provides a massive MIMO user scheduling and beamforming method based on MAB, including defining a combined multi-armed bandit problem (2D-CMAB) acting in two-dimensional space on the basis of a multi-armed bandit problem (MAB), and using the combined multi-armed bandit problem (2D-CMAB) to describe the pre-beamforming matrix and user scheduling design problem in a two-stage beamforming scheme (TSB); further dividing the objective function of the 2D-CMAB problem into two sub-problems, selecting a gambling arm (arm) in the first sub-problem to find zero PAS, and selecting an arm in the second sub-problem to maximize the received energy; for the application of a single-ring channel, an optimal linear UCB algorithm (O-LUCB) and a suboptimal linear UCB algorithm (S-LUCB) are proposed to solve the 2D-CMAB problem; in a multi-scattering cluster (MSC) channel, it is proved that DTL is equal to the color number of a graph to reduce the computational complexity, and a suboptimal linear S-LUCB algorithm with low computational complexity is proposed.
[0206] Embodiment 2, as Figure 6 As shown, a massive MIMO user scheduling and beamforming system based on MAB, the system comprising:
[0207] Modeling module, used to establish the communication model of TSB scheme and obtain the design problem of maximizing spectrum efficiency;
[0208] A conversion module, used for converting the obtained design problem of maximizing spectrum efficiency into a combinatorial multi-armed bandit problem in two-dimensional space;
[0209] A problem decomposition module is used to further divide the objective function of the two-dimensional combined multi-armed bandit problem into two sub-problems, in which the gambling arms are selected to find the zero value in the power angular spectrum in the first sub-problem, and in the second sub-problem, the gambling arms are selected to maximize the received energy;
[0210] The calculation module is used to solve the combined multi-armed bandit problem in two-dimensional space according to two sub-problems by using the optimal linear UCB algorithm and the suboptimal linear UCB algorithm to obtain a beamforming matrix.
[0211] The present invention defines a combined multi-armed bandit facing a two-dimensional problem, which is used to describe the pre-beamforming matrix and user scheduling design problem in the two-stage beamforming scheme (TSB); the problem is further divided into two sub-problems, first selecting an arm to find zero PAS, and then selecting an arm to maximize the received energy, thereby improving the spectrum efficiency. At the same time, the present invention proposes an optimal linear UCB algorithm and a suboptimal UCB algorithm to solve the MAB problem for the application of a single-ring channel; in the multi-scattering cluster (MSC) channel, the present invention proves that DTL is equal to the color number of the graph, and proposes a suboptimal linear UCB algorithm with low computational complexity. The simulation results verify the effectiveness of the proposed scheme, and through a horizontal comparison with the existing methods, it shows its advantages in improving the effective spectrum efficiency and reducing the feedback length and downlink training length.
[0212] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for massive MIMO user scheduling and beamforming based on MAB, characterized in that: The method comprises: Establish the TSB scheme communication model and obtain the design problem of maximizing spectrum efficiency; The obtained design problem of maximizing spectrum efficiency is converted into a combinatorial multi-armed bandit problem in two-dimensional space; The objective function of the two-dimensional combinatorial multi-armed bandit problem is further divided into two sub-problems. In the first sub-problem, the bandit arms are selected to find the zero value in the power angular spectrum, and in the second sub-problem, the bandit arms are selected to maximize the received energy. According to the two sub-problems, the optimal linear UCB algorithm and the suboptimal linear UCB algorithm are used to solve the two-dimensional combined multi-armed bandit problem and obtain the beamforming matrix. The objective function of the two-dimensional combined multi-armed bandit problem is further divided into two sub-problems, including: The first sub-problem is to standardize the likelihood ratio in the interval from 0 to 1 using a standardized utility function, and the standardized utility function is as follows: ; in, Indicates The channel vector of user The likelihood ratio of the signal energy after the operation of the DFT vector, Then The standardized form of the likelihood ratio The bigger, If the signal is transmitted with the same probability, then Its mapping non-zero; The first sub-problem is transformed into how to choose and Can make The sum is the largest, and the formula of the first sub-problem is expressed as follows: ; ; ; in, Indicates The channel vector of user The normalized likelihood ratio of the signal energy after the operation of the DFT vector, The set of indices representing the selected user, represents the set of indices of the selected DFT vector, is the limiting constant of the downlink training length, In this subproblem, selected users in a time window, In this subproblem, The selected DFT vector in the time window; because The likelihood ratio of the group is zero in the remaining set, There are multiple solutions, in each solution In, there is Make have or Make have , it is necessary to remove rows and columns with all zero elements from the solution so that the sum of the likelihood ratios is the maximum value; When the solution , select the remaining arms to maximize the received energy, then the second sub-problem is formulated as follows: ; ; ; in, Indicated in selected users in a time window, Indicated in The selected DFT vector in the time window, The set of indices representing the selected user, represents the set of indices of the selected DFT vector, For the User in The channel vector in the time window is The index is The DFT vector, is the limiting constant of the downlink training length, Indicated in In the sub-problem selected users in a time window, Indicated in In the sub-problem The selected DFT vector in the time window; According to the two sub-problems, the optimal linear UCB algorithm or the suboptimal linear UCB algorithm is used to solve the combined multi-armed bandit problem in two-dimensional space to obtain a beamforming matrix, including: based on the application of a single-loop channel, the optimal linear UCB algorithm and the suboptimal linear UCB algorithm are used to solve the combined multi-armed bandit problem in two-dimensional space to obtain a beamforming matrix; In the application scenario based on the single-loop channel, the optimal linear UCB algorithm is used to solve the combined multi-armed bandit problem in two-dimensional space to obtain the beamforming matrix, including: In the first subproblem, the utility function As indexed by the DFT vector With user index Gambling Arm The reward of a basic gambling arm in the optimal linear algorithm The UCB value is defined as follows: ; in, Represents the gambling arm in the optimal algorithm The UCB value, is the number of users, is the number of DFT vectors, Indicates The number of times this gambling arm has been selected in iterations, is the undetermined gambling arm set corresponding to the effective channel vector in the user channel, For the current gambling arm The modulus of the reward value obtained in the iteration; in order to select the super arm that minimizes the UCB value, a vector To indicate whether the user has been selected, use Indicates whether a beam has been selected. Representative beams have been selected, and 0 is the opposite. The problem of selecting the super arm of the slot machine will be expressed as follows: ; ; in, represents the transpose operation, is a 0-1 vector indicating whether each user has been selected. To indicate the A logical value indicating whether the user has been selected. A 0-1 vector indicating whether each DFT vector has been selected. For all channels, For the Channel vector for each user; Based on the application of single-loop channel, the suboptimal linear UCB algorithm is used to solve the combined multi-armed bandit problem in two-dimensional space, and the beamforming matrix is obtained, including: Consider the first sub-problem, assuming that In the iteration, and is the index set of the selected user and beam, and takes action in the next iteration The rewards will be defined as follows: ; in, For the The gambling arm Rewards, , is the union of the user set and the DFT vector set in the current gambling arm and the i-1th iteration, To indicate the The channel vector of user The normalized likelihood ratio of the signal energy after the operation of the DFT vector; because , easy to prove Therefore, in the suboptimal algorithm, In each gambling arm The UCB value is defined as follows: ; in, Represents the gambling arm in the first sub-problem of the suboptimal algorithm The UCB value, For the The gambling arm The upper limit of the reward value, Indicates The number of times this gambling arm has been selected in iterations, For the current gambling arm The modulus of the reward value obtained in the iteration; maximize the UCB value in the The original problem in this iteration will be listed as follows: ; ; To solve the problem, first calculate the UCB value of each unselected gambling arm , such as In , and select the maximum value from them; When considering the second sub-problem, let In the iteration, and The index set of the selected user and vector, which will be acted on in the next iteration The rewards will be defined as follows: ; in, For the The gambling arm Rewards, , is the union of the user set and the DFT vector set in the current gambling arm and the i-1th iteration, For the User in The channel vector in the time window is The index is DFT vector of award It will present a sub-exponential distribution, so the UCB value of each gambling arm will be defined as follows; ; in, Represents the gambling arm in the second sub-problem of the suboptimal algorithm The UCB value, is the number of users, For the gambling arm, Indicates The number of times this gambling arm has been selected in iterations, For the current gambling arm The modulus of the reward values obtained in the iteration, is the total reward that the current system has obtained in the i-th iteration; The original problem in this iteration will be listed as follows: ; ; Calculate the UCB value of each unselected gambling arm , and select the maximum value from them.
2. The method for massive MIMO user scheduling and beamforming based on MAB according to claim 1, characterized in that: The TSB scheme communication model is established to obtain the design problem of maximizing spectrum efficiency, including: Using DFT matrix To design the pre-beamforming matrix, , Indicates the unprocessed signal at the sender. Indicated in The selected users in the time window are Indicated in The selected DFT vector in the time window, then the pre-beamforming matrix is expressed as , in The effective channel matrix in the time window is expressed as ; set up for The base number, represents a complex set, L is the number of base station array elements, then when the base station sends a signal, the pilot matrix is At that time, The signal received by each user is: ; in, represents the conjugate transposed matrix, Indicates Users in The signal received at any time, and For the The channel vector of the user Noise in a time window; In the applicable channel, the problem is transformed into designing a and , so that the spectrum efficiency is maximized, according to the derivation of Shannon's formula, the energy of the received signal When the maximum value is reached, the spectrum efficiency index will be improved accordingly, because , the objective function of the design problem of the pre-beamforming matrix is: ; in, Indicated in selected users in a time window, Indicated in The selected DFT vector in the time window, The set of indices representing the selected user, represents the set of indices of the selected DFT vector, For the User in The channel vector in the time window is The index is The DFT vector, The length of downlink training The limiting constant.
3. The MAB-based massive MIMO user scheduling and beamforming method according to claim 2, characterized in that: The design problem of maximizing the spectrum efficiency is converted into a two-dimensional combined multi-armed bandit problem, including: The two-dimensional combinatorial multi-armed bandit problem has two sets of variables and , and represents the number of elements in each of these two variable sets, is the number of users, is the number of DFT vectors, the combination of each pair of elements of the two variable sets All have an unknown mean, and each mean satisfies independent and identical distribution; The problem It is applicable to the problem of a combined multi-armed bandit in a two-dimensional space corresponding to two sets of variables. The set of indices representing the selected user, , represents the set of indices of the selected DFT vector, , the super arm of the multi-armed slot machine is The corresponding reward is .
4. The method for massive MIMO user scheduling and beamforming based on MAB according to claim 1, characterized in that: Also includes: Based on the multi-scattering cluster channel, the calculation of the downlink training length in the constraint condition is converted into a graph coloring problem. The suboptimal linear UCB algorithm is used to solve the color number of the graph to obtain the beamforming matrix. The calculation of the downlink training length in the constraint condition is converted into a graph coloring problem, including: Use the color number of the graph to reduce the computational complexity and satisfy The minimum DTL value is ,in, For users The effective channel matrix is represents transpose, is a 0-1 matrix representing the relationship between the DFT vector and the user pairing, express The chromatic number of the graph of the adjacency matrix is for A solution of , the pilot matrix is designed as ,in, It is arbitrary Unitary matrix of size No. List.
5. The method for massive MIMO user scheduling and beamforming based on MAB according to claim 4, characterized in that: Based on the multi-scattering cluster channel, the downlink training length is converted into the color number of the graph, and the suboptimal linear UCB algorithm is used to solve the color number of the graph to obtain the beamforming matrix, including: Suppose that in the suboptimal linear algorithm In the iteration, and is the index set of the selected users and vectors, The reward is , in Actions in iterations The rewards will be defined as follows: ; in, Represents the gambling arm in the first subproblem of the MSC channel The UCB value, For the The gambling arm The upper limit of the reward value, Indicates The number of times this gambling arm has been selected in iterations, For the current gambling arm The modulus of the reward values obtained in iterations; The problem of maximizing the UCB value of the selected action is The iterations will be listed as follows: ; ; To solve the first sub-problem, first calculate the UCB value of each unselected gambling arm , and then the largest gambling arm of UCB was selected and satisfied , and then use the backtracking algorithm to calculate the color number; Solve the second sub-problem, reward It will present a sub-exponential distribution, and the UCB value of each gambling arm will be defined as follows: ; in, Represents the gambling arm in the second sub-problem of the MSC channel The UCB value, is the number of users, For the gambling arm, Indicates The number of times this gambling arm has been selected in iterations, For the current gambling arm The modulus of the reward values obtained in the iteration, is the total reward that the current system has obtained in the i-th iteration; The original problem in this iteration will be listed as follows: ; ; Calculate under constraints The UCB value of each unselected gambling arm , the number of colors in the constraint will be obtained by the backtracking algorithm, and the gambling arm with the largest UCB value will be selected from it.
6. A massive MIMO user scheduling and beamforming system based on MAB, characterized in that: The system comprises: Modeling module, used to establish the communication model of TSB scheme and obtain the design problem of maximizing spectrum efficiency; A conversion module, used for converting the obtained design problem of maximizing spectrum efficiency into a combinatorial multi-armed bandit problem in two-dimensional space; A problem decomposition module is used to further divide the objective function of the two-dimensional combined multi-armed bandit problem into two sub-problems, in which the gambling arms are selected to find the zero value in the power angular spectrum in the first sub-problem, and in the second sub-problem, the gambling arms are selected to maximize the received energy; A calculation module is used to solve the combined multi-armed bandit problem in two-dimensional space according to two sub-problems by using an optimal linear UCB algorithm and a suboptimal linear UCB algorithm to obtain a beamforming matrix; The objective function of the two-dimensional combined multi-armed bandit problem is further divided into two sub-problems, including: The first sub-problem is to standardize the likelihood ratio in the interval from 0 to 1 using a standardized utility function, and the standardized utility function is as follows: ; in, Indicates The channel vector of user The likelihood ratio of the signal energy after the operation of the DFT vector, Then The standardized form of the likelihood ratio The bigger, If the signal is transmitted with the same probability, then Its mapping non-zero; The first sub-problem is transformed into how to choose and Can make The sum is the largest, and the formula of the first sub-problem is expressed as follows: ; ; ; in, Indicates The channel vector of user The normalized likelihood ratio of the signal energy after the operation of the DFT vector, The set of indices representing the selected user, represents the set of indices of the selected DFT vector, is the limiting constant of the downlink training length, In this subproblem, selected users in a time window, In this subproblem, The selected DFT vector in the time window; because The likelihood ratio of the group is zero in the remaining set, There are multiple solutions, in each solution In, there is Make have or Make have , it is necessary to remove rows and columns with all zero elements from the solution so that the sum of the likelihood ratios is the maximum value; When the solution , select the remaining arms to maximize the received energy, then the second sub-problem is formulated as follows: ; ; ; in, Indicated in selected users in a time window, Indicated in The selected DFT vector in the time window, The set of indices representing the selected user, represents the set of indices of the selected DFT vector, For the User in The channel vector in the time window is The index is The DFT vector, is the limiting constant of the downlink training length, Indicated in In the sub-problem selected users in a time window, Indicated in In the sub-problem The selected DFT vector in the time window; According to the two sub-problems, the optimal linear UCB algorithm or the suboptimal linear UCB algorithm is used to solve the combined multi-armed bandit problem in two-dimensional space to obtain a beamforming matrix, including: based on the application of a single-loop channel, the optimal linear UCB algorithm and the suboptimal linear UCB algorithm are used to solve the combined multi-armed bandit problem in two-dimensional space to obtain a beamforming matrix; In the application scenario based on the single-loop channel, the optimal linear UCB algorithm is used to solve the combined multi-armed bandit problem in two-dimensional space to obtain the beamforming matrix, including: In the first subproblem, the utility function As indexed by the DFT vector With user index Gambling Arm The reward of a basic gambling arm in the optimal linear algorithm The UCB value is defined as follows: ; in, Represents the gambling arm in the optimal algorithm The UCB value, is the number of users, is the number of DFT vectors, Indicates The number of times this gambling arm has been selected in iterations, is the undetermined gambling arm set corresponding to the effective channel vector in the user channel, For the current gambling arm The modulus of the reward value obtained in the iteration; in order to select the super arm that minimizes the UCB value, a vector To indicate whether the user has been selected, use Indicates whether a beam has been selected. Representative beams have been selected, and 0 is the opposite. The problem of selecting the super arm of the slot machine will be expressed as follows: ; ; in, represents the transpose operation, is a 0-1 vector indicating whether each user has been selected. To indicate the A logical value indicating whether the user has been selected. A 0-1 vector indicating whether each DFT vector has been selected. For all channels, For the Channel vector for each user; Based on the application of single-loop channel, the suboptimal linear UCB algorithm is used to solve the combined multi-armed bandit problem in two-dimensional space, and the beamforming matrix is obtained, including: Consider the first sub-problem, assuming that In the iteration, and is the index set of the selected user and beam, and takes action in the next iteration The rewards will be defined as follows: ; in, For the The gambling arm Rewards, , is the union of the user set and the DFT vector set in the current gambling arm and the i-1th iteration, To indicate the The channel vector of user The normalized likelihood ratio of the signal energy after the operation of the DFT vector; because , easy to prove Therefore, in the suboptimal algorithm, In each gambling arm The UCB value is defined as follows: ; in, Represents the gambling arm in the first sub-problem of the suboptimal algorithm The UCB value, For the The gambling arm The upper limit of the reward value, Indicates The number of times this gambling arm has been selected in iterations, For the current gambling arm The modulus of the reward value obtained in the iteration; maximize the UCB value in the The original problem in this iteration will be listed as follows: ; ; To solve the problem, first calculate the UCB value of each unselected gambling arm , such as In , and select the maximum value from them; When considering the second sub-problem, let In the iteration, and The index set of the selected user and vector, which will be acted on in the next iteration The rewards will be defined as follows: ; in, For the The gambling arm Rewards, , is the union of the user set and the DFT vector set in the current gambling arm and the i-1th iteration, For the User in The channel vector in the time window is The index is DFT vector of award It will present a sub-exponential distribution, so the UCB value of each gambling arm will be defined as follows; ; in, Represents the gambling arm in the second sub-problem of the suboptimal algorithm The UCB value, is the number of users, For the gambling arm, Indicates The number of times this gambling arm has been selected in iterations, For the current gambling arm The modulus of the reward values obtained in the iteration, is the total reward that the current system has obtained in the i-th iteration; The original problem in this iteration will be listed as follows: ; ; Calculate the UCB value of each unselected gambling arm , and select the maximum value from them.
Citation Information
Patent Citations
Two-stage precoding method based on MAB under multi-scattering-cluster channel
CN115865155A
Passive beam forming method based on deep reinforcement learning and RIS partitioning
CN118646461A