A Two-Stage Beamforming Method Based on Combinatorial Multi-Armed Bandit

By combining the multi-arm gambling machine to design the pre-beamforming matrix, the pre-beamforming matrix design is transformed into an arm selection problem, and the 0-1 integer linear planning and branch delimiting method are used to solve the problem of large overhead of the pre-beamforming matrix in the prior art, and the significant improvement of spectrum efficiency and the reduction of channel estimation are achieved.

CN115996078BActive Publication Date: 2025-07-08NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211474774.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-23
Publication Date
2025-07-08
Estimated Expiration
2042-11-23

AI Technical Summary

Technical Problem

The two-stage beamforming method based on the channel covariance matrix in the prior art costs a large overhead and low efficiency when designing the pre-beam formation matrix, especially the overhead problems caused by frequent changes in CCM when the user moves quickly.

Method used

The pre-beamforming matrix is designed using a combined multi-arm gambling machine. By converting the pre-beamforming matrix design problem into the arm selection problem in the combined multi-arm gambling machine and converting it into a 0-1 integer linear planning problem, the super arm is selected using a linear UCB strategy, and the solution is combined with the branch delimiting method, the precoder is designed and the angle spectrum of each user is determined to reduce channel estimation overhead.

Benefits of technology

Significantly improve spectrum efficiency and reduce channel estimation overhead, proving that it regrets to increase logarithmicly over time, converges to optimal action, and achieves efficient spectrum utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115996078B_ABST
    Figure CN115996078B_ABST
Patent Text Reader

Abstract

The present invention discloses a two-stage beamforming method based on a combinatorial multi-armed bandit, which includes designing a pre-beamforming matrix by using a combinatorial multi-armed bandit; transmitting a pilot matrix with a cyclic structure and receiving an estimated equivalent channel matrix of users according to the pilot matrix; designing a precoder and determining the angular spectrum of each user according to the equivalent channel matrix; and correcting the pre-beamforming matrix by using the precoder and the angular spectrum of each user. The present invention transforms the pre-beamforming matrix design problem into a CMAB problem, designs a precoder and determines the angular spectrum of each user according to the equivalent channel matrix, and corrects the pre-beamforming matrix by using the precoder and the angular spectrum of each user. Compared with the previous methods, it can significantly improve the spectral efficiency, significantly reduce the channel estimation overhead and improve the spectral efficiency without the need for CCM.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technologies, and particularly to a two-stage beamforming method based on a combinatorial multi-armed bandit. Background Art

[0002] Large-scale multiple-input multiple-output (MIMO) systems show great potential in 6G wireless communication. Since channel reciprocity does not exist in frequency-division duplexing (FDD), and the downlink training length (DTL) for obtaining instantaneous channel state information (CSI) and CSI feedback impose a heavy burden on FDD systems, making FDD large-scale MIMO challenging. Considering that statistical CSI remains unchanged within intervals much longer than the symbol transmission time scale, a two-stage beamforming (TSB) scheme based on the channel covariance matrix (CCM), namely JSDM, was proposed, which addresses DTL and CSI feedback. In the first stage, JSDM groups users and designs a pre-beamformer based on statistical CSI to reduce the dimension of the effective channel matrix, resulting in reduced DTL and CSI feedback. In the second stage, a multi-user precoder is designed to reduce interference. Many works have been dedicated to improving the performance of the TSB scheme. The JSDM scheme based on the minimum mean square error (MMSE) designs the pre-beamformer and multi-user precoder according to the MMSE criterion and the weighted MMSE criterion, respectively. To address the grouping problem, a neighbor-based JSDM (NJSDM) that fully utilizes the signal space was proposed to achieve higher spectral efficiency. The above TSB schemes assume an available CCM. Khalilsarai et al. estimated the CCM and proposed a TSB scheme called ACS that uses the DFT matrix to design the pre-beamformer. In practice, the overhead of obtaining the CCM is significant, especially when users move fast and the CCM changes frequently. Summary of the Invention

[0003] The purpose of the present invention is to provide a two-stage beamforming method based on a combinatorial multi-armed bandit to solve the problems of high overhead and low efficiency in designing the pre-beamforming matrix using the CCM in the prior art.

[0004] To achieve the above object, the present invention is implemented by the following technical solutions:

[0005] A two-stage beamforming method based on a combinatorial multi-armed bandit, comprising:

[0006] Designing a pre-beamforming matrix using a combinatorial multi-armed bandit;

[0007] Transmitting a pilot matrix with a cyclic structure and receiving the equivalent channel matrix estimated by the user according to the pilot matrix;

[0008] Design a precoder according to the equivalent channel matrix and determine the angular spectrum of each user;

[0009] Modify the pre-beamforming matrix by using the precoder and the angular spectrum of each user.

[0010] Furthermore, designing the pre-beamforming matrix by using the combinatorial multi-armed bandit includes:

[0011] Convert the design problem of the pre-beamforming matrix into an arm selection problem in the combinatorial multi-armed bandit;

[0012] Convert the arm selection problem in the combinatorial multi-armed bandit into a 0-1 integer linear programming problem.

[0013] Furthermore, converting the design problem of the pre-beamforming matrix into an arm selection problem in the combinatorial multi-armed bandit includes:

[0014] The pre-beamforming matrix consists of the DFT matrix D = [d1, d2, …, d M ;

[0015] where is the basic arm, i = 1, 2, 3, …, the vector is the direction vector, θ is the incident angle, M is the number of antennas, λ c is the wavelength, d is the antenna spacing, T is the matrix transpose symbol, and j is the imaginary part symbol;

[0016] Define the instantaneous channel state information interval as a time block. Assume that the angular range AR of user k at the th time block is and are the lower and upper limits of the angle of user k respectively. At the th time block, set a large constant Δ such that the AR of user k belongs to Let the set be the set of i, is the channel vector at the tth time slot in the tth time block As the number of antennas approaches infinity, there is:

[0017]

[0018] where d l is the lth column of the DFT matrix D;

[0019] Use to represent the selected set of DFT vectors, and use to indicate the pre-beamforming matrix. There is:

[0020]

[0021] Among them, ∩ refers to the intersection symbol of sets, refers to and the intersection of, matrix is composed of

[0022] the columns of the DFT matrix constitute;

[0023] Let the circular structure of the pilot matrix be Among them, refers to the set the number of elements in, refers to the set of complex matrices of dimension , L is the downlink training length, [x1, x2, …, x L , x1, x2, …] refers to the matrix composed of the vectors x1, x2, …, x L , x1, x2, … arranged in sequence, L ≥ max k d k , max k d k refers to the largest value among all d k in, d k is the number of elements in, and the matrix [x1, x2, …, x L is a unitary matrix. The received signal of user k in channel estimation is:

[0024]

[0025] Among them, n H (t) represents the noise at the t-th moment, H represents the conjugate transpose; Substituting (4) into (5) and removing the zero terms, we get:

[0026]

[0027] Among them, is a high unitary matrix composed of partial vectors x i ;

[0028] Use the symbol to represent the in formula (6), where t refers to the t-th time slot and is estimated using the least squares criterion; According to the least squares criterion, multiply formula (6) by to get:

[0029]

[0030] And is estimated as and fed back to the base station, using and formula (4), the base station obtains the estimated value of, and represents the estimated value of as where t is the t-th time slot;

[0031] In the case of a limited downlink training length, design to maximize the spectral efficiency and maximize the received signal energy, that is to obtain a high spectral efficiency; considering formula (4), let the number of elements in be d k , and set the downlink training length to κmax k d k , where κ is a constant, and the pre-beamforming matrix design problem is:

[0032]

[0033] s.t. max k d k <τ. (8)

[0034] where DTL is the pilot length; in , multiple DFT vectors are required, so it is an arm selection problem in the combinatorial multi-armed bandit.

[0035] Furthermore, transforming the arm selection problem in the combinatorial multi-armed bandit into a 0-1 integer linear programming problem includes:

[0036] M basic arms d1, d2, …, d M are labeled from 1 to M, and let be the reward of the i-th basic arm, and the super arm is composed of the basic arm a i , i = 1, 2, ···, where a i ∈ {1, 2, ··· M}, given a super arm A a(t) , the pre-beamforming matrix is and the reward is the power of the received signal given by the following formula:

[0037]

[0038] Using the linear UCB strategy to select the super arm in each time slot, considering the chi-square distribution of the reward, the UCB value of the basic arm in the t-th time slot is defined as:

[0039]

[0040] where is an arbitrary positive constant. m i,t is the number of times action i is selected. Then the UCB value of the super-arm is Prove that it converges to the optimal action of the latter;

[0041] In each time slot, use x to represent whether a base arm is selected in the super-arm, where x i = 1 means the i-th base arm is selected, and z i = 0 means the i-th base arm is not selected; Let u = [u1(y), u2(t) ···]. Define the matrix A, where the l-th element of its k-th row a k is:

[0042]

[0043] where is the element in the k-th row and l-th column. Then the problem of maximizing the UCB value is:

[0044]

[0045] s.t. Ax ≤ τe (15)

[0046] where e is a vector with all elements being 1, and τ is a given value used to describe the pilot length;

[0047] When the of all users decreases to convergence, is regarded as the actual AR, and the downlink training length is minimized, resulting in having the largest solution; The actions are divided into two groups, one is the candidate set involving the edge vectors and the other is the remaining set containing the remaining vectors x is divided into x_1 and x_2, where x_1 represents the actions in the candidate set and x_2 represents the remaining actions; A is divided into A_1 and A_2, that is, A_1 is the set of vectors and a i is the i-th column of the matrix A, and A2 is the set of vectors ; The vector u is divided into u1 and u2, that is, u1 is the set of vectors and u2 is the set of vectors ;

[0048] Then is divided into two sub-problems. The first method is to select the base arms in to maximize the UCB value:

[0049]

[0050] s.t. A1x1 ≤ τe,, (17)

[0051] Among them, without considering the arms in, once x1 is obtained, we replace x1 with Then consider the basic arms in; The second sub-problem formula is:

[0052]

[0053] s.t. A2x2 ≤ τe - A2x1 (19)

[0054] and are linear 0-1 integer programming problems.

[0055] Furthermore, it also includes:

[0056] In each time slot, update the rewards of each basic arm;

[0057] Use the likelihood ratio to calculate the power of each basic arm in the candidate set. If the basic arm in the candidate set has zero or non-zero power, update the marginal vector; The expected regret of the CMAB algorithm is O(lnt). And it can be proved that its regret grows logarithmically with time.

[0058] Furthermore, use the branch and bound method to solve the 0-1 integer linear programming problem.

[0059] Furthermore, it also includes using the maximum likelihood method to detect the power of the angle spectrum.

[0060] According to the above technical solutions, the embodiments of the present invention have at least the following effects: The present invention transforms the pre-beamforming matrix design problem into a CMAB problem, designs a pre-coder and determines the angle spectrum of each user according to the equivalent channel matrix; modifies the pre-beamforming matrix using the pre-coder and the angle spectrum of each user. The present invention proves that the regret grows logarithmically with time, making the proposed scheme converge to the optimal action, and the present invention can significantly improve the spectrum efficiency compared with the previous methods and can significantly reduce the channel estimation overhead and improve the spectrum efficiency without the need for CCM. Description of the Drawings

[0061] Figure 1 It is a schematic diagram of the ESE simulation of different schemes under different SNRs for an embodiment of the present invention;

[0062] Figure 2 It is a schematic diagram of the ESE simulation of different schemes under different numbers of users for an embodiment of the present invention;

[0063] Figure 3Schematic diagram of ESE simulation for different schemes under different DTLs in an embodiment of the present invention;

[0064] Figure 4 Schematic diagram of ESE simulation of the CMAB scheme in each time slot in an embodiment of the present invention. Detailed implementation manners

[0065] To make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific implementation manners.

[0066] For a single-cell FDD massive MIMO system, the present invention proposes a two-stage beamforming (TSB) method based on combinatorial multi-armed bandit (CMAB). Different from other TSB schemes, the pre-beamforming matrix in this application is designed by CMAB and does not require CCM.

[0067] In the first stage of the TSB scheme, we transform the pre-beamforming design problem into a CMAB problem and regard the pre-beamforming matrix as a super-arm. We use the linear UCB strategy to select the super-arm and transform the super-arm selection problem into a linear integer 0-1 programming problem. The branch and bound method is used to solve this programming problem.

[0068] In the second stage of the TSB scheme, a multi-user precoder is designed to reduce interference. In the CMAB process, we use the maximum likelihood detection method to determine the power angle spectrum (PAS) and iteratively reduce the angle range. It is proved that this method can achieve results that grow logarithmically with time. The simulation results also verify the good performance of this scheme.

[0069] This application can significantly reduce the channel estimation overhead and improve the spectral efficiency without the need for CCM. Specifically, the method includes the following steps: using a combinatorial multi-armed bandit to design the pre-beamforming matrix; transmitting a pilot matrix with a cyclic structure and receiving the equivalent channel matrix estimated by the user according to the pilot matrix; designing a precoder based on the equivalent channel matrix and determining the angular power spectrum of each user; using the precoder and the angular power spectrum of each user to correct the pre-beamforming matrix.

[0070] To better explain the technical solution of the present invention, preliminary work needs to preprocess relevant parameters and data. Among them:

[0071] Consider a single-cell FDD massive MIMO system, where the base station (BS) serves K single-antenna user equipments (UEs). The BS uses a ULA array with M elements. During downlink transmission, let be the downlink channel vector from the BS to user k, be the received signal, where yK is the received signal for user k. Let B and W be the pre-beamforming matrix and multi-user precoding matrix in the TSB scheme respectively. We have:

[0072] y = H H BW s + n, (1)

[0073] where the downlink channel vector is the transmitted signal, and the variance of the transmitted signal is I is the identity matrix, indicates that the Gaussian noise follows a uniformly symmetric complex Gaussian distribution. In the first stage of the TSB scheme, the pre-beamforming matrix B is designed to make the effective channel matrix H H B have a special structure, so as to reduce the DTL and channel feedback required for estimating H H B. In the second stage, the multi-user precoder is designed to eliminate interference, where is the pseudo-inverse operation.

[0074] The single-ring channel model is considered, where user k has an azimuth center angle θ k and an angular spread (AS) Δ k . Through θ k and Δ k , the CCM is calculated as:

[0075]

[0076] where is the steering vector, () * is the conjugate, λ c is the wavelength, is the spacing between two antenna elements. γ k (θ) is the channel power angular spectrum (PAS), η k is the normalized density function image Using the Karhunen-Loeve representation, the channel vector of user k is expressed as where is the small-scale fading.

[0077] Next, the TSB technical solution based on CMAB proposed by the present invention will be introduced, where:

[0078] In the first aspect, relevant inferences are made on the formula problem, including:

[0079] In the first stage of the proposed TSB scheme, the pre-beamforming matrix consists of the DFT matrix D = [d1,..., d M . Among them, is the basic arm, i = 1, 2, 3,..., vector is the direction vector, θ is the incident angle, M is the number of antennas, and λ c is the wavelength, d is the antenna spacing, T is the matrix transpose symbol, and j is the imaginary part symbol.

[0080] We define the statistical CSI interval that remains unchanged as the time block (TB). Assume that the AR of user k at the th TB is At the th TB, we set a large constant Δ such that the AR of user k belongs to Let the set be the set of i, and is the channel vector at the t-th time slot in the t-th TB As the number of antennas approaches infinity, we have:

[0081]

[0082] Denote the selected DFT vector set by , and the pre-beamforming matrix is indicated by . Then,

[0083]

[0084] is a sparse vector. Among them, ∩ refers to the intersection symbol of sets, refers to and 's intersection. The matrix is composed of the columns of the DFT matrix. Let the circular structure of the pilot matrix be Among them, refers to the number of elements in the set , refers to the set of complex matrices with dimension , [x1, x2, …, x L , x1, x2, …] refers to the matrix composed of the vectors x1, x2, …, x L , x1, x2, … arranged in sequence. L ≥ max k d k , max k d k refers to the largest value among all d k , d k is 's number of elements, and the matrix [x1, x2, …, x L is a unitary matrix. In channel estimation, the received signal of user k is:

[0085]

[0086] where n H (t) represents the noise at the t-th moment, H denotes the conjugate transpose; substituting (4) into (5) and removing the zero terms, we obtain

[0087]

[0088] where, is also a high unitary matrix composed of partial vectors x i , then can be estimated using the least squares (LS) criterion, DTL is L, and L is the downlink training length. According to the LS criterion, we multiply (6) by and then obtain

[0089]

[0090] while is estimated as and fed back to the base station. Using and (4), the BS can obtain In the second stage, the multi-user precoder W is designed as to mitigate interference.

[0091] The main problem is to design to maximize the spectral efficiency (SE) under the condition of limited DTL. We maximize the received signal energy, that is to obtain a high SE. Let the number of elements in be d k . To better estimate the channel, set DTL to κmax k d k , where κ is a constant.

[0092] The problem of pre-beamforming matrix design is

[0093]

[0094] s.t. max k d k <τ. (8)

[0095] where DTL is the pilot length, in , multiple DFT vectors are required, so is the arm selection problem in CMAB. Next, we show a method to reduce the user AR, thereby improving the accuracy of AR and the effective SE (ESE).

[0096] On the second aspect, the corresponding adjustment is made to its adaptive angle range, including:

[0097] In the previous subsection, it is large due to the large constant Δ, resulting in a long DTL and a low ESE. To improve the ESE, we propose an adaptive adjustment method to make closer to the actual AR, thereby reducing the DTL.

[0098] Considering (3), assume If the power of user k on the edge vector (the basic arm (selection action) corresponding to the lower limit of the angle of user k) or is zero, where represents the lo k -th column of the DFT matrix, represents the up k -th column of the DFT matrix. Then the PAS at the corresponding angle is zero, and the AR can be reduced. Therefore, we iteratively reduce the angle range and consider the edge vector or the power of user k at the vector (selection action) d m as follows. We assume that if p m,k is less than the constant ζ 2 , then the power is zero. Otherwise, the power is non-zero. Considering (7), if the vector d m is included in , then considering d m , we have:

[0099]

[0100] where the channel of user k at the m-th vector at the t-th moment can be expressed as the noise of user k at the m-th vector at the t-th moment and the received signal y of user k at the k,m (t)-th vector are the elements corresponding to d m in and respectively. Since is a Gaussian variable with zero mean and variance p m,k . Therefore, is a Gaussian variable with zero mean and variance If the vector d m is selected in n time slots, we have y k,m , where is the received signal of user k at the first moment in the m-th vector, is the received signal of user k at the n-th moment in the m-th vector. We can use y k,m to detect the variance and obtain p m,k . Considering (9), the log-likelihood ratio (LLR) is:

[0101]

[0102] Let where Γ(·) is the gamma function. Calculate it numerically and obtain If LLR≥ε, the power is detected as non-zero, where ε is a constant. If , the power is detected as 0. In other cases, the power is uncertain. ε is a constant used to represent the probability of correctly detecting the power. The larger ε is, the greater the probability of correctly detecting the power.

[0103] In the third aspect, a TSB algorithm based on CMAB is proposed, including:

[0104] The present invention proposes an algorithm based on CMAB to solve . The DFT vectors are regarded as base arms, and M base arms d1, d2, …, d M are labeled from 1 to M. Let be the reward of the i-th base arm. The super arm is composed of the base arms a i , i = 1, 2, ···, where a i ∈{1, 2, ··· M}. Given a super arm the pre-beamforming matrix is Use to represent the reward of the super arm at the t-th moment, and the reward is given by the following formula:

[0105]

[0106] where refers to the square of the Frobenius norm of matrix A. We use the linear UCB strategy to select the super arm in each time slot. Considering the chi-square distribution of the reward, the UCB value of the base arm in the t-th time slot is defined as:

[0107]

[0108] where refers to the average value of the rewards at the previous t - 1 moments, is any positive constant, m i,tis the number of times action i is selected. Then the super arm has a UCB value of Prove that it converges to the optimal action of the latter.

[0109] In each time slot, we use x to represent whether a base arm is selected in the super arm, where x i = 1 indicates that the i-th base arm is selected, and x i = 0 indicates that the i-th base arm is not selected. Let u = [u1(t), u2(t) ··· u M (t)]. Define the matrix A, where the l-th element of its k-th row a k is

[0110]

[0111] The problem of maximizing the UCB value is

[0112]

[0113] s.t. Ax ≤ τe (15)

[0114] where e is a vector with all elements equal to 1.

[0115] To reduce DTL and thus improve ESE, an accurate AR should be found When the of all users decreases to convergence, is regarded as the actual AR, and the DTL is minimized, thus obtaining The solution of is the largest. Therefore, the action corresponding to the edge vector has a higher priority. Therefore, we divide the actions into two groups. One is the candidate set involving the edge vector The other is the remaining set containing the remaining vectors x is divided into x1 and x2, where x1 represents the actions in the candidate set and x2 represents the remaining actions. Correspondingly, A is divided into A1 and A2, that is, A1 is the set of vectors and a i is the i-th column of matrix A, and A2 is the set of vectors ; the vector u is divided into u1 and u2, that is, u1 is the set of vectors ; u2 is the set of vectors .

[0116] Then is divided into two sub-problems. The first method is to select the base arms in to maximize the UCB value

[0117]

[0118] s.t. A1x1 ≤ τe, (17)

[0119] Disregarding the arm in Then consider the base arm in

[0120]

[0121] s.t. A2x2 ≤ τe - A2x1. (19)

[0122] and is a linear 0 - 1 integer programming problem and can be solved by the branch - and - bound method.

[0123] We have shown the process of obtaining the super - arm in each time slot. In each time slot, the reward of each basic arm can be updated. Then, we use the LLR to calculate the power of the basic arms in the candidate set. If it is determined that the base arm in the candidate set has zero or non - zero power, the marginal vector should be updated. The details are given in Algorithm 1.

[0124] Algorithm 1: TSB Algorithm Based on CMAB

[0125] Data:

[0126] Result:

[0127] Step 1: Initialize m i (t) = 0;

[0128] Step 2: For each k, calculate and

[0129] Step 3: Calculate u i (t) using (12);

[0130] Step 4: Solve and according to the proposed algorithm to obtain

[0131] Step 5: Calculate the likelihood ratio and update lo k , up k and m i (t), and go back to Step 3.

[0132] Theorem 1: The expected regret of the CMAB algorithm is O(ln t).

[0133] Proof: In (10), there exists a constant Make the AR of all users remain unchanged after time slots. Next, we consider the latter time slots. Define the variable Z i,t . In the t-th time slot, if the optimal action a * is selected, then Z i,t remains unchanged. If a non-optimal action a is selected, then Z i,t is incremented by 1, where i = arg min j∈a m j,t (if there are multiple solutions, we can arbitrarily choose one of them). The number of non-optimal actions selected is and Z i,t ≤ m i,t . Let I i,t denote the indicator function, and if Z i,t increases at time t + 1, the indicator function is 1. Let l be any positive integer. Then

[0134]

[0135] where 1(x) is 1 when the event x is true and 0 when it is false. When I i,t = 1, a non-optimal action a(p) is selected, where Then

[0136]

[0137] where Let Q * denote the number of elements in the set and Q p denote the number of elements in the set at the (p + 1)-th moment. Thus, l ≤ Z i,p ≤ m i,p ,

[0138]

[0139] denotes that at least one of the following events must be true.

[0140] Event ε1:

[0141] Event ε2:

[0142] Event ε3:

[0143] where the reward corresponding to the optimal action a * and and R iThe mean reward for action i. For ε1, there is at least one j such that We have

[0144]

[0145] In (9), follows a chi-square distribution and is a sub-exponential function, where

[0146] It can be seen that is a sub-exponential function. Applying the Chernoff-Hoeffding bound for this sub-exponential function, the upper bound for each term is:

[0147]

[0148] where And, results in Since

[0149]

[0150] Therefore,

[0151]

[0152] Thus where As a result,

[0153]

[0154] Moreover, the average value of the probability of event ε1 is

[0155]

[0156] Similarly, the average value of the probability of event ε2 is

[0157]

[0158] Let, and set

[0159]

[0160] And

[0161]

[0162] Since In (28) and (29),

[0163] makes E3 always satisfied. Note

[0164]

[0165]

[0166] And

[0167]

[0168]

[0169] Therefore, the value Z of action i at time t i,t The mean value satisfies Where Is the mean value of l, and O(lnt) refers to the order of magnitude of lnt. As a result, the regret

[0170] In this application, in the CMAB process, the maximum likelihood detection method is used to determine the power angle spectrum (PAS), and the angle range is iteratively reduced. It is proved that this method can achieve results that grow logarithmically with time. The simulation results also verify the good performance of this scheme.

[0171] The method in the embodiment of the present invention will be described below in combination with specific embodiments for simulation verification.

[0172] In this simulation example, we prove the efficiency of the proposed CMAB-based scheme through simulation. The number of antennas is 128. The central angles of the users are uniformly distributed in . The changes in the central angles of all users are in [0°, 5°]. The AS of each user is uniformly distributed in [15°, 30°]. In the CMAB algorithm, ε is set to 0.01, and ζ is set to In the UCB value, Is set to 1, and κ is set to 1.5. We compare the CMAB scheme with the ACS and N-JSDM schemes. In the ACS and N-JSDM schemes, the CCM is obtained through 500 time slots, and the TB includes 1000 time slots, which corresponds to 5 meters at a speed of 10 meters per second in the LTE system. With T c The effective spectral efficiency (ESE) R of user k in the coherent block with symbols k Is given by:

[0173]

[0174] Where τ represents the DTL of user k, and T c Is set to 100, and SINR k Is the signal-to-noise interference ratio of user k.

[0175] Figure 1 The ESEs of different schemes under different SNRs are compared. The number of users is 10 and 16 respectively. Δ is set to 15°. The CMAB-based scheme achieves a higher ESE than the NJSDM and ACS schemes. This is because ACS and NJSDM spend many time slots to estimate the CCM. For 10 users, the CMAB, NJSDM, and ACS schemes are respectively lower than the CMAB, NJSDM, and ACS schemes for 16 users. Figure 2 The ESEs of different schemes under different numbers of users are shown. The SNR is 10 dB. Δ is 15°. Our CGTSB scheme has a higher ESE than other schemes. By increasing the number of users, the ESE of the CGTSB scheme first increases and then decreases. The reasons are as follows.

[0176] For a small number of users, the interference between users is not large, and the multi-user gain improves the ESE. However, when the number of users is large, the interference between users is high, thus reducing the ESE. Figure 3 The ESEs of different schemes under different DTLs are compared. The SNR is 10 dB. Δ is 15°. CMAB has higher performance than ACS. The ESEs of the CMAB and ACS schemes first increase and then decrease. This is because, for small DTL, the influence of DTL decreases. By increasing DTL, the number of DFT vectors will increase, thus improving the spectral efficiency. Therefore, the ESE increases. However, when DTL is large, DTL will be important. By increasing DTL, the spectral efficiency will not increase, so the ESE will decrease. The ESE of the NJSDM scheme remains unchanged because NJSDM is independent of DTL.

[0177] Figure 4 The ESE of the CMAB scheme in each time slot is shown. The signal-to-noise ratio is 10 dB and the DTL is 30. Due to the characteristics of the UCB strategy, the expected ESE of the CMAB scheme first increases and then converges as the number of time slots increases. When Δ is large, it will take more time slots to exploit the actions, so the ESE convergence speed is slower.

[0178] The present invention transforms the pre-beamforming matrix design problem into a CMAB problem. The pre-beamforming matrix design within each time slot is regarded as an arm selection problem in CMAB, and then the arm selection problem is transformed into a 0-1 integer linear programming problem and solved by the branch and bound method. During the training process, the maximum likelihood method is used to detect the power of the angle spectrum and adaptively determine the angle range of the users. The present invention proves that the regret grows logarithmically with time, enabling the proposed scheme to converge to the optimal action, and the present invention can significantly improve the spectral efficiency compared with the previous methods and can significantly reduce the channel estimation overhead and improve the spectral efficiency without the need for CCM.

[0179] As is known by common technical knowledge, the present invention can be implemented by other embodiments that do not depart from its spiritual essence or essential features. Therefore, the above-disclosed embodiments are illustrative in all aspects and not exclusive. All changes within the scope of the present invention or within the scope equivalent to the present invention are encompassed by the present invention.

Claims

1. A two-stage beamforming method based on a combinatorial multi-armed bandit, characterized in that, Including: Adopting a combinatorial multi-armed bandit design for the pre-beamforming matrix; Transmitting a pilot matrix with a cyclic structure and receiving the equivalent channel matrix estimated by the user based on the pilot matrix; Designing a precoder and determining the angular spectrum of each user according to the equivalent channel matrix; Correcting the pre-beamforming matrix using the precoder and the angular spectrum of each user; Adopting a combinatorial multi-armed bandit design for the pre-beamforming matrix includes: Converting the design problem of the pre-beamforming matrix into an arm selection problem in the combinatorial multi-armed bandit; Converting the arm selection problem in the combinatorial multi-armed bandit into a 0-1 integer linear programming problem; Converting the design problem of the pre-beamforming matrix into an arm selection problem in the combinatorial multi-armed bandit includes: The pre-beamforming matrix consists of the DFT matrix D = d1, d2, …, d M which is composed of; Among them, is the basic arm, i = 1, 2, 3, … Vector is the direction vector, θ is the incident angle, M is the number of antennas, λ c is the wavelength, d is the antenna spacing, T is the matrix transpose symbol, and j is the imaginary part symbol; Define the instantaneous channel state information interval as a time block. Assume that the angular range (AR) of user k at the th time block is and which are the lower and upper angular limits of user k respectively. In the th time block, set a large constant Δ such that the AR of user k belongs to Let the set be the set of i, is the channel vector at the t-th time slot in the t-th time block As the number of antennas approaches infinity, we have: where d l is the l-th column of the DFT matrix D; Use to represent the selected DFT vector set, and use to indicate the pre-beamforming matrix, we have: Among them, ∩ refers to the intersection symbol of sets, denotes and 's intersection. The matrix is composed of the columns d l , of the DFT matrix; Let the circular structure of the pilot matrix be where denotes the number of elements in the set , denotes the set of complex matrices of dimension , L is the downlink training length, [x1, x2, …, x L , x1, x2, …] denotes the matrix formed by arranging the vectors x1, x2, …, x L , x1, x2, … in sequence, L ≥ max k d k , max k d k denotes the maximum value among all d k , d k is the number of elements in, and the matrix [x1, x2, …, x L is a unitary matrix. The received signal of user k in channel estimation is: where n H (t) represents the noise at the t-th moment, and H represents the conjugate transpose; substituting (4) into (5) and removing the zero terms, we get: Among them, is a high unitary matrix composed of partial vectors x i ; Denoted by the symbol in equation (6), where t refers to the t-th time slot and is estimated using the least squares criterion; according to the least squares criterion, multiply equation (6) by to obtain: while estimated as and fed back to the base station. Using and formula (4), the base station obtains the estimated value of, and represents the estimated value of as where t is the t-th time slot; In the case of a limited downlink training length, design to maximize the spectral efficiency and maximize the received signal energy, that is to obtain high spectral efficiency; considering Equation (4), let the number of elements in be d k , and set the downlink training length to κmax k d k , where κ is a constant, and the pre-beamforming matrix design problem is as follows: s.t.max k d k <τ (8) Among them, DTL is the pilot length; in multiple DFT vectors are required, so it is the arm selection problem in the combinatorial multi-armed bandit; Converting the arm selection problem in the combinatorial multi-armed bandit into a 0-1 integer linear programming problem includes: M basic arms d1, d2, …, d M are labeled from 1 to M. Let be the reward of the i-th basic arm. The super arm consists of basic arms a i , i = 1, 2, ···, where a i ∈{1, 2, ··· M}. Given a super arm A a(t) , the pre-beamforming matrix is and the reward is the power of the received signal given by the following formula: Using the linear UCB strategy to select the super-arm in each time slot. Considering the chi-square distribution of the rewards, the UCB value of the basic arm in the t-th time slot is defined as: Among them, is an arbitrary positive constant, m i,t is the number of times action i is selected. Then the UCB value of the super-arm is Prove that it converges to the optimal action of the latter; In each time slot, use \(x\) to indicate whether the base arm is selected in the super arm, where \(x\) i = 1 indicates that the \(i\)-th base arm is selected, and \(x\) i = 0 indicates that the \(i\)-th base arm is not selected; let \(u = [u_1(t), u_2(t)\cdots]\), and define the matrix \(A\), where the \(l\)-th element of its \(k\)-th row \(a\) k is: where is the element in the k-th row and l-th column, then the problem of maximizing the UCB value is: s.t. Ax ≤ τe(15) where e is a vector with all elements being 1, and τ is a given value used to describe the pilot length; When decreases to convergence for all users, is regarded as the actual AR, and the minimum downlink training length gives the maximum solution; the actions are divided into two groups, one is the candidate set involving the edge vectors, and the other is the remaining set containing the remaining vectors. x is divided into x1 and x2, where x1 represents the actions in the candidate set and x2 represents the remaining actions; A is divided into A1 and A2, that is, A1 is the set of vectors i where a is the i-th column of matrix A, and A2 is the set of vectors ; the vector u is divided into u1 and u2, that is, u1 is the set of vectors u, and u2 is the set of vectors u. Then it is divided into two sub-problems. The first method is to select the basic arm in to maximize the UCB value: s.t. A1x1 ≤ τe(17) Among them, without considering the arm in Once we get x1, we replace x1 with Then consider the basic arm in The formula for the second sub-problem is: s.t. A2x2 ≤ τe - A2x1(19) and is a linear 0-1 integer programming problem.

2. The two-stage beamforming method based on a combinatorial multi-armed bandit according to claim 1, wherein, Also including: Updating the reward of each basic arm in each time slot; Calculating the power of each basic arm in the candidate set using the likelihood ratio. If the basic arm in the candidate set has zero or non-zero power, update the marginal vector; The expected regret of the CMAB algorithm is O(lnt).

3. The two-stage beamforming method based on a combinatorial multi-armed bandit according to claim 1, wherein Solving the 0-1 integer linear programming problem using the branch and bound method.

4. The two-stage beamforming method based on a combinatorial multi-armed bandit according to claim 1, wherein Also including detecting the power of the angular spectrum using the maximum likelihood method.