Massive MIMO terminal scheduling method based on combined multi-arm bandit

CN116367318BActive Publication Date: 2026-09-08NARI NANJING CONTROL SYSTEM CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211731963.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2026-09-08
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

然而,这些方法均需要获取信道协方差矩阵,在环境变化比较快的场景中,信道协方差矩阵的变化也较快,此时获取实时的信道协方差矩阵将十分困难

Benefits of technology

[0060] This invention transforms the terminal scheduling problem in a massive MIMO system into a CMAB (Content Management Algorithm) problem. During the selection of the superarm, a confidence upper bound (UCB) value is defined, and a greedy algorithm with a linear confidence upper bound (UCB) is proposed, enabling rapid selection of terminals with high spectral efficiency. Furthermore, after the terminal sends orthogonal pilot signals to the base station, this invention employs least squares for channel estimation and MMSE (Multi-Level Multi-Screen Encoding) for precoding, effectively improving the system's spectral efficiency and reducing pilot overhead in massive MIMO.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116367318B_ABST
    Figure CN116367318B_ABST
Patent Text Reader

Abstract

The application discloses a large-scale MIMO system terminal scheduling method based on a combined multi-arm gambling machine, which comprises the following steps: acquiring channel matrix information of a MIMO system, taking each terminal for selection as a basic arm, and constructing a combined multi-arm gambling machine problem model corresponding to a terminal scheduling problem; solving the constructed CMAB problem model by using a linear confidence upper bound UCB method, obtaining a basic arm reward and a plurality of terminals corresponding to a maximum super arm, and taking the terminals as terminals needing to be scheduled; acquiring an orthogonal pilot signal sent by the terminals needing to be scheduled, performing channel estimation based on the orthogonal pilot signal; and detecting a terminal signal based on a result of the channel estimation. The terminal scheduling problem is modeled as a combined multi-arm gambling machine problem, a terminal with high spectral efficiency can be quickly selected, better system spectral efficiency can be obtained in communication, and a large-scale MIMO system with high spectral efficiency is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication technology, and in particular to a terminal scheduling method for a large-scale MIMO system based on a combined multi-armed gambling machine. Background Technology

[0002] Massive MIMO has become a key technology in 5G mobile communication systems. This technology significantly improves system capacity by configuring a large number of antennas at the base station and / or terminal. To fully utilize the gains from massive MIMO, precoding of each terminal at the base station is typically required during downlink communication. This precoding necessitates the base station acquiring effective Channel State Information (CSI). In Time Division Duplex (TDD) systems, due to the reciprocity between uplink and downlink channels, the base station can obtain downlink CSI through uplink pilots and channel estimation. The system pilot length is proportional to the number of serving terminals. When the number of terminals is large, pilot and feedback overhead becomes significant. Furthermore, interference between terminals leads to poor system performance, thus requiring terminal scheduling to improve system performance.

[0003] Compared to instantaneous CSI, statistical CSI changes more slowly, making it easier to obtain. Directly using statistical CSI for terminal scheduling can reduce system overhead and improve system performance. However, all these methods require obtaining the channel covariance matrix. In scenarios with rapidly changing environments, the channel covariance matrix also changes rapidly, making it very difficult to obtain the real-time channel covariance matrix. Summary of the Invention

[0004] The purpose of this invention is to provide a terminal scheduling method for large-scale MIMO systems based on a combined multi-armed gambling machine. By modeling the terminal scheduling problem as a combined multi-armed gambling machine problem, terminals with high spectral efficiency can be selected quickly, thereby achieving better system spectral efficiency in communication. The technical solution adopted by this invention is as follows.

[0005] On one hand, the present invention provides a terminal scheduling method for a large-scale MIMO system based on a combined multi-armed gambling machine, comprising:

[0006] Obtain the channel matrix information of the MIMO system;

[0007] Based on the channel matrix information and the pre-built terminal scheduling problem model, each selectable terminal is taken as the basic arm, and all selected terminals are taken as the super arms. The expression of the basic arm reward is determined according to the channel state information of each terminal in each time slot, and a combined multi-armed gambling machine (CMAB) problem model corresponding to the terminal scheduling problem is constructed.

[0008] The constructed CMAB problem model is solved using the linear confidence upper bound UCB method to obtain the base arm reward and the multiple terminals corresponding to the largest superarm, which are then used as the terminals to be scheduled.

[0009] Obtain the orthogonal pilot signal sent by the terminal that needs to be scheduled, and perform channel estimation based on the orthogonal pilot signal;

[0010] The terminal signal is detected based on the channel estimation results.

[0011] Optionally, the pre-built terminal scheduling problem model is represented as:

[0012]

[0013] In the formula, Let represent the equivalent channel energy of U terminals selected from all terminals at time t. Let U represent the channel matrix of the U terminals at time t, where the channel corresponding to any terminal k is represented as:

[0014]

[0015] Wherein, ρ(θ) l,k ) indicates that the angle of incidence is θ l,k The path complex gain, ρ(θ) l.k )~CN(0,σ l,k I); Let a(θ) be the direction vector. l The i-th element in ) is L represents the number of channel paths, σ represents the complex gain variance, and I represents the identity matrix.

[0016] Optionally, the construction of the combined multi-armed gambling machine (CMAB) problem model corresponding to the terminal scheduling problem includes:

[0017] Let all K terminals be denoted as the K basic arms of the CMAB problem, and let h be the channel state information of terminal k in time slot t. k,t ,Will Record this as a reward for the basic arm;

[0018] In each time slot, U terminals are selected for communication, and the set of U terminals is denoted as superarm A. t The reward and sum of the superarm are then represented as:

[0019]

[0020] Solving the corresponding terminal scheduling problem is the combinatorial multi-armed gambling machine (CMAB) problem, which involves selecting terminals that communicate with the base station in each time slot such that the sum of the rewards of the combined arms formed by the selected terminals is maximized. maximum.

[0021] Optionally, solving the constructed CMAB problem model using the linear confidence upper bound UCB method includes:

[0022] Determine the distribution of reward values: the reward value for terminal k is Where, |ρ(θ) l,k )| 2 It is a chi-square distribution with 1 degree of freedom, and its parameter is... If the exponential distribution is followed, then It is a parameter of The sub-exponential distribution;

[0023] Determine the UCB value for each base arm as follows:

[0024]

[0025] in, m i,t Select the number of times the basic arm i has performed its action at time t. The average reward of action i at time t; considering the finite energy of each path, here we will... Normalization Then at this time

[0026] The superarm A is determined based on the UCB values ​​of each base arm. t The UCB value is:

[0027] Superarm A will be selected in each time slot. t The included basic arms allow the superarm's reward and biggest problem to be transformed into an objective function:

[0028] Superarm A is obtained by sorting and selecting U base arms with the largest UCB value. t The terminals corresponding to each basic arm are designated as the terminals that need to be scheduled.

[0029] Optionally, the channel estimation based on orthogonal pilot signals includes:

[0030] Based on the orthogonal pilot signals X sent to the base station by each selected terminal, the signal Y(t) received at the base station in the t-th time slot is expressed as:

[0031]

[0032] in, Let N(t) represent the channel matrix of the selected U terminals at time t, and let N(t) be the noise of the t-th time slot, which is additive white Gaussian noise.

[0033] The uplink channel is estimated using the channel trajectory method based on the zero-forcing criterion, and the channel estimation results are as follows:

[0034]

[0035] Optionally, the detection of the terminal signal based on the channel estimation results adopts the minimum mean square error (MMSE) method, and the equalization matrix is ​​represented as follows:

[0036]

[0037] Where ρ is the signal-to-noise ratio and I is the identity matrix;

[0038] According to the equalization matrix, the signal transmitted by the terminal is W. MMSE Y(t).

[0039] Secondly, the present invention provides a terminal scheduling device for a large-scale MIMO system based on a combined multi-armed gambling machine, comprising:

[0040] The channel matrix information acquisition module is configured to acquire the channel matrix information of the MIMO system.

[0041] The CMAB problem model construction module is configured to, based on the channel matrix information and the pre-built terminal scheduling problem model, take each selectable terminal as the base arm and all selected terminals as the super arm, determine the expression of the base arm reward according to the channel state information of each terminal in each time slot, and construct a combined multi-armed gambling machine CMAB problem model corresponding to the terminal scheduling problem.

[0042] The optimization solution module is configured to solve the constructed CMAB problem model using the linear confidence upper bound UCB method, obtain the base arm reward and the multiple terminals corresponding to the largest superarm, and use them as the terminals to be scheduled.

[0043] The channel estimation module is configured to acquire the orthogonal pilot signal emitted by the terminal to be scheduled, and perform channel estimation based on the orthogonal pilot signal;

[0044] Additionally, a terminal signal detection module is configured to detect the terminal signal based on the channel estimation result.

[0045] Optionally, the CMAB problem model building module constructs a combined multi-armed gambling machine CMAB problem model corresponding to the terminal scheduling problem, including:

[0046] Let all K terminals be denoted as the K basic arms of the CMAB problem, and let h be the channel state information of terminal k in time slot t. k,t ,Will Record this as a reward for the basic arm;

[0047] In each time slot, U terminals are selected for communication, and the set of U terminals is denoted as superarm A. t The reward and sum of the superarm are then represented as:

[0048]

[0049] Solving the corresponding terminal scheduling problem is the combinatorial multi-armed gambling machine (CMAB) problem, which involves selecting terminals that communicate with the base station in each time slot such that the sum of the rewards of the combined arms formed by the selected terminals is maximized. maximum.

[0050] Optionally, the optimization solution module uses the linear confidence upper bound UCB method to solve the constructed CMAB problem model, including:

[0051] Determine the distribution of reward values: the reward value for terminal k is Where, |ρ(θ) l,k )| 2 It is a chi-square distribution with 1 degree of freedom, and its parameter is... If the exponential distribution is followed, then It is a parameter of The sub-exponential distribution;

[0052] Determine the UCB value for each base arm as follows:

[0053]

[0054] in, m i,t Select the number of times the basic arm i has performed its action at time t. The average reward of action i at time t;

[0055] The superarm A is determined based on the UCB values ​​of each base arm. t The UCB value is:

[0056] Superarm A will be selected in each time slot. t The included basic arms allow the superarm's reward and biggest problem to be transformed into an objective function:

[0057] Superarm A is obtained by sorting and selecting U base arms with the largest UCB value. t The terminals corresponding to each basic arm are designated as the terminals that need to be scheduled.

[0058] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the terminal scheduling method for a large-scale MIMO system based on a combined multi-armed gambling machine as described in the first aspect.

[0059] Beneficial effects

[0060] This invention transforms the terminal scheduling problem in a massive MIMO system into a CMAB (Content Management Algorithm) problem. During the selection of the superarm, a confidence upper bound (UCB) value is defined, and a greedy algorithm with a linear confidence upper bound (UCB) is proposed, enabling rapid selection of terminals with high spectral efficiency. Furthermore, after the terminal sends orthogonal pilot signals to the base station, this invention employs least squares for channel estimation and MMSE (Multi-Level Multi-Screen Encoding) for precoding, effectively improving the system's spectral efficiency and reducing pilot overhead in massive MIMO. Attached Figure Description

[0061] Figure 1 The diagram shown is a flowchart of the method of the present invention;

[0062] Figure 2 The diagram shows a performance comparison of the system's effective spectral efficiency as a function of the signal-to-noise ratio (SNR).

[0063] Figure 3 The diagram shows a performance comparison of the system's effective spectral efficiency as a function of the number of terminals. Detailed Implementation

[0064] The following description, in conjunction with the accompanying drawings and specific embodiments, provides further details.

[0065] Example 1

[0066] This embodiment introduces a terminal scheduling method for a large-scale MIMO system based on a combined multi-armed gambling machine, which includes the following steps:

[0067] Obtain the channel matrix information of the MIMO system;

[0068] Based on the channel matrix information and the pre-built terminal scheduling problem model, each selectable terminal is taken as the basic arm, and all selected terminals are taken as the super arms. The expression of the basic arm reward is determined according to the channel state information of each terminal in each time slot, and a combined multi-armed gambling machine (CMAB) problem model corresponding to the terminal scheduling problem is constructed.

[0069] The constructed CMAB problem model is solved using the linear confidence upper bound UCB method to obtain the base arm reward and the multiple terminals corresponding to the largest superarm, which are then used as the terminals to be scheduled.

[0070] Obtain the orthogonal pilot signal sent by the terminal that needs to be scheduled, and perform channel estimation based on the orthogonal pilot signal;

[0071] The terminal signal is detected based on the channel estimation results.

[0072] refer to Figure 1 The flowchart shown below illustrates the implementation process of this embodiment.

[0073] I. Obtain the channel matrix of the MIMO system and model the large-scale MIMO terminal scheduling problem.

[0074] This section first introduces the communication model of a large-scale MIMO system, and then performs mathematical modeling of the terminal scheduling problem.

[0075] 1.1 Multi-terminal communication model for large-scale MIMO systems

[0076] Consider a single-cell, multi-terminal massive MIMO system where one base station serves K single-antenna terminals. The base station is configured with a linear array of M antennas. Let the channel from the base station to terminal k in time slot t be denoted as... The signal sent to terminal k is s k,t Then the received signal y of terminal k k,t for

[0077]

[0078] in, The precoding vector of terminal j in time slot t, n k,t This refers to the noise at terminal k. Here, we assume n... k,t Obey n k,t ~CN(0,σ).

[0079] Consider a commonly used Saleh Valenzuela massive MIMO channel model, in which the channel between terminal k and the base station can be described as...

[0080]

[0081] Wherein, ρ(θ) l,k () is the incident angle θ l,k Path complex gain, Let a(θ) be the direction vector. l The i-th element in ) is ρ(θ l.k )~CN(0,σ l,k I).

[0082] The statistical CSI changes with the terminal's position. When the terminal moves slowly, the statistical CSI will remain consistent for a relatively long period of time. However, in high-speed moving scenarios, the statistical CSI changes rapidly, making it difficult to obtain parameters such as angle and angle power spectrum.

[0083] 1.2 Mathematical modeling of the terminal scheduling problem.

[0084] The scheduling problem addressed in this invention is to select U terminals from K terminals for communication in each time slot, aiming to achieve a high transmission rate for the selected terminals. Since the transmission rate of a terminal is related to its receiving power, to improve the system's transmission rate, the channel power of the selected U terminals must be maximized. Assume... Let U be the channel matrix of the selected U terminals at time t. Then, when the selected terminals have equivalent channel energy... A larger capacity can increase the system's transmission rate. Because... At this point, the terminal scheduling problem in a large-scale MIMO system can be described as follows:

[0085]

[0086] II. Constructing a combined multi-armed gambling machine (CMAB) problem model corresponding to the terminal scheduling problem.

[0087] This part of the task is to transform the terminal scheduling problem into a combined multi-armed gambling machine problem. Specifically, all K terminals in the system are denoted as the K basic arms of the CMAB problem, labeled 1, 2, ..., K. Each terminal (basic arm) corresponds to a channel, and the channel state information of terminal k in time slot t is denoted as h. k,t ,Will Record this as a reward for the basic arm.

[0088] During terminal scheduling, U terminals are selected for communication in each time slot. These U terminals are grouped into a combined arm, and the set of U terminals is denoted as superarm A. t The reward for the combined arm at this time is:

[0089]

[0090] To ensure the equivalent channel energy of the selected terminal The problem of converting it into a combined multi-armed gambling machine (CMAB) is as follows: Select the terminal that communicates with the base station in each time slot so that the superarm A... t Total reward and maximum.

[0091] 3. Using an action selection algorithm based on the linear confidence upper bound (UCB) method, select the base arm to obtain the set of terminals to be scheduled.

[0092] This embodiment employs an action selection algorithm based on the linear confidence upper bound (UCB) method to solve the combined arm optimization problem, which can accelerate convergence and improve system performance. Details are as follows.

[0093] 3.1 Analysis of reward value distribution and determination of UCB value.

[0094] Based on the foregoing introduction, in the SV channel model, the reward value for terminal k is:

[0095]

[0096] Due to ρ(θ) l )~CN(0,σ l I), |ρ(θ) l )| 2 It is a chi-square distribution with 1 degree of freedom, and its parameter is... The sub-exponential distribution. From the concept of sub-exponential, we can see that... It is a parameter of The exponential distribution is as follows. Since the energy of each path is finite, we will... Normalization

[0097] Considering that the UCB strategy can provide better performance, this embodiment adopts the UCB strategy to implement the basic arm selection.

[0098] Considering that the reward of each arm follows a sub-exponential distribution, based on the concentration inequality, the UCB value of each arm is defined as follows:

[0099]

[0100] in m i,t Let t be the number of times action i has been executed. The average reward for action i at time t.

[0101] because at this time

[0102] 3.2 Obtain terminal scheduling results using the linear UCB algorithm based on the base arm UCB value.

[0103] Based on the definition of the UCB value of the basic arm, superarm A t The UCB value is:

[0104]

[0105] In the t-th time slot, select set A t To maximize the reward, the problem can be described as follows:

[0106]

[0107] For this problem, since there are no constraints, to maximize the objective function, we only need to sort the rewards and then select the U largest UCB values. The base arm set A... t The choice of which can be made by referring to the following theorem and proof process.

[0108] Theorem 2: The average regret value of the CMAB method is O(lnt).

[0109] Proof: Define variable T i,t In time slot t, if the optimal combination arm a is selected... * So T i,t Nothing changes. If a non-optimal combination arm 'a' is chosen, then T... i,t The value is incremented by 1, where i = min j∈a m j,t Therefore, in time slot t, the number of times non-optimal actions are selected is... Define variable Q i,t If T i,t Then add 1 to the value of Q. i,t =1; if T i,t If the value of Q remains unchanged, then i,t Let l be any positive integer, then 0.

[0110]

[0111] Where 1{a} is an indicator function. When a is a true event, 1{a} = 1; when a is a false event, 1{a} = 0. i,p =1, a non-optimal action was selected, and:

[0112]

[0113] in:

[0114]

[0115] make Q p =|A a(p+1) |. Since l≤T i,p ≤m i,p ,

[0116]

[0117] This means that at least one of the following three events must be true:

[0118]

[0119]

[0120]

[0121] in, R i Let be the mean reward for action i. For event ε1, we have:

[0122]

[0123] Due to rewards Obtain the parameter as The exponential distribution, in which Using Hofding's inequality, we have:

[0124]

[0125] As can be seen from the above, So because achievable also So From this we can obtain Similarly, let set up

[0126]

[0127] So:

[0128]

[0129] And there are:

[0130]

[0131] Therefore, it can be concluded that At the same time thus

[0132]

[0133] Therefore, P{ε3} = 0. Note that...

[0134]

[0135] therefore So Where Δ max =max a Δ a .

[0136] IV. Channel estimation using the least squares (LS) method

[0137] After obtaining the scheduled terminals, channel estimation and precoding design are performed. Assume the selected terminal set is A. t At this time, each terminal sends an orthogonal pilot signal X to the base station. Then, the signal Y(t) received at the base station in the t-th time slot is...

[0138]

[0139] Here, N(t) represents the noise in the t-th time slot, which is additive white Gaussian noise. Therefore, the uplink channel can be estimated using the zero-forcing criterion channel estimation method, i.e., multiplying the signal by an orthogonal matrix X. H At this point, we can obtain:

[0140]

[0141] Then the channel estimation result is Y(t)X H .

[0142] V. Detecting terminal signals using the minimum mean square error (MMSE) method.

[0143] After obtaining the channel estimate, applying MMSE detection to the terminals can reduce inter-terminal interference and improve system spectral efficiency. When using MMSE, the equalization matrix is ​​designed as follows:

[0144]

[0145] Where ρ is the signal-to-noise ratio and I is the identity matrix. Then the signal transmitted by the terminal is W. MMSE Y(t).

[0146] Effect verification

[0147] Taking the extended Saleh Valenzuela model as an example, in this model, the channel matrix can be described as clusters and paths, where each cluster contains multiple transmission paths that propagate along the cluster. The channel matrix h between the terminal and the base station... k It can be represented as:

[0148]

[0149] Assuming a uniform linear antenna array (ULA) is used in the simulation, the antenna array response vector can be defined as:

[0150]

[0151] Where d is the distance between adjacent antenna elements, and λ is the wavelength of the millimeter wave.

[0152] The simulation was performed using Matlab. It is assumed that there are 100 high-speed mobile terminals in the cell. The base station is equipped with 128 antennas. Due to the high-speed movement, the coherence time length for statistical CSI is set to T = 300 time slots (at a speed of 100 km / h, 300 time slots correspond to a range of 4.16 meters). After 300 time slots, the terminal's statistical CSI changes.

[0153] The method of this invention is compared with a method based on statistical CSI, and it is assumed that the statistical CSI is estimated using conventional methods in the first coherent time, and the precoding design is performed using the method of this invention in subsequent coherent times.

[0154] The statistical CSI-based method performs precoding based on the statistical CSI obtained at the first coherent time. During subsequent movements, the same statistical CSI is used for precoding design. Although this process does not require additional pilot feedback overhead, the statistical CSI changes due to terminal movement, resulting in a loss of precoding performance.

[0155] Figure 2 This paper describes a comparison between the system spectral efficiency as a function of SNR using the method of this embodiment and the traditional method. Figure 3 This paper compares the effective spectral efficiency of the system with the number of users in the method of this invention and traditional methods. It can be seen that the MAB-based precoding method of this invention has higher spectral efficiency than the statistical CSI-based precoding method. This is because using the estimated CCM for precoding design consumes more time slots for CCM estimation, resulting in poorer average spectral efficiency for CCM-based methods. The method of this invention avoids estimating the terminal CCM, effectively reducing system pilot overhead and improving system spectral efficiency. Furthermore, it can be observed that, within a certain range, the system exhibits higher spectral efficiency as the number of terminals increases when using the method of this invention.

[0156] Example 2

[0157] Based on the same inventive concept as Embodiment 1, this embodiment introduces a terminal scheduling device for a large-scale MIMO system based on a combined multi-armed gambling machine, which includes:

[0158] The channel matrix information acquisition module is configured to acquire the channel matrix information of the MIMO system.

[0159] The CMAB problem model construction module is configured to, based on the channel matrix information and the pre-built terminal scheduling problem model, take each selectable terminal as the base arm and all selected terminals as the super arm, determine the expression of the base arm reward according to the channel state information of each terminal in each time slot, and construct a combined multi-armed gambling machine CMAB problem model corresponding to the terminal scheduling problem.

[0160] The optimization solution module is configured to solve the constructed CMAB problem model using the linear confidence upper bound UCB method, obtain the base arm reward and the multiple terminals corresponding to the largest superarm, and use them as the terminals to be scheduled.

[0161] The channel estimation module is configured to acquire the orthogonal pilot signal emitted by the terminal to be scheduled, and perform channel estimation based on the orthogonal pilot signal;

[0162] Additionally, a terminal signal detection module is configured to detect the terminal signal based on the channel estimation result.

[0163] The specific implementation of each of the above functional modules is described in Example 1.

[0164] Example 3

[0165] This embodiment incorporates a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the terminal scheduling method for a large-scale MIMO system based on a combined multi-armed gambling machine as described in Embodiment 1.

[0166] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0167] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0168] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0169] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0170] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A terminal scheduling method for a large-scale MIMO system based on a combined multi-armed gambling machine, characterized in that, include: Obtain the channel matrix information of the MIMO system; Based on the channel matrix information and the pre-built terminal scheduling problem model, each selectable terminal is taken as the basic arm, and all selected terminals are taken as the super arms. The expression of the basic arm reward is determined according to the channel state information of each terminal in each time slot, and a combined multi-armed gambling machine (CMAB) problem model corresponding to the terminal scheduling problem is constructed. The constructed CMAB problem model is solved using the linear confidence upper bound UCB method to obtain the base arm reward and the multiple terminals corresponding to the largest superarm, which are then used as the terminals to be scheduled. Obtain the orthogonal pilot signal sent by the terminal that needs to be scheduled, and perform channel estimation based on the orthogonal pilot signal; The terminal signal is detected based on the channel estimation results.

2. The method according to claim 1, characterized in that, The pre-built terminal scheduling problem model is represented as follows: In the formula, = This indicates that the device was selected from all terminals. Each terminal The equivalent channel energy at time t. Indicates the Each terminal The channel matrix at time t, where any terminal The corresponding channel representation is as follows: in, Indicates the angle of incidence as Path complex gain, ; It is a direction vector. M This indicates the number of linear arrays of antennas configured in the base station. The Middle The elements are , Indicates the number of channel paths. Represents the variance of the complex gain. Represents the identity matrix.

3. The method according to claim 2, characterized in that, The construction of the combined multi-armed gambling machine (CMAB) problem model corresponding to the terminal scheduling problem includes: All of The problem is denoted as CMAB. A basic arm, a terminal exist The channel state information corresponding to the time slot is ,Will Record this as a reward for the basic arm; In each time slot, select Each terminal communicates, and records The collection of terminals constitutes a superarm. The reward and sum of the superarm are then represented as: Solving the corresponding terminal scheduling problem is the combinatorial multi-armed gambling machine (CMAB) problem, which involves selecting terminals that communicate with the base station in each time slot such that the sum of the rewards of the combined arms formed by the selected terminals is maximized. maximum.

4. The method according to claim 3, characterized in that, The method of solving the constructed CMAB problem model using the upper bound of linear confidence (UCB) includes: Determine the distribution of reward values: Terminal The reward value is ,in, It is a chi-square distribution with 1 degree of freedom, and its parameter is... If the exponential distribution is followed, then It is a parameter of The sub-exponential distribution; Determine the UCB value for each base arm as follows: in, , Select the base arm at time t The number of times the action has been performed. Action at time t The average reward M This indicates the number of linear arrays of antennas configured in the base station; The superarm is determined based on the UCB value of each base arm. The UCB value is: ; Superarms will be selected in each time slot. The included basic arms allow the superarm's reward and biggest problem to be transformed into an objective function: ; Select by sorting A superarm is obtained from a base arm with the highest UCB value. The terminals corresponding to each basic arm are designated as the terminals to be scheduled.

5. The method according to claim 1, characterized in that, The channel estimation based on orthogonal pilot signals includes: Based on the orthogonal pilot signals sent to the base station by each selected terminal In the Signal received at each time slot base station Represented as: in, Indicates the selected Each terminal Channel matrix at time 10:00 For the first The noise in each time slot is additive white Gaussian noise; The uplink channel is estimated using the least squares (LS) method, and the channel estimation result is as follows: 。 6. The method according to claim 2, characterized in that, The detection of terminal signals based on channel estimation results employs the minimum mean square error (MMSE) method, and the equalization matrix is ​​represented as follows: in For signal-to-noise ratio, It is the identity matrix; According to the equalization matrix, the signal sent by the terminal is .

7. A terminal scheduling device for a large-scale MIMO system based on a combined multi-armed gambling machine, characterized in that, include: The channel matrix information acquisition module is configured to acquire the channel matrix information of the MIMO system. The CMAB problem model construction module is configured to, based on the channel matrix information and the pre-built terminal scheduling problem model, take each selectable terminal as the base arm and all selected terminals as the super arm, determine the expression of the base arm reward according to the channel state information of each terminal in each time slot, and construct a combined multi-armed gambling machine CMAB problem model corresponding to the terminal scheduling problem. The optimization solution module is configured to solve the constructed CMAB problem model using the linear confidence upper bound UCB method, obtain the base arm reward and the multiple terminals corresponding to the largest superarm, and use them as the terminals to be scheduled. The channel estimation module is configured to acquire the orthogonal pilot signal emitted by the terminal to be scheduled, and perform channel estimation based on the orthogonal pilot signal; Additionally, a terminal signal detection module is configured to detect the terminal signal based on the channel estimation result.

8. The terminal scheduling device for a large-scale MIMO system based on a combined multi-armed gambling machine according to claim 7, characterized in that, The CMAB problem model building module constructs a combined multi-armed gambling machine CMAB problem model corresponding to the terminal scheduling problem, including: All of The problem is denoted as CMAB. A basic arm, a terminal exist The channel state information corresponding to the time slot is ,Will Record this as a reward for the basic arm; In each time slot, select Each terminal communicates, and records The collection of terminals constitutes a superarm. The reward and sum of the superarm are then represented as: Solving the corresponding terminal scheduling problem is the combinatorial multi-armed gambling machine (CMAB) problem, which involves selecting terminals that communicate with the base station in each time slot such that the sum of the rewards of the combined arms formed by the selected terminals is maximized. maximum.

9. The terminal scheduling device for a large-scale MIMO system based on a combined multi-armed gambling machine according to claim 8, characterized in that, The optimization solution module uses the linear confidence upper bound UCB method to solve the constructed CMAB problem model, including: Determine the distribution of reward values: Terminal The reward value is ,in, It is a chi-square distribution with 1 degree of freedom, and its parameter is... If the exponential distribution is followed, then It is a parameter of The sub-exponential distribution; Determine the UCB value for each base arm as follows: in, , Select the base arm at time t The number of times the action has been performed. Action at time t The average reward M This indicates the number of linear arrays of antennas configured in the base station; The superarm is determined based on the UCB value of each base arm. The UCB value is: ; Superarms will be selected in each time slot. The included basic arms allow the superarm's reward and biggest problem to be transformed into an objective function: ; Select by sorting A superarm is obtained from a base arm with the highest UCB value. The terminals corresponding to each basic arm are designated as the terminals to be scheduled.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the terminal scheduling method for a large-scale MIMO system based on a combined multi-armed gambling machine as described in any one of claims 1-6.