A 5G multi-base station antenna weight joint optimization method and device

By using deep Q network and UCB exploration methods to jointly find antenna weights in the optimization model of multiple 5G base stations, multiple 5G base station beam coordination optimization problems have been solved, and beam coordination optimization effects with high performance and low overhead have been achieved.

CN115190510BActive Publication Date: 2025-05-23TSINGHUA UNIVERSITY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210714231.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-22
Publication Date
2025-05-23
Estimated Expiration
2042-06-22

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently realize beam coordination optimization of multiple 5G base stations, resulting in exponential growth of search space and making it difficult to achieve high-performance and low-overhead beam coordination optimization.

Method used

The optimization model containing multiple deep Q networks is adopted, combined with UCB exploration method, and the antenna weights are jointly optimized for multiple 5G base stations to obtain the optimal antenna weights.

Benefits of technology

Effectively search for the optimal coordinated beam in a huge search space, improve the optimization quality and sample efficiency of the coordinated beam, and achieve high-performance and low-overhead 5G multi-base station beam coordination optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115190510B_ABST
    Figure CN115190510B_ABST
Patent Text Reader

Abstract

The present invention relates to a 5G multi-base station antenna weight joint optimization method and device, including: obtaining the state information of multiple 5G base stations to be optimized; based on the state information, using the UCB exploration method on an optimization model containing multiple deep Q networks to jointly optimize the antenna weights of the multiple 5G base stations, and obtaining the optimal antenna weights of the multiple 5G base stations. The present invention is based on deep integrated uncertainty estimation and a sample effective confidence upper bound (UCB) search strategy to effectively search for the optimal coordinated beam in a very large search space, thereby improving the optimization quality and sample efficiency of the coordinated beam.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technologies, and in particular to a method and device for joint optimization of 5G multi-base station antenna weights. Background Art

[0002] Massive MIMO (Massive Multiple Input Multiple Output) technology is a key technology for improving system capacity, network coverage and spectrum efficiency in 5G (5th Generation Mobile Communication Technology). It adjusts the antenna weight parameters of hundreds of antennas in the base station (BS) to form a narrow and directional high-gain signal beam. In the Massive MIMO system, there is inevitably coupling effect (inter-cell interference, ICI) between multiple densely deployed 5G base stations. To reduce this effect, it is necessary to coordinate and optimize the beamforming of multiple base stations. However, due to the range of antenna weight parameters and the number of antennas, the search space for antenna weight parameters is very large when a single 5G base station beam is formed. The coordinated optimization of multiple 5G base station beams will lead to an exponential growth in the search space, which is very difficult to implement.

[0003] At present, beamforming optimization methods often adopt the following three methods:

[0004] The first category is wave velocity formation optimization methods that rely on analytical modeling and convex or non-convex optimization. This type of method provides a good mathematical modeling framework, but has high complexity and has difficulties in incorporating complex real-world environment geometry or blockage.

[0005] The second category is beamforming optimization methods that use machine learning techniques such as deep learning and reinforcement learning. For example, deep neural networks have been used to directly fit the optimal beamforming in a supervised learning manner, or combined with Monte Carlo methods to search for the optimal beamforming vector. The disadvantage of this type of method is that it requires time-consuming sample data collection and model training overhead.

[0006] The third category is beamforming optimization methods that use a reinforcement learning (RL) framework based on a value function. This type of method is usually implemented together with a 3D ray tracing simulator for MIMO channel modeling and evaluation, so that the environmental impact can be realistically captured. Although this type of method has achieved great success in complex decision-making tasks, RL methods also rely on extensive interactive exploration and data collection in simulated or real environments. Since beamforming vector optimization is a single-stage optimization problem (different operations do not change the state of the environment), using sequential decision-making methods such as RL may be overly destructive and lack sample efficiency. The cost of running a high-fidelity 3D ray tracing simulator simultaneously is relatively high. When extended to beamforming collaborative optimization of multiple base stations, the computational cost used for simulation may be very expensive and it is impossible to support large-scale practical implementation.

[0007] In short, there is an urgent need to provide a high-performance and low-overhead implementation method for the coordinated optimization of multiple 5G base station beams. Summary of the invention

[0008] The purpose of the present invention is to provide a 5G multi-base station antenna weight joint optimization method and device to solve the problem that multiple 5G base station beam coordination optimization is difficult to implement in the prior art, and to achieve high-performance and low-overhead 5G multi-base station beam coordination optimization.

[0009] In a first aspect, the present invention provides a 5G multi-base station antenna weight joint optimization method, the method comprising:

[0010] Obtain status information of multiple 5G base stations to be optimized;

[0011] Based on the state information, the UCB exploration method is used on an optimization model including multiple deep Q networks to jointly optimize the antenna weights of the multiple 5G base stations to obtain the optimal antenna weights of the multiple 5G base stations.

[0012] According to the 5G multi-base station antenna weight joint optimization method provided by the present invention, the antenna weights of the multiple 5G base stations are jointly optimized based on the state information and in an optimization model including multiple deep Q networks using a UCB exploration method to obtain the optimal antenna weights of the multiple 5G base stations, including:

[0013] Step A: Initializing each deep Q network and the number of iterations in the optimization model, and randomly setting the optimal antenna weight of each 5G base station in the multiple 5G base stations;

[0014] Step B: In this iteration, the multiple 5G base stations are randomly sorted to obtain a corresponding arrangement list;

[0015] Step C: taking the first 5G base station in the arrangement list that has not performed antenna weight optimization as the target 5G base station;

[0016] Step D: Utilizing the state information and the optimal antenna weights of other 5G base stations among the multiple 5G base stations except the target 5G base station, optimizing the antenna weights of the target 5G base station using a UCB exploration method on the optimization model;

[0017] Step E: updating the optimization model based on the antenna weight optimization result of the target 5G base station, and updating the optimal antenna weight of the target 5G base station to the antenna weight optimization result of the target 5G base station;

[0018] Step F: If the target 5G base station is the last 5G base station in the arrangement list, then updating the optimal antenna weight of each of the multiple 5G base stations again according to the optimal antenna weight of each of the multiple 5G base stations at the end of the previous iteration; otherwise, return to step C;

[0019] Step G: If the current number of iterations is equal to the preset maximum number of iterations, then output the optimal antenna weight of each 5G base station in the multiple 5G base stations at this time; otherwise, increase the number of iterations by 1 and return to step B.

[0020] According to the 5G multi-base station antenna weight joint optimization method provided by the present invention, step D comprises:

[0021] The optimal antenna weights of other 5G base stations in the plurality of 5G base stations except the target 5G base station are used to form

[0022] The Input into each deep Q network of the optimization model so that each deep Q network of the optimization model learns according to the preset value function The Q value for each antenna weight in the pre-stored antenna weight action set;

[0023] calculate The mean and standard deviation of the Q value output by the deep Q network in the optimization model for each antenna weight in the pre-stored antenna weight action set;

[0024] Based on the mean and standard deviation, a preset UCB exploration formula is used to explore within the antenna weight action set.

[0025] Where S = {S 1 , …, S i , …, S n},S i represents the status information of the i-th 5G base station among the multiple 5G base stations, Indicates the antenna weight optimization result of the i-th 5G base station among the multiple 5G base stations, that is, the antenna weight optimization result of the target 5G base station; represents the antenna weight of the i-th 5G base station among the multiple 5G base stations, It indicates that the antenna weight of the f-th 5G base station to be optimized is its optimal antenna weight, f≠i, f∈(1~n), and n represents the total number of 5G base stations to be optimized.

[0026] According to the 5G multi-base station antenna weight joint optimization method provided by the present invention, the cost function is used to characterize Q k (S, A) infinitely approaches r(S, A); where r(S, A) represents the global reward under (S, A), Q k (S, A) represents the Q value output by the kth deep Q network of the optimization model under (S, A), k∈(1~K), K is the total number of deep Q networks in the optimization model;

[0027] The expression of r(S,A) is as follows:

[0028] r(S,A)=1-αWSC(h S,A ,ρ)-(1-α)ICI(h S,A , ρ)

[0029] In the above formula, α represents the weight factor, WSC(h S,A , ρ) represents the overall weak signal coverage, ICI(h S,A , ρ) represents the inter-cell interference, h S,A represents the channel determined by S and A, ρ represents the user density,

[0030] The expression of the UCB exploration formula is as follows:

[0031]

[0032] In the above formula, express is the antenna weight action set A dzj In and Respectively The mean and standard deviation of the Q-values ​​output by the deep Q-network in the optimization model described below, β represents a hyperparameter that controls the aggressiveness of exploration.

[0033] According to the 5G multi-base station antenna weight joint optimization method provided by the present invention, the optimization model is updated based on the antenna weight optimization result of the target 5G base station, including:

[0034] Check if there is a global reward in the result buffer

[0035] No global reward exists in the result buffer In the case of Input into the high-fidelity MIMO simulator for simulation and get global rewards

[0036] Take advantage of global rewards Each deep Q network in the optimization model is updated by using regression fitting method.

[0037] According to the 5G multi-base station antenna weight joint optimization method provided by the present invention, the optimal antenna weight of each 5G base station in the multiple 5G base stations is updated again according to the optimal antenna weight of each 5G base station in the multiple 5G base stations at the end of the previous iteration, including:

[0038] Compare Global Rewards and global rewards size;

[0039] If the global reward Greater than global rewards Then the optimal antenna weight of the i-th 5G base station among the multiple 5G base stations is updated to

[0040] Otherwise, the optimal antenna weight of the i-th 5G base station among the multiple 5G base stations is updated to

[0041] While updating the optimal antenna weight of each of the multiple 5G base stations according to the optimal antenna weight of each of the multiple 5G base stations at the end of the previous iteration, it also includes:

[0042] Global Rewards Added into the result buffer.

[0043] According to the 5G multi-base station antenna weight joint optimization method provided by the present invention, the antenna weights include: uptilt angle, azimuth angle, horizontal beam width and vertical beam width;

[0044] The process of acquiring the antenna weight action set includes:

[0045] The multidimensional beam shape related variables are mapped into one-dimensional continuous variables by using multidimensional scaling analysis method.

[0046] Discretize the one-dimensional continuous variables, azimuth angle and uptilt angle, and then obtain corresponding discrete action quantities;

[0047] The antenna weight action set is generated using the discrete action quantity.

[0048] In a second aspect, the present invention further provides a 5G multi-base station antenna weight joint optimization device, the device comprising:

[0049] An acquisition module, used to obtain status information of multiple 5G base stations to be optimized;

[0050] A joint optimization module is used to perform joint optimization of antenna weights for the multiple 5G base stations based on the state information and using a UCB exploration method on an optimization model including multiple deep Q networks to obtain the optimal antenna weights for the multiple 5G base stations.

[0051] In the third aspect, the present invention also discloses an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the 5G multi-base station antenna weight joint optimization method as described in the first aspect is implemented.

[0052] In a fourth aspect, the present invention further discloses a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the 5G multi-base station antenna weight joint optimization method as described in the first aspect is implemented.

[0053] The present invention provides a 5G multi-base station antenna weight joint optimization method and device, which obtains the state information of multiple 5G base stations to be optimized; based on the state information, the UCB exploration method is used on the optimization model containing multiple deep Q networks to jointly optimize the antenna weights of the multiple 5G base stations to obtain the optimal antenna weights of the multiple 5G base stations. Based on the deep integrated uncertainty estimation and the sample effective confidence upper bound (UCB) search strategy, the present invention effectively searches for the optimal coordinated beam in a very large search space, thereby improving the optimization quality and sample efficiency of the coordinated beam. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0055] Figure 1 It is a flow chart of a 5G multi-base station antenna weight joint optimization method provided by the present invention;

[0056] Figure 2 It is a schematic diagram of a 5G multi-base station antenna weight joint optimization framework provided by the present invention;

[0057] Figure 3It is a comparison chart of the best rewards between the antenna weight joint optimization method provided by the present invention and the traditional method;

[0058] Figure 4 It is a comparison diagram of the best WCS between the antenna weight joint optimization method provided by the present invention and the traditional method;

[0059] Figure 5 It is a comparison diagram of the best ICI between the antenna weight joint optimization method provided by the present invention and the traditional method;

[0060] Figure 6 This is a structural diagram of a 5G multi-base station antenna weight joint optimization device provided by the present invention;

[0061] Figure 7 It is a schematic diagram of the structure of an electronic device for implementing a 5G multi-base station antenna weight joint optimization method provided by the present invention;

[0062] Reference numerals: A: 5G multi-base station antenna weight joint optimization method provided by the present invention

[0063] B: Traditional multi-base station beam coordination optimization method based on MASAC;

[0064] C: Traditional multi-base station beam coordination optimization method based on MADDPG;

[0065] D: Traditional multi-base station beam coordination optimization method based on GP-UCB. DETAILED DESCRIPTION

[0066] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0067] Massive multiple-input multiple-output (MIMO) is one of the core technologies of the fifth-generation (5G) cellular networks and beamforming plays an important role in MIMO communications. Correctly setting the beam of a 5G base station can greatly improve the communication service quality of mobile users and has great practical value.

[0068] Here are some traditional MIMO system beamforming methods:

[0069] The first one: beamforming based on angle of arrival;

[0070] Beamforming based on angle of arrival (AOA) is a commonly used method in MIMO systems. Specifically, for a base station with L rows and M columns of antenna elements (AEs), according to AOA (Ψ k ,φ k ) configures the weights W on the antenna elements (AEs) k (L×M matrix) to form a beam. Here Ψ k and φ k are the azimuth and elevation angles of the kth user equipment (UE), respectively:

[0071] W k =[ω mk ·ξ lk ]l=1,…,L;m=1,…,M

[0072]

[0073]

[0074] Among them, d h and d v are the row distance and column distance in antenna elements (AEs), respectively. In AOA-based beamforming, the weight W k Depends on (Ψ k ,φ k ), which requires accurate channel state information (CSI) to estimate (Ψ k ,φ k ). This process can be very complex and becomes impractical when jointly modeling the coupling effects of multiple BSs.

[0075] Second: Search-based beamforming:

[0076] In order to reduce the modeling complexity, compared with AOA estimation, search-based beamforming methods have become another promising direction, that is, finding the antenna weights under the optimal beam for a single BS. This is close to searching all possible antenna weights (including azimuth and downtilt angles) and finding the optimal antenna weights. For example, given the channel h and UE density ρ, the antenna weights are searched with the goal of minimizing the overall weak signal coverage WSC (h, ρ) and inter-cell interference ICI (h, ρ).

[0077]

[0078] Among them, WSC(h, p) and ICI(h, p) can be defined as the probability that the reference signal received power (RSRP) and signal-to-noise ratio (SINR) are less than some predefined thresholds in a specific cell:

[0079] WSC(h,p)=P r(RSRP<T r |h,p)

[0080] ICI(h, p) = P r (sinr<T s |h,p)

[0081] make Indicates that from The first BS serves The signal strength at the cell is The RSRP and SINR at each cell can be evaluated as:

[0082]

[0083]

[0084] in, is the number of BSs, and B(h, p) is the background noise. In practice, the UE density ρ can be obtained by modeling empirical data in real scenarios and can be evaluated in real environments or high-fidelity MIMO simulations (Ψ k ,φ k ) under the WSC and ICI.

[0085] In most real-world 5G MIMO systems, for a single antenna in a single 5G base station, the factors that affect beamforming include not only the optimized azimuth and uptilt angles, but also the beam shape properties, which results in the antenna weights of each antenna containing three elements, namely, A i =(p i ,Ψ i ,φ i ). Because there are many antennas in a single 5G base station, the search space for antenna weights of a single 5G base station is very huge. Due to the complex real-world environment (such as location, terrain, and building obstacles) that has a great impact on signal propagation, the optimal beams of different base stations are very different. If we want to expand the single-base station beamforming optimization problem to the beamforming problem of multiple densely deployed 5G base stations, we need to jointly optimize multiple 5G base stations with the goal of minimizing the overall WSC and ICI of all cells in a specific area, so that each base station can find the optimal antenna weight.

[0086] It can be seen that the present invention actually aims to solve the multi-agent beam coordination optimization problem, as shown below:

[0087]

[0088] Among them, r(S, A) is a global reward function used to capture the coupling effect of all base stations on all cells in a specific area, S = {S1 , …, S i , …, S n},S i ∈ S,S i A contains, for example, altitude, longitude, latitude and other BS-specific static information; A = {A 1 , …, A i , …, A n}, A i ∈Antenna weight action set.

[0089] r(S,A)=1-1-αWSC(h,ρ)-(1-α)ICI(h,ρ)

[0090] Here, h depends on A in all base stations.

[0091] The multi-agent beam coordination optimization problem allows for simultaneous coordination and optimization of multiple base station beams to reduce the complex coupling effects between different base stations (i.e., suppress ICI), thereby providing greater performance gains than optimizing each base station beam separately in real scenarios. However, solving the above multi-agent beam coordination optimization problem faces the problem of action space (|A| n ), which is time-consuming and costly. Therefore, an optimization algorithm that can effectively explore the huge action space is the key to successful deployment.

[0092] Combine the following Figure 1-Figure 7 The present invention describes a 5G multi-base station antenna weight joint optimization method and device.

[0093] In a first aspect, the present invention provides a 5G multi-base station antenna weight joint optimization method, such as Figure 1 As shown, the method includes:

[0094] S11. Obtain status information of multiple 5G base stations to be optimized;

[0095] S12. Based on the state information, a UCB exploration method is used on an optimization model including multiple deep Q networks to jointly optimize the antenna weights of the multiple 5G base stations to obtain the optimal antenna weights of the multiple 5G base stations.

[0096] There is an inevitable coupling effect between multiple densely deployed 5G base stations, so it is a complex task to simultaneously determine the optimal beams of multiple 5G base stations. To this end, this paper proposes an efficient multi-agent optimization framework, whose core algorithm relies on deep integration of neural networks and a search strategy based on sample valid confidence upper bound (UCB). This method can effectively search for the optimal coordinated beam in a very large search space, and is superior to traditional multi-agent reinforcement learning methods in terms of optimization quality and sample efficiency.

[0097] The present invention provides a 5G multi-base station antenna weight joint optimization method, which is based on deep integrated uncertainty estimation and a sample effective confidence upper bound (UCB) search strategy, and effectively searches for the optimal coordinated beam in a very large search space, thereby improving the optimization quality and sample efficiency of the coordinated beam.

[0098] On the basis of the above embodiments, as an optional embodiment, the method of jointly optimizing the antenna weights of the multiple 5G base stations based on the state information and using the UCB exploration method on the optimization model including multiple deep Q networks to obtain the optimal antenna weights of the multiple 5G base stations includes:

[0099] Step A: Initializing each deep Q network and the number of iterations in the optimization model, and randomly setting the optimal antenna weight of each 5G base station in the multiple 5G base stations;

[0100] Step B: In this iteration, the multiple 5G base stations are randomly sorted to obtain a corresponding arrangement list;

[0101] Step C: taking the first 5G base station in the arrangement list that has not performed antenna weight optimization as the target 5G base station;

[0102] Step D: Utilizing the state information and the optimal antenna weights of other 5G base stations among the multiple 5G base stations except the target 5G base station, optimizing the antenna weights of the target 5G base station using a UCB exploration method on the optimization model;

[0103] Step E: updating the optimization model based on the antenna weight optimization result of the target 5G base station, and updating the optimal antenna weight of the target 5G base station to the antenna weight optimization result of the target 5G base station;

[0104] Step F: If the target 5G base station is the last 5G base station in the arrangement list, then updating the optimal antenna weight of each of the multiple 5G base stations again according to the optimal antenna weight of each of the multiple 5G base stations at the end of the previous iteration; otherwise, return to step C;

[0105] Step G: If the current number of iterations is equal to the preset maximum number of iterations, then output the optimal antenna weight of each 5G base station in the multiple 5G base stations at this time; otherwise, increase the number of iterations by 1 and return to step B.

[0106] It is important to understand that 3D ray tracing has become a popular technique for radio frequency (RF) analysis and MIMO channel modeling and has been applied in many wireless shaping studies in recent years. One of the applications is the high-fidelity MIMO simulator, which implements the launch and bounce ray (SBR) tracing method based on the real-world UE distribution ρ estimated by empirical data and real environment characteristics (such as 3D terrain maps and buildings) to evaluate the high-fidelity MIMO channel characteristics h. The SBR method can capture the impact of signal surface reflection and is effective in the frequency range of 100MHz to 100GHz. The benefit of using the ray tracing method is that it has high sensitivity in the surrounding environment, can accurately reflect the real channel information, and does not require a large number of and expensive field measurements. On the other hand, high-fidelity ray tracing simulations also incur high computational costs, especially for complex scenarios with a large number of reflective surfaces and many MIMO systems with large antenna arrays (such as 5G BSs).

[0107] The present invention provides Figure 2 The 5G multi-base station antenna weight joint optimization framework shown in the figure is used to coordinate the MIMO beam optimization problem. Under this framework, the present invention uses a three-dimensional ray tracing simulator based on a real environment (i.e., a high-fidelity MIMO simulator) for channel modeling and performance evaluation during the optimization process. Because for a single base station, the possible action space is already very large. The beam coordination optimization problem of multiple base stations directly causes the action space to grow exponentially (i.e., |A| n ), the action search space is even larger. If the optimization algorithm is not sample efficient when interacting with the high-fidelity simulator, the high computational cost of high-fidelity ray tracing simulation will make the training cost unaffordable. Therefore, the present invention proposes a lightweight but efficient multi-agent optimization algorithm, which is constructed based on deep integration of neural networks and exploration based on sample efficient confidence upper bounds, which can minimize the number of interactions with the simulation environment on a large number of deployed 5G base stations, while retaining sufficient implicitness for actual deployment.

[0108] The core of this algorithm lies in the following two points:

[0109] First: Set the value function Q k (S, A), k = {1, ..., K}, a set of deep Q networks, Q k (S,A) approximates the global reward function r(S,A);

[0110] Second: adopt the UCB exploration procedure to greedily search for the best possible action and incorporate uncertainty information.

[0111] For the first point, the deep Q-network ensemble learns Q k (S, A), then calculate Q k(S, A), k = {1, ..., K}, and then use the mean and standard deviation to perform a UCB-style greedy search. In addition, there are several benefits to using a deep Q-network ensemble: First, the ensemble method of multi-model prediction is a common technique in machine learning that helps improve prediction performance and compensate for the inadequate learning of a single model. Second, deep neural network ensembles have recently been shown to be one of the most effective methods for uncertainty estimation (e.g., evaluating the mean μ(S, A) and standard deviation σ(S, A), and constructing an efficient UCB-style exploration based on the evaluated μ(S, A) and σ(S, A).

[0112] Regarding the second point, UCB is a class of efficient algorithms that handle the exploration-exploitation tradeoff in online decision making with partial information feedback, and has achieved great success in multi-armed bandit problems and even some RL problems. The idea of ​​UCB-style exploration is to take optimistic actions in the face of uncertainty. The confidence upper bound can be evaluated as the empirical average return plus a term proportional to the uncertainty of the action. This strategy can show superior sample efficiency and has strong theoretical guarantees.

[0113] We develop a multi-base station UCB exploration based on a deep integration of previously learned value functions. To reduce the overall action space and improve efficiency, we do not perform a multi-base station UCB exploration on the joint action space (|A| n ), but evaluates the action Ai of a single base station in each iteration and uses the maximum UCB action Determine the optimal value of Ai and obtain the optimal action of all base stations in this iteration, which changes the UCB evaluation quantity from |A| n It is reduced to n×|A|. It is worth mentioning that in order to prevent the iterative action update strategy from causing the exploration to fall into the local optimal solution, we further introduced a randomization scheme to arrange the update order of base stations in each iteration round (for example: ζ = {1, ..., n}; 1, 2, .....n are randomly arranged), and then perform UCB exploration in sequence according to the order specified in ζ.

[0114] In summary, the simple and efficient multi-agent optimization framework proposed in this embodiment is used for coordinated beamforming involving multiple base stations in actual 5G cellular networks. The core algorithm in this framework relies on deep integrated uncertainty estimation and exploration based on upper confidence bound (UCB), exploring strong multi-agent reinforcement learning (MARL), achieving superior optimization performance and better sample efficiency.

[0115] Based on the above embodiments, as an optional embodiment, step D includes:

[0116] The optimal antenna weights of other 5G base stations in the plurality of 5G base stations except the target 5G base station are used to form

[0117] The Input into each deep Q network of the optimization model, so that each deep Q network of the optimization model learns according to the preset value function The Q value for each antenna weight in the pre-stored antenna weight action set;

[0118] calculate The mean and standard deviation of the Q value output by the deep Q network in the optimization model for each antenna weight in the pre-stored antenna weight action set;

[0119] Based on the mean and standard deviation, a preset UCB exploration formula is used to explore within the antenna weight action set.

[0120] Where S = {S 1 , …, S i , …, S n},S i represents the status information of the i-th 5G base station among the multiple 5G base stations, Indicates the antenna weight optimization result of the i-th 5G base station among the multiple 5G base stations, that is, the antenna weight optimization result of the target 5G base station; represents the antenna weight of the i-th 5G base station among the multiple 5G base stations, It indicates that the antenna weight of the f-th 5G base station to be optimized is its optimal antenna weight, f≠i, f∈(1~n), and n represents the total number of 5G base stations to be optimized.

[0121] It is understandable that when the present invention performs UCB exploration on a base station, the optimal actions of other base stations are kept unchanged so that the deep network only learns the actions of the base station. k (S, A), k = {1, ..., K} and the following calculation formula to obtain the mean μ(S, A) and standard deviation σ(S, A):

[0122] μ(S, A) = mean(Q k (S, A)

[0123] σ(S, A) = std(Q k (S, A)

[0124] In addition, it is important to note that due to the generalization ability of the deep Q-network, the action A explored by the optimization model UCB imay not be in the antenna weight action set. Here, an action mapping procedure is introduced to convert the action A explored by UCB i Mapped to the closest discretized action in the antenna weight action set To enforce feasibility conditions.

[0125] The relevant formula is as follows;

[0126]

[0127] The present invention improves the effectiveness of action exploration by presetting a value function and a UCB exploration formula.

[0128] Based on the above embodiments, as an optional embodiment, the value function is used to characterize Q k (S, A) infinitely approaches r(S, A); where r(S, A) represents the global reward under (S, A), Q k (S, A) represents the Q value output by the kth deep Q network of the optimization model under (S, A), k∈(1~K), K is the total number of deep Q networks in the optimization model;

[0129] The expression of r(S, A) is as follows:

[0130] r(S, A)=1-αWSC(h S,A ,ρ)-(1-α)ICI(h S,A , ρ)

[0131] In the above formula, α represents the weight factor, WSC(h S,A , ρ) represents the overall weak signal coverage, ICI(h S,A , ρ) represents the inter-cell interference, h S,A represents the channel determined by S and A, ρ represents the user density,

[0132] The expression of the UCB exploration formula is as follows:

[0133]

[0134] In the above formula, express is the antenna weight action set A dzj In and Respectively The mean and standard deviation of the Q-values ​​output by the deep Q-network in the optimization model described below, β represents a hyperparameter that controls the aggressiveness of exploration.

[0135] It can be understood that the optimization model of the present invention searches for the best action with the goal of minimizing the overall weak signal coverage (WSC) and inter-cell interference (ICI). In the UCB exploration of the present invention, β>0 is set; since μ(S, A) and σ(S, A) are evaluated by deep neural networks, the UCB values ​​of all possible action configurations of a single base station (size |A|) can be efficiently calculated in parallel in one or more batches, which is very effective in modern deep learning frameworks such as Tensorfflow or Pytorch.

[0136] Based on the above embodiments, as an optional embodiment, updating the optimization model based on the antenna weight optimization result of the target 5G base station includes:

[0137] Check if there is a global reward in the result buffer

[0138] No global reward exists in the result buffer In the case of Input into the high-fidelity MIMO simulator for simulation and get global rewards

[0139] Take advantage of global rewards Each deep Q network in the optimization model is updated by using regression fitting method.

[0140] It should be noted that the present invention Q k (S, A) is close to the global reward function r(S, A), so the deep Q network is updated by regression fitting. The update formula is as follows:

[0141]

[0142] The present invention uses r(S, A) simulated by a three-dimensional ray tracing simulator based on a real environment to perform regression fitting on the deep Q network, so that the deep Q network tends to be accurate as the number of UCB explorations increases, thereby achieving efficient UCB exploration.

[0143] Based on the above embodiments, as an optional embodiment, updating the antenna weight of each 5G base station in the multiple 5G base stations again according to the optimal antenna weight of each 5G base station in the multiple 5G base stations at the end of the previous iteration includes:

[0144] Compare Global Rewards and global rewards size;

[0145] If the global reward Greater than global rewards Then the optimal antenna weight of the i-th 5G base station among the multiple 5G base stations is updated to

[0146] Otherwise, the optimal antenna weight of the i-th 5G base station among the multiple 5G base stations is updated to

[0147] While updating the optimal antenna weight of each of the multiple 5G base stations according to the optimal antenna weight of each of the multiple 5G base stations at the end of the previous iteration, it also includes:

[0148] Global Rewards Added into the result buffer.

[0149] The present invention constructs a result buffer B for storing unique simulation samples r(S, A) to avoid the high-fidelity MIMO simulator from resimulating already simulated actions, thereby reducing unnecessary resource consumption.

[0150] Based on the above embodiments, as an optional embodiment, the antenna weights include: an uptilt angle, an azimuth angle, a horizontal beam width, and a vertical beam width;

[0151] The process of acquiring the antenna weight action set includes:

[0152] The multidimensional beam shape related variables are mapped into one-dimensional continuous variables by using multidimensional scaling analysis method.

[0153] Discretize the one-dimensional continuous variables, azimuth angle and uptilt angle, and then obtain corresponding discrete action quantities;

[0154] The antenna weight action set is generated using the discrete action quantity.

[0155] It should be noted that real-world 5G base stations typically describe the beam shape and how the antenna radiates energy into space based on a set of pre-configured antenna patterns. Each base station antenna pattern contains parameters such as azimuth, uptilt, beamwidth, and amplifier gain. The azimuth and uptilt have specific adjustable ranges. This leads to an extremely complex and discontinuous optimization search space. In order to reduce model complexity and further reduce simulation costs, we discretize the effective azimuth and uptilt of each pattern to 1 degree, and only simulate these discrete valid action values. This results in a total of about 5,000 valid operating configurations for a single base station (Huawei 5G RAN equipment |a| = 5091. Other 5G devices may be different).

[0156] The present invention is explained in detail below by taking K=3 as an example.

[0157] We use K=3 to represent the number of networks in the optimized model’s depth Q network. Three networks already produce good uncertainty estimates, which helps reduce computational cost compared to using more networks. k It is implemented as a five-layer fully connected neural network with RELU activation. The number of units in the network layer is set to [6n, 512, 512, 128, 1], where n is the total number of base stations. The Adam optimizer with a learning rate of 0.0001 is used to train Q k The β used to calculate the UCB value is set to 2.

[0158] The following are the actual steps:

[0159] Step 1: Randomly initialize each deep Q network Q k , and randomly set the optimal action for each 5G base station, the result buffer

[0160] Step 2: For iteration number 1 to the maximum iteration number, generate a randomly arranged ordered list of multiple base stations;

[0161] Step 3: The first 5G base station in the ordered list that has not performed UCB exploration is used as the target 5G base station;

[0162] Step 4: Perform UCB exploration on the target 5G base station;

[0163] Step 5: Update Q according to the global reward corresponding to the UCB exploration action of the target 5G base station k and updating the optimal action of the target 5G base station to the UCB exploration action of the target 5G base station;

[0164] Step 6: Repeat steps 3 to 5 until there are no 5G base stations that have not been explored by UCB in the ordered list;

[0165] Step 7: If the action sequence consisting of the UCB exploration actions of all 5G base stations is not in B, use a high-fidelity MIMO simulator to simulate the global reward corresponding to the action sequence and fill it into B accordingly; at the same time, based on the optimal action of each 5G base station obtained in the previous iteration, update the optimal action of each 5G base station in this iteration again;

[0166] Step 8: If the current number of iterations is equal to the preset maximum number of iterations, then output the optimal antenna weights of each of the multiple 5G base stations at this time and end; otherwise, return to step 2.

[0167] Under the same conditions, multi-base station beam coordination optimization based on GP-UCB, MADDPG and MASAC was performed. Figure 3A comparison diagram of the best reward between the method of the present invention and the traditional multi-base station beam coordination optimization methods based on GP-UCB, MADDPG and MASAC is illustrated; Figure 4 A comparison diagram of the best WCS between the method of the present invention and the traditional multi-base station beam coordination optimization methods based on GP-UCB, MADDPG and MASAC is illustrated; Figure 5 The comparison diagrams of the best ICI between the method of the present invention and the traditional multi-base station beam coordination optimization methods based on GP-UCB, MADDPG and MASAC are illustrated; in these three diagrams, A represents the method of the present invention, B represents the multi-base station beam coordination optimization method based on MASAC, C represents the multi-base station beam coordination optimization method based on MADDPG, and D represents the multi-base station beam coordination optimization method based on GP-UCB.

[0168] From these three figures, we can see that the algorithm of the present invention converges to the best reward of 0.764 within 500 steps, while the other algorithms are still exploring the environment, and it seems difficult to approach the best reward achieved by the algorithm of the present invention within a limited number of steps. A huge obstacle for MARL algorithms is that they require a large number of samples and expensive training steps to cover the optimal solution, however, 2000 training steps are difficult to meet their requirements. This is in line with our expectations that MASAC is more effective than MADDPG, because MADDPG simply adds random noise to the action to encourage exploration, while MASAC is equipped with entropy loss of the action, thus encouraging efficient exploration. WCS optimization from 11% to 9%, but ICI optimization from 40% to 28%, ICI has great potential, and it makes sense to continuously optimize ICI. From an industry perspective, when WCS reaches a certain threshold, network devices will remain connected, and there is no need to further optimize this item. However, it is very important to optimize ICI as much as possible. As ICI decreases, the probability of data transmission errors will become lower and lower. In addition, the algorithm of the present invention has achieved significant performance in single-agent GP-UCB, which can enhance communication and cooperation between agents, thereby improving overall performance. In conclusion, the algorithm of the present invention is highly efficient and can achieve optimal performance.

[0169] In summary, the present invention proposes a simple and effective multi-agent algorithm for beam coordination optimization involving multiple base stations in realistic 5G cellular networks. The use of MIMO simulators comes at the expense of a large amount of computing resources and samples, and the biggest highlight of the algorithm of the present invention is the introduction of an integrated value network to reduce variance and naturally adopt UCB to promote valuable exploration to improve sample efficiency. According to experience, the algorithm of the present invention can efficiently explore beneficial coordination settings and obtain optimal performance with limited training steps. In addition, the algorithm of the present invention has strong scalability to the number of agents and can achieve better coordinated beamforming in a growing number of agents.

[0170] Secondly, the 5G multi-base station antenna weights joint optimization device provided by the present invention is described. The 5G multi-base station antenna weights joint optimization device described below and the 5G multi-base station antenna weights joint optimization method described above can refer to each other. Figure 6 A structural schematic diagram of a 5G multi-base station antenna weight joint optimization device is illustrated, and the device in the figure includes:

[0171] An acquisition module 21 is used to obtain status information of multiple 5G base stations to be optimized;

[0172] The joint optimization module 22 is used to perform joint optimization of antenna weights for the multiple 5G base stations based on the state information and using a UCB exploration method on an optimization model including multiple deep Q networks to obtain the optimal antenna weights for the multiple 5G base stations.

[0173] The present invention provides a 5G multi-base station antenna weight joint optimization device, which is based on deep integrated uncertainty estimation and a sample effective confidence upper bound (UCB) search strategy, and effectively searches for the optimal coordinated beam in a very large search space, thereby improving the optimization quality and sample efficiency of the coordinated beam.

[0174] The 5G multi-base station antenna weights joint optimization device provided in the embodiment of the present invention specifically executes the above-mentioned 5G multi-base station antenna weights joint optimization method embodiment processes. For details, please refer to the contents of the above-mentioned 5G multi-base station antenna weights joint optimization method embodiments, which will not be repeated here.

[0175] Thirdly, Figure 7 The following is a schematic diagram of the physical structure of an electronic device. Figure 7 As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730 and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call the logic instructions in the memory 730 to execute a 5G multi-base station antenna weight joint optimization method, the method comprising: obtaining status information of multiple 5G base stations to be optimized; based on the status information, using a UCB exploration method on an optimization model including multiple deep Q networks to jointly optimize the antenna weights of the multiple 5G base stations, and obtaining the optimal antenna weights of the multiple 5G base stations.

[0176] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0177] In a fourth aspect, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, it executes a 5G multi-base station antenna weight joint optimization method, the method comprising: obtaining status information of multiple 5G base stations to be optimized; based on the status information, using a UCB exploration method on an optimization model containing multiple deep Q networks to jointly optimize the antenna weights of the multiple 5G base stations, so as to obtain the optimal antenna weights of the multiple 5G base stations.

[0178] In a fifth aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon for executing a 5G multi-base station antenna weight joint optimization method, the method comprising: obtaining status information of multiple 5G base stations to be optimized; based on the status information, using a UCB exploration method on an optimization model comprising multiple deep Q networks to jointly optimize the antenna weights of the multiple 5G base stations, to obtain the optimal antenna weights of the multiple 5G base stations.

[0179] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0180] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A joint optimization method for 5G multi-base station antenna weights, It is characterized in that The method comprises: Obtain status information of multiple 5G base stations to be optimized; Based on the state information, a UCB exploration method is used on an optimization model including multiple deep Q networks to jointly optimize the antenna weights of the multiple 5G base stations to obtain the optimal antenna weights of the multiple 5G base stations; The method of jointly optimizing the antenna weights of the multiple 5G base stations based on the state information and using the UCB exploration method on the optimization model including multiple deep Q networks to obtain the optimal antenna weights of the multiple 5G base stations includes: Step A: Initializing each deep Q network and the number of iterations in the optimization model, and randomly setting the optimal antenna weight of each 5G base station in the multiple 5G base stations; Step B: In this iteration, the multiple 5G base stations are randomly sorted to obtain a corresponding arrangement list; Step C: taking the first 5G base station in the arrangement list that has not performed antenna weight optimization as the target 5G base station; Step D: Utilizing the state information and the optimal antenna weights of other 5G base stations among the multiple 5G base stations except the target 5G base station, optimizing the antenna weights of the target 5G base station using the UCB exploration method on the optimization model; Step E: updating the optimization model based on the antenna weight optimization result of the target 5G base station, and updating the optimal antenna weight of the target 5G base station to the antenna weight optimization result of the target 5G base station; Step F: If the target 5G base station is the last 5G base station in the arrangement list, then updating the optimal antenna weight of each of the multiple 5G base stations again according to the optimal antenna weight of each of the multiple 5G base stations at the end of the previous iteration; otherwise, return to step C; Step G: If the current number of iterations is equal to the preset maximum number of iterations, then output the optimal antenna weight of each 5G base station in the multiple 5G base stations at this time; otherwise, increase the number of iterations by 1 and return to step B.

2. According to claim 1, the 5G multi-base station antenna weight joint optimization method, It is characterized in that The step D comprises: Using the state information and the optimal antenna weights of the multiple 5G base stations to form ; The Input into each deep Q network of the optimization model so that each deep Q network of the optimization model learns according to the preset value function The Q value for each antenna weight in the pre-stored antenna weight action set; calculate The mean and standard deviation of the Q value output by the deep Q network in the optimization model for each antenna weight in the pre-stored antenna weight action set; Based on the mean and standard deviation, a preset UCB exploration formula is used to explore within the antenna weight action set. ; in, , Indicates the first of the multiple 5G base stations Status information of 5G base stations, Indicates the first of the multiple 5G base stations The antenna weight optimization result of each 5G base station, that is, the antenna weight optimization result of the target 5G base station; Indicates the first of the multiple 5G base stations The antenna weights of a 5G base station, Indicates the first The 5G base station antenna weight is its optimal antenna weight, , n represents the total number of 5G base stations to be optimized.

3. According to claim 2, the 5G multi-base station antenna weight joint optimization method, It is characterized in that The value function is used to characterize Infinite Approach ;in, express The global reward for express The optimization model The Q value output by the deep Q network, , is the total number of deep Q networks in the optimization model; Said The expression is as follows: ; In the above formula, represents the weight factor, Indicates overall weak signal coverage, represents the inter-cell interference, Indicated by and The jointly determined channel, represents the user density, .

4. The 5G multi-base station antenna weight joint optimization method according to claim 2, It is characterized in that The updating of the optimization model based on the antenna weight optimization result of the target 5G base station includes: Check if there is a global reward in the result buffer ; No global reward exists in the result buffer In the case of Input into the high-fidelity MIMO simulator for simulation and get global rewards ; Take advantage of global rewards , each deep Q network in the optimization model is updated by regression fitting.

5. According to claim 4, the 5G multi-base station antenna weight joint optimization method, It is characterized in that The updating again the optimal antenna weight of each of the multiple 5G base stations according to the optimal antenna weight of each of the multiple 5G base stations at the end of the previous iteration includes: Compare Global Rewards and global rewards The size of represents the optimal antenna weight of the i-th 5G base station antenna to be optimized at the end of the previous iteration; If the global reward Greater than global rewards , then the first of the multiple 5G base stations The optimal antenna weights of 5G base stations are updated to ; Otherwise, the first of the multiple 5G base stations The optimal antenna weights of 5G base stations are updated to ; While updating the optimal antenna weight of each of the multiple 5G base stations according to the optimal antenna weight of each of the multiple 5G base stations at the end of the previous iteration, it also includes: Global Rewards Added into the result buffer.

6. The 5G multi-base station antenna weight joint optimization method according to claim 2, It is characterized in that The antenna weights include: uptilt angle, azimuth angle, horizontal beam width and vertical beam width; The process of acquiring the antenna weight action set includes: The multidimensional beam shape related variables are mapped into one-dimensional continuous variables by using multidimensional scaling analysis method. Discretize the one-dimensional continuous variables, azimuth angle and uptilt angle, and then obtain corresponding discrete action quantities; The antenna weight action set is generated using the discrete action quantity.

7. A 5G multi-base station antenna weight joint optimization device, It is characterized in that The device comprises: An acquisition module, used to obtain status information of multiple 5G base stations to be optimized; A joint optimization module, configured to perform joint optimization of antenna weights for the multiple 5G base stations based on the state information and using a UCB exploration method on an optimization model including multiple deep Q networks to obtain optimal antenna weights for the multiple 5G base stations; The method of jointly optimizing the antenna weights of the multiple 5G base stations based on the state information and using the UCB exploration method on the optimization model including multiple deep Q networks to obtain the optimal antenna weights of the multiple 5G base stations includes: Step A: Initializing each deep Q network and the number of iterations in the optimization model, and randomly setting the optimal antenna weight of each 5G base station in the multiple 5G base stations; Step B: In this iteration, the multiple 5G base stations are randomly sorted to obtain a corresponding arrangement list; Step C: taking the first 5G base station in the arrangement list that has not performed antenna weight optimization as the target 5G base station; Step D: Utilizing the state information and the optimal antenna weights of other 5G base stations among the multiple 5G base stations except the target 5G base station, optimizing the antenna weights of the target 5G base station using the UCB exploration method on the optimization model; Step E: updating the optimization model based on the antenna weight optimization result of the target 5G base station, and updating the optimal antenna weight of the target 5G base station to the antenna weight optimization result of the target 5G base station; Step F: If the target 5G base station is the last 5G base station in the arrangement list, then updating the optimal antenna weight of each of the multiple 5G base stations again according to the optimal antenna weight of each of the multiple 5G base stations at the end of the previous iteration; otherwise, return to step C; Step G: If the current number of iterations is equal to the preset maximum number of iterations, then output the optimal antenna weight of each 5G base station in the multiple 5G base stations at this time; otherwise, increase the number of iterations by 1 and return to step B.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, it implements the 5G multi-base station antenna weight joint optimization method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the 5G multi-base station antenna weight joint optimization method as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Antenna weight parameter optimization method and device

    CN114519295A