A Multi-City Distributed Scheduling Method and System Based on Multi-Agent Reinforcement Learning

By training a deep Q-learning network for each base station in a multi-cell cellular network and using channel state and queue length information for reward training, the problems of inter-cell interference and information overhead are solved, and a high-performance distributed scheduling strategy is achieved, which is close to the effect of a centralized scheme.

CN116347477BActive Publication Date: 2026-04-03XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-27
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies suffer from inter-cell interference in multi-cell cellular networks, which limits the increase in system capacity. Centralized scheduling schemes increase signaling exchange overhead and user latency, while existing distributed algorithms have unsatisfactory performance, and existing methods based on multi-agent reinforcement learning have not been effectively applied to user service demand interval scenarios.

Method used

A multi-agent reinforcement learning approach is adopted, in which each base station trains its own deep Q-learning network (DQN) with input channel state information and queue length information. The instantaneous indicators of the transformation are used as rewards for training, the optimal user index set is scheduled, and the low-information-overhead, high-performance scheduling strategy is output.

Benefits of technology

With low information overhead, it achieves performance close to that of centralized algorithms, with performance improvements of over 90% in single-user scenarios and over 80% in multi-user scenarios, reducing system interaction feedback and improving network performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116347477B_ABST
    Figure CN116347477B_ABST
Patent Text Reader

Abstract

A multi-cell distributed scheduling method and system based on multi-agent reinforcement learning is disclosed. The scheduling method includes setting metrics for multi-cell, multi-user QoS based on a pre-established multi-cell cellular network system model according to 3GPP; establishing an optimization target model based on the metrics for multi-cell, multi-user QoS to maximize system experience rate; and solving the optimization target model to obtain the optimal strategy for multi-cell distributed scheduling. This invention allows each base station to train its own deep Q-learning network (DQN). The input information includes channel state information and queue length information. The transformed instantaneous metrics are used as rewards for training, and the optimal set of scheduled user indices is output to achieve low information overhead and high performance. This invention's scheduling method first designs three network elements for system experience rate in a single-user scheduling scenario, then extends the design of these three elements to a multi-user scheduling scenario, and designs a more effective network structure for the multi-user scheduling scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of distributed scheduling technology for wireless multi-cell systems, specifically relating to a multi-cell distributed scheduling method and system based on multi-agent reinforcement learning. Background Technology

[0002] With the rapid increase in the number of wireless terminals and the surge in mobile data traffic, cellular networks face serious challenges. However, inter-cell interference in cellular networks restricts the increase in system capacity. Therefore, eliminating inter-cell interference in multi-cell cellular systems is an ongoing research topic. The cooperative scheduling technology of the 3GPP system, which involves joint scheduling by multiple cell base stations, has become an effective solution to solve inter-cell interference and improve spectrum efficiency. In the implementation of cooperative scheduling technology, the user's control and data signals are entirely provided by a single base station, i.e., the user's serving base station; only the scheduling decisions for the user are jointly determined by multiple base stations. Traditional centralized cooperative scheduling schemes not only require significant signaling exchange overhead but also increase user latency and reduce service quality. Furthermore, existing distributed algorithms are not satisfactory in terms of performance. Therefore, there is an urgent need for a low-overhead, high-performance distributed algorithm to implement cooperative scheduling.

[0003] Currently, research on using reinforcement learning for collaborative scheduling to improve user service quality is in its early stages. Xu Chunmei from Southeast University proposed a distributed algorithm based on multi-agent reinforcement learning to implement joint user scheduling and beam selection strategies. This strategy minimizes long-term average network latency costs while ensuring the instantaneous user service quality for each user. However, this method assumes that users have service needs in every time slot. In real-world scenarios, there are certain intervals between different user service needs, so this method cannot be directly applied to service scenarios with arrival intervals. Summary of the Invention

[0004] The purpose of this invention is to address the problems in the prior art by providing a multi-cell distributed scheduling method and system based on multi-agent reinforcement learning. Each base station trains its own deep Q-learning network (DQN), with inputs including channel state information and queue length information. The transformed instantaneous index is used as a reward for training, and the optimal set of scheduled user indices is output, thereby achieving the goal of low information overhead and high performance.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A multi-cell distributed scheduling method based on multi-agent reinforcement learning includes:

[0007] The 3GPP sets metrics for measuring multi-cell multi-user QoS based on pre-established multi-cell cellular network system models.

[0008] An optimization target model is established to maximize the system experience rate by measuring the QoS of multiple cells and multiple users;

[0009] Solve the optimization objective model to obtain the optimal strategy for multi-cell distributed scheduling.

[0010] Preferably, the multi-cell cellular network system model has J cells, and each cell has one equipped with N cells. t A base station with one transmitting antenna and K base stations equipped with N r N users with receiving antennas t ≥N r Each user shares the same frequency band with users in other cells. If a base station serves only a single user in each time slot, the base station uses SVD precoding. If a base station can serve multiple users in each time slot, the base station uses block diagonalization (BD) precoding. The received signal of the k-th user at the j-th base station is:

[0011]

[0012] in, This represents the channel state information from base station j′ to the k-th user served by base station j. This represents the precoding matrix sent by base station j to the k-th user. This represents the symbol vector sent by base station j to the k-th user. Let the noise at the k-th user served by base station j follow a pattern with mean 0 and variance σ. 2 The complex Gaussian distribution; P j,k,i This represents the power of the i-th stream allocated by base station j to the k-th user; The channel matrix from base station j to all serving users except the k-th user is expressed as follows:

[0013]

[0014] in, Let j be the total number of users scheduled by the current time slot base station, and let the number of antennas satisfy the constraints. For the channel Perform SVD decomposition using the following formula:

[0015]

[0016] in, yes An orthogonal basis, the equivalent channel from base station j to the k-th user served by base station j is Using the BD precoding algorithm Perform the following SVD decomposition:

[0017]

[0018] in, The first N r The column contains The right-hand singular phasor, defined as V j,k The first N r Listed as This yields the precoding matrix. diagonal elements N is the matrix r There are singular values; at the receiving end, the k-th user of the j-th base station uses a matrix. We then perform weighted reception, and the weighted received signal is:

[0019]

[0020] The first term on the right side of the above equation represents the useful signal, the second term represents inter-cell interference, and the third term represents noise. Therefore, the signal-to-interference-plus-noise ratio (SIR) of the i-th stream served by the k-th user at base station j is:

[0021]

[0022] Where σ 2 Let ||| represent noise power, and |||2 represent the 2-norm of the vector. This represents the interference matrix from base station j′ to the k-th user served by base station j. Representation matrix The element in the i-th row is the interference power vector of all flows from base station j′ to the i-th flow of the k-th user served by base station j; base station j performs power allocation through power watering, therefore,

[0023]

[0024] For the symbol (x) + If we take the larger value between 0 and x, then:

[0025]

[0026] Among them, P j Let J represent the maximum available power of base station j, and let J represent the rate at which the k-th user is served by base station j.

[0027]

[0028] Let Q j,k Let represent the queue length of the k-th user served by base station j. Then, the actual user rate is:

[0029]

[0030] This indicates that the user's actual rate does not exceed the user queue length;

[0031] The user queue update rules are as follows:

[0032]

[0033] Among them, a j,k (t) represents the number of data packets arriving in the t-th time slot for the k-th user served by base station j, a j,k The magnitude of (t) follows a Poisson distribution, and the arrival interval follows an exponential distribution.

[0034] As a preferred approach, the metrics for measuring multi-cell multi-user QoS are derived by combining the key performance indicators of 3GPP for visible user service quality with metrics for measuring latency.

[0035] The key performance indicator for 3GPP's Quality of Service (QoS) for users, IPThroughput, is calculated using the following expression:

[0036]

[0037] Where ThpVol is the size of the successfully transmitted data packet, and ThpTime is the time from the start of transmission to when the user queue is empty;

[0038] The 3GPP latency metric IPLatencyDL is calculated as follows:

[0039]

[0040] Where, N Samples It represents the number of samples, and TLat represents the time from when the data is first scheduled.

[0041] The metric for measuring QoS across multiple cells and multiple users is calculated using the following expression:

[0042]

[0043] Among them, T j,k (t) represents the total transmission time and waiting time of the kth user served by the base station j in the t-th time slot, that is, the time when the queue of the kth user served by the base station j in the t-th time slot is not empty.

[0044] As a preferred embodiment, the optimization target model established by measuring the QoS of multiple cells and multiple users to maximize the system experience rate includes:

[0045] The optimization goal is to maximize the system's user experience speed, and the calculation expression is as follows:

[0046]

[0047] Maximizing system experience speed is equivalent to minimizing total transmission latency:

[0048]

[0049] Without considering future data arrivals, let the initial length of the queue for the i-th user who has completed transmission in the current system be b. i The instantaneous rate of the i-th user who has finished transmitting in the n-th time slot is The average rate of the i-th user after transmission from the n1-th time slot to the n2-th time slot is: The time required for the i-th user to transmit after the (i-1)-th user has finished transmitting is t. i The time taken for the first user to complete the transmission is... The time taken for the second user to complete the transmission was The time taken for the i-th user to complete the transmission is Therefore, the total transmission delay of the entire system is:

[0050]

[0051] in, Therefore, the target model for optimization is:

[0052]

[0053] Preferably, in the step of solving the optimization objective model to obtain the optimal strategy for multi-cell distributed scheduling, if each base station only schedules a single user in each time slot, then each base station sets up a pair of DQN networks, including the following steps:

[0054] Action settings: Actions are the user indexes scheduled for each base station. The action settings for base station j are as follows:

[0055] a j (t)=I j (t)

[0056] Among them, I j (t) is the user index scheduled by base station j in the t-th time slot;

[0057] State settings: The state includes channel state information and queue length information for all users. The virtual transmission time of all users is used as the state input to the network. The state of base station j is set as follows:

[0058]

[0059] in, This represents the historical average transmission rate of the k-th user served by base station j. In the formula, the denominator of each element contains the channel state information of the corresponding user, and the numerator contains the queue length information.

[0060] Reward settings: The instantaneous reward is the sum of two parts: the first part is the independent reward for each base station action, and the second part is the sum of the independent rewards for all base station actions; the optimization objective is:

[0061]

[0062] When the t-th i After -1+1 time slots are transmitted, the optimization objective is:

[0063]

[0064] In the formula, the numerator is the queue length of the i-th user who completes transmission in the current time slot, and the denominator is the future t′. i -t′ i-1 The average transmission rate of the queue length of the i-th user who completes transmission in -1 time slots is determined by future decisions, and (JK-i+1) is the priority factor of the queue length of the i-th user who completes transmission.

[0065] Based on the optimization objective, the instantaneous reward of the j-th base station in the t-th time slot is determined to obtain the corresponding scheduling user's action.

[0066] Preferably, the network reward is determined by a greedy method, and the action that minimizes the optimization objective is selected in each time slot;

[0067] The denominator of the objective expression is optimized by approximating the historical average rate of users, as shown in the following expression:

[0068]

[0069] In the formula, w j,k The priority factor for the k-th user serving the j-th base station;

[0070] In the above formula As the tth i-1 The independent action reward for scheduling the k-th user at the j-th base station in the +1 time slot, and the instantaneous reward for the j-th base station in the t-th time slot are:

[0071]

[0072] In the formula, The priority factor for the aj(t)-th user serving the j-th base station. The rate of the aj(t)th user serving base station j. The historical average transmission rate of the aj(t)th user serving base station j.

[0073] Preferably, the user's priority factor w j,k Obtained through the following methods:

[0074] Calculate the estimated transmission time for all users whose buffers are not empty.

[0075] Sort the estimated transmission times of all users whose buffers are not empty in descending order;

[0076] The sorted user index is the user priority factor.

[0077] Preferably, in the step of solving the optimization objective model to obtain the optimal strategy for multi-cell distributed scheduling, if each base station serves multiple users, each base station is configured with... For DQN networks, the following steps are included:

[0078] Action settings: Given the number of users whose queues at current time slot base station j are not empty, each base station executes... There are several DQN networks, each outputting a scheduling user index. Combining all network actions yields the corresponding set of scheduling users for base station j. The network outputs for base station j are as follows:

[0079]

[0080] In the formula, a j,k (t) represents the action output by the k-th network at base station j, I j,k (t) is the scheduling user index of the k-th network output of base station j;

[0081] State settings: The state of a multi-user DQN network includes channel state information, queue length information, and the index of the user previously scheduled by the corresponding base station for all users. The state of the i-th DQN network of base station j is:

[0082]

[0083] in, This represents the historical average transmission rate of the k-th user served by base station j. In the formula, the denominator of each element contains the channel state information of the corresponding user, and the numerator contains the queue length information.

[0084] Reward settings: The reward for the k-th DQN network of base station j is:

[0085]

[0086] In the formula, The a-th base station serving the j-th base station j,k (t) priority factors for users, The a-th service of base station j j,k The rate of (t) users, The a-th service of base station j j,k (t) historical average transmission rate of users.

[0087] A multi-cell distributed scheduling system based on multi-agent reinforcement learning includes:

[0088] The QoS indicator setting module is used to set indicators for measuring the QoS of multiple cells and multiple users based on the pre-established multi-cell cellular network system model of 3GPP.

[0089] The optimization target model building module is used to build an optimization target model based on metrics that measure QoS across multiple cells and multiple users to maximize system experience rate.

[0090] The optimal scheduling strategy solution module is used to solve the optimization target model to obtain the optimal strategy for multi-cell distributed scheduling.

[0091] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the multi-cell distributed scheduling method based on multi-agent reinforcement learning.

[0092] Compared with the prior art, the present invention has at least the following beneficial effects:

[0093] Based on the 3GPP's established multi-cell cellular network system model, a metric for measuring QoS (Quality of Service) in multi-cell, multi-user scenarios is set, namely the system experience rate. Scheduling decisions for each time slot are determined locally by the base station. Base stations only need to exchange a small amount of scalar information, and there is no need for channel state information feedback between base stations and users. The total amount of interaction feedback in the entire system is independent of antenna configuration and is far less than that of centralized algorithms. The scheduling method of this invention sets up multiple network structures for each base station in multi-user scenarios. Simulations demonstrate that this network structure helps improve network performance. Simulations show that the performance of the method of this invention is far superior to other distributed algorithms, achieving over 90% of the performance of centralized algorithms in single-user scenarios and over 80% in multi-user scenarios.

[0094] Furthermore, when solving the optimization target model to obtain the optimal strategy for multi-cell distributed scheduling, this invention faces challenges because the future data arrival information and channel state information are unknown, making the optimization of the target model a non-convex optimization problem. This invention's multi-cell distributed scheduling method based on multi-agent reinforcement learning allows each base station to train its own deep Q-learning network (DQN). The input information includes channel state information and queue length information. The transformed instantaneous indicators are used as rewards for training, and the optimal set of scheduled user indices is output, achieving low information overhead and high performance. This invention's scheduling method first designs three network elements for system experience rate in single-user scheduling scenarios, then extends the three-element design to multi-user scheduling scenarios, and designs a more effective network structure for multi-user scheduling scenarios. Compared to other distributed intelligent algorithms, the three elements and network structure design proposed in this invention effectively improve network performance. Attached Figure Description

[0095] Figure 1 This is a schematic diagram of the downlink multi-cell cellular network system model structure with service arrival according to an embodiment of the present invention;

[0096] Figure 2 A schematic diagram of the grouping layer model for business operations;

[0097] Figure 3 Example diagram of the 3GPP indicator IPThroughput;

[0098] Figure 4 Example diagram of the 3GPP indicator IPLatencyDL;

[0099] Figure 5 The time T during which the k-th user queue serving base station j is not empty j,k Example image;

[0100] Figure 6 A schematic diagram of a single-user, multi-agent reinforcement learning network model;

[0101] Figure 7 A performance comparison curve for instantaneous reward and long-term reward;

[0102] Figure 8 A schematic diagram of a multi-user, multi-agent reinforcement learning network model;

[0103] Figure 9 A graph showing the system experience rate comparison of different reinforcement learning schemes in a single-user, low-noise scenario;

[0104] Figure 10 A graph showing the system experience speed comparison between the proposed solution and different heuristic solutions in a single-user low-noise scenario;

[0105] Figure 11 A graph showing the system experience speed comparison of different reinforcement learning schemes in a single-user high-noise scenario;

[0106] Figure 12 A graph showing the system experience speed comparison between the proposed solution and different heuristic solutions in a single-user high-noise scenario;

[0107] Figure 13 A graph showing the system experience speed comparison between the proposed solution and different heuristic solutions in a multi-user, low-noise scenario;

[0108] Figure 14 A graph showing the system experience speed comparison between the proposed solution and different heuristic solutions in a multi-user, high-noise scenario;

[0109] Figure 15 This is a comparison curve of system experience rates for base stations with multiple networks and a single network in a multi-user scenario. Detailed Implementation

[0110] The present invention will now be described in further detail with reference to the accompanying drawings.

[0111] like Figure 1 As shown, consider a downlink multi-cell cellular network system model with service arrival. The system has J cells, and each cell has one equipped with N t A base station with one transmitting antenna and K base stations equipped with N r N users with receiving antennas t ≥N r Each user shares the same frequency band with users in other cells, so inter-cell interference exists. If the base station serves only a single user in each time slot, the base station uses SVD precoding; if the base station serves multiple users in each time slot, the base station uses block diagonalization (BD) precoding, so there is no intra-cell interference. Therefore, the received signal of the k-th user at the j-th base station can be expressed as:

[0112]

[0113] in, This represents the channel state information from base station j′ to the k-th user served by base station j. This represents the precoding matrix sent by base station j to the k-th user. This represents the symbol vector sent by j to the k-th user. Let the noise at the k-th user served by base station j follow a pattern with mean 0 and variance σ. 2 The complex Gaussian distribution. P j,k,i This represents the power allocated by base station j to the i-th stream of the k-th user. (Definition) Let represent the channel matrix from base station j to all serving users except the k-th user, i.e.:

[0114]

[0115] in, Given the total number of users scheduled by base station j in the current time slot, the number of antennas must satisfy the constraint. For the channel SVD decomposition yields:

[0116]

[0117] in It can be seen as An orthogonal basis, the equivalent channel from base station j to the k-th user served by base station j is

[0118] Then, the BD precoding algorithm is used to... Perform SVD decomposition:

[0119]

[0120] in, The first N r The column contains The right-hand singular phasor, defined as V j,k The first N r Listed as From this, the precoding matrix can be obtained. diagonal elements N is the matrix r There are several singular values. At the receiving end, the k-th user of the j-th base station uses a matrix. We then perform weighted reception, and the weighted received signal is:

[0121]

[0122] The first term on the right-hand side of the above equation represents the useful signal, the second term represents inter-cell interference, and the third term represents noise. Therefore, the signal-to-interference-plus-noise ratio (SIR) of the i-th stream for the k-th user served by base station j can be expressed as:

[0123]

[0124] Where, σ 2 Let ||| represent noise power, and |||2 represent the 2-norm of the vector. This represents the interference matrix from base station j′ to the k-th user served by base station j. Representation matrix The element in the i-th row is the interference power vector for all flows from base station j′ to the i-th flow of the k-th user served by base station j. Base station j performs power allocation through power watering, therefore:

[0125]

[0126] Where (x) + This indicates taking the larger value between 0 and x:

[0127]

[0128] Where P j Let J represent the maximum available power of base station j, and let J represent the rate at which the k-th user is served by base station j.

[0129]

[0130] Let Q j,k Let represent the queue length of the k-th user served by base station j. Then, the actual user rate is:

[0131]

[0132] This means that the actual user rate does not exceed the user queue length. The user queue update rules are as follows:

[0133]

[0134] Where a j,k (t) represents the number of data packets arriving in the t-th time slot for the k-th user served by base station j. The packet size follows a Poisson distribution, and the arrival interval follows an exponential distribution, such as... Figure 2 As shown.

[0135] 3GPP has proposed IP Throughput, a key performance indicator that reflects the quality of user service, which is defined as:

[0136]

[0137] Where ThpVol is the size of the successfully transmitted data packet, and ThpTime is the time from the start of transmission until the user queue is empty, such as... Figure 3 As shown.

[0138] At the same time, 3GPP also proposed IPLatencyDL, an indicator for measuring latency, which is defined as:

[0139]

[0140] Where N Samples It represents the number of samples, and TLat represents the time from when the data is first scheduled. Figure 4 As shown.

[0141] Combining the two metrics proposed by 3GPP, this invention proposes a metric for measuring QoS in multi-cell, multi-user environments, namely, system experience rate, defined as:

[0142]

[0143] Where T j,k (t) represents the total transmission time and waiting time up to the time slot t is used by base station j to serve the k-th user, i.e., the time up to the time slot t is used by base station j to ensure the k-th user queue is not empty. For simplicity, we do not consider transmission failures. T j,k like Figure 5 As shown.

[0144] The optimization goal is to maximize the system experience speed:

[0145]

[0146] Maximizing system experience speed is equivalent to minimizing total transmission latency:

[0147]

[0148] Without considering future data arrivals, let the initial length of the queue for the i-th user who has completed transmission in the current system be b. i The instantaneous rate of the i-th user who has finished transmitting in the n-th time slot is The average rate of the i-th user after transmission from the n1-th time slot to the n2-th time slot is: The time required for the i-th user to transmit after the (i-1)-th user has finished transmitting is t. i The time taken for the first user to complete the transmission is... The time taken for the second user to complete the transmission was The time taken for the i-th user to complete the transmission is Therefore, the total transmission delay of the entire system is:

[0149]

[0150] in Therefore, the optimization problem is modeled as follows:

[0151]

[0152] As can be seen from the above optimization objective, it is impossible to obtain the optimal scheduling decision without knowing the data arrival information and channel state information at future times. Furthermore, the optimization objective includes optimization parameters in both the numerator and denominator, indicating a non-convex optimization problem. Therefore, optimizing such an objective function is quite challenging.

[0153] Currently, no distributed algorithm has been studied to maximize this metric in multi-cell scenarios. Classic scheduling algorithms such as Proportional Fair (PF), Max Weight (MW), and Shortest-Queue-First (SF) either require global channel state information, leading to excessive system overhead, or only use local channel state information, resulting in poor performance in multi-cell scenarios. The method in this invention achieves a trade-off between system overhead and performance through a multi-agent reinforcement learning algorithm, achieving near-centralized scheme performance with minimal information overhead.

[0154] If a base station schedules only a single user per time slot, then each base station sets up a pair of DQN networks, using a multi-agent reinforcement learning scheme for training and execution. The network structure under a single user is as follows: Figure 6 As shown.

[0155] Action settings: The action is the user index scheduled for each base station. The action for base station j is:

[0156] a j (t)=I j (t)

[0157] Where I j (t) is the user index scheduled by base station j in the t-th time slot.

[0158] State Setup: The state should include channel state information and queue length information for all users. Directly inputting all user channel state information and queue length information results in an excessively large input vector, making training difficult to converge. Here, the virtual transmission time of all users is used as the state input to the network, and the state of base station j is:

[0159]

[0160] in This represents the historical average transmission rate of the k-th user served by base station j. The denominator of each element in the state vector contains the channel state information for that user, and the numerator contains the queue length information.

[0161] Reward Settings: This embodiment of the invention sets two types of rewards: long-term rewards and instantaneous rewards. The long-term reward is set as the negative value of the number of users whose current queue is not empty, i.e.:

[0162]

[0163] Where || ||0 represents the zero norm of the vector;

[0164] Let the discount factor of the network be γ, and let γ = 1. The long-run discount payoff G of the network is:

[0165]

[0166] It is evident that maximizing the long-term discount benefit of the network is equivalent to minimizing the objective to be optimized.

[0167] The instantaneous reward is the sum of two parts: the first part is the independent reward for each base station's action, and the second part is the sum of the independent rewards for all base station actions. The independent reward for each base station's action is to more accurately determine the reward for that base station's action, while the sum of the rewards for all base station actions is to enable all base stations to collaboratively optimize the objective. The optimization objective of the method in this invention is:

[0168]

[0169] Without loss of generality, assume that the t′-th time... i-1 After the transmission of +1 time slot is completed, the optimization objective becomes:

[0170]

[0171] Where the numerator is the queue length of the i-th user who completes transmission in the current time slot, and the denominator is the future t′ i -t′ i-1 The average transmission rate of the queue length of the i-th user who completes transmission in time slot -1 is determined by future decisions, and (JK-i+1) is the priority factor of the queue length of the i-th user who completes transmission. A greedy method is used to determine the network reward, that is, selecting the action that minimizes the objective in each time slot, i.e., maximizing its difference. Since the denominator is determined by future decisions and its exact value cannot be obtained, the historical average rate of users is used to approximate the denominator, i.e.:

[0172]

[0173] Where w j,k w is the user's priority factor. j,k Obtained through the following methods:

[0174] 1) Calculate the estimated transmission time for all users whose buffers are not empty.

[0175] 2) Sort the estimated transmission times of all users whose buffers are not empty in descending order;

[0176] 3) The sorted user index is the user priority factor.

[0177] Therefore, the above formula will be... As the tth i-1The independent action reward for the j-th base station scheduling the k-th user in the +1 time slot, and the instantaneous reward for the j-th base station in the t-th time slot are:

[0178]

[0179] from Figure 7 It can be seen that instantaneous rewards perform better. This is because although the long-term discount benefit of long-term rewards is equivalent to the optimization problem, the physical meaning of its reward is the negative value of the number of users in the current non-empty queue, which does not provide clear guidance to the network. When the service just arrives, the number of users in the non-empty queue is large, and the reward value of any action is relatively small. When the service is about to finish transmitting, the number of users in the non-empty queue is small, and the reward value of any action is relatively large. This situation means that even though the long-term discount benefit of long-term rewards is equivalent to the optimization objective, its reward is not suitable for network training, resulting in a significant performance loss after convergence. Based on the above results and analysis, instantaneous rewards are adopted as the reward setting.

[0180] If each base station can serve multiple users, each base station is configured with... For the DQN network, a multi-agent reinforcement learning scheme is used for training and execution. The multi-user network structure is as follows: Figure 8 As shown, Figure 8 Only the execution process of the base station network is shown; the network training process is the same as that of the single-user scenario.

[0181] Action settings: Given the number of users whose queues at current time slot base station j are not empty, each base station executes... There are several DQN networks, each outputting a scheduling user index. Combining all network actions forms the set of scheduling users for that base station. The network outputs for base station j are:

[0182]

[0183] Where a j,k (t) represents the action output by the k-th network at base station j, I j,k (t) is the scheduling user index of the kth network output of base station j.

[0184] State settings: In addition to channel state information and queue length information for all users, the state of a multi-user DQN network should also include the index of users previously scheduled by that base station. Therefore, the state of the i-th DQN network for base station j is:

[0185]

[0186] Reward settings: Similar to the reward settings in a single-user network, the reward for the k-th DQN network of base station j is:

[0187]

[0188] To verify the performance of the distributed scheduling algorithm of this invention, the following simulations were performed:

[0189] The proposed algorithms use centralized and distributed schemes based on three classic scheduling algorithms: SF, PF, and MW, for comparison. The centralized scheme requires global CSI and queue status, while the distributed scheme only requires local CSI and queue status. The proposed algorithms require base stations to interact with each other across all the users they serve. and The total amount of interactive information in the entire system is 2J(J-1)K, which is independent of the base station and user antenna configurations. In contrast, the centralized algorithm requires the interaction of channel state information and queue information, resulting in a total amount of interactive information of J(J-1)K(N). t N r +1). Simulations were performed in a scenario where J=3 and K=6.

[0190] Case 1: Single-user simulations were conducted under high and low signal-to-noise ratios. To demonstrate the effectiveness of the proposed algorithm in reward and state settings, in Figure 9 The performance was compared under different reward and state settings. GS indicates using the estimated transmission time of all users as the state, DS indicates using only the estimated transmission time of the local user as the state. DR indicates using the independent action reward as the final reward, CR indicates using the sum of the independent action rewards of all base stations as the final reward, and GR indicates using the sum of the proposed two reward components as the final reward. Figure 9 It can be seen that the GS+GR and GS+CR schemes perform best under high noise conditions.

[0191] In low-noise environments, interference has a significant impact, making coordination between base stations crucial. Therefore, relying solely on local state DS or local reward DR will result in substantial performance degradation. Figure 10 As can be seen, under low noise conditions, the proposed algorithm outperforms other comparative schemes (excluding Con-SF and Con-PF) after a very short episode, and surpasses the Con-PF scheme after 4000 episodes. Upon convergence, it achieves 95% of the performance of the optimal Con-SF scheme. Simulations were also conducted under high noise conditions. Figure 11 It can be seen that the GS+GR scheme performs best; using DS or CR / DR in high-noise conditions will result in performance loss. From Figure 12It can be seen that, under high noise conditions, the proposed algorithm performs almost identically to Con-PF after convergence, and its performance surpasses that of other comparative schemes except Con-SF and Con-PF after a very short episode. After convergence, it can achieve 95% of the performance of the optimal Con-SF scheme.

[0192] Scenario 2: Multi-user simulations were performed in high and low noise scenarios. From Figure 13 and Figure 14 It can be seen that, under different signal-to-noise ratios in multi-user scenarios, the proposed algorithm can achieve 80% of the performance of the optimal Con-SP scheme after convergence.

[0193] When scheduling multiple users at a base station, a DQN network is set up at each base station. Figure 15 This demonstrates the performance of setting up multiple DQN networks and a single DQN network at each base station. When each base station is configured with a single DQN network, the network directly outputs the index of the scheduled user set. From Figure 15 This demonstrates the effectiveness of the multi-network structure.

[0194] Therefore, as can be seen from the above, the distributed scheduling method proposed in this invention can approach the performance of a centralized scheme with low information overhead.

[0195] Another embodiment of the present invention proposes a multi-cell distributed scheduling system based on multi-agent reinforcement learning, comprising:

[0196] The QoS indicator setting module is used to set indicators for measuring the QoS of multiple cells and multiple users based on the pre-established multi-cell cellular network system model of 3GPP.

[0197] The optimization target model building module is used to build an optimization target model based on metrics that measure QoS across multiple cells and multiple users to maximize system experience rate.

[0198] The optimal scheduling strategy solution module is used to solve the optimization target model to obtain the optimal strategy for multi-cell distributed scheduling.

[0199] Another embodiment of the present invention also proposes a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the multi-cell distributed scheduling method based on multi-agent reinforcement learning.

[0200] For example, the instructions stored in the memory can be divided into one or more modules / units. These modules / units are stored in a computer-readable storage medium and executed by the processor to complete the multi-cell distributed scheduling method based on multi-agent reinforcement learning according to the present invention. The one or more modules / units can be a series of computer-readable instruction segments capable of performing specific functions, which describe the execution process of the computer program in the server.

[0201] The electronic device may be a smartphone, laptop, PDA, or cloud server, among other computing devices. It may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the electronic device may also include more or fewer components, or combinations of certain components, or different components; for example, it may also include input / output devices, network access devices, buses, etc.

[0202] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0203] The memory can be an internal storage unit of the server, such as a hard drive or RAM. Alternatively, it can be an external storage device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory can include both internal and external storage units. The memory is used to store computer-readable instructions and other programs and data required by the server. It can also be used to temporarily store data that has been output or will be output.

[0204] It should be noted that the information interaction and execution process between the above-mentioned module units are based on the same concept as the method embodiment. For details on their specific functions and technical effects, please refer to the method embodiment section. They will not be repeated here.

[0205] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0206] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0207] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0208] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A multi-cell distributed scheduling method based on multi-agent reinforcement learning, characterized in that, include: The 3GPP sets metrics for measuring multi-cell multi-user QoS based on pre-established multi-cell cellular network system models. An optimization target model is established to maximize the system experience rate by measuring the QoS of multiple cells and multiple users; Solving the optimization objective model yields the optimal strategy for multi-cell distributed scheduling; The multi-cell cellular network system model mentioned above includes There are several residential communities, each equipped with one [facility / equipment]. Base stations with one transmitting antenna and Each equipment One user receiving antenna, Each user shares the same frequency band with users in other cells. If the base station serves only a single user in each time slot, the base station uses SVD precoding; if the base station can serve multiple users in each time slot, the base station uses block diagonalization BD precoding. The first base station The received signal for each user is: in, Indicates base station to base station The service Channel state information for each user Indicates base station Send to the Precoding matrix of each user Indicates base station Send to the Symbol vectors for each user, Indicates base station The service The noise at each user location follows a mean of 0 and a variance of . The complex Gaussian distribution; , Indicates base station Distributed to the first The first user's The power of each stream; Indicates base station Except for the first The channel matrix for all service users other than the individual user is expressed as follows: in, For the current time slot base station The total number of scheduled users and the number of antennas satisfy the constraints. , For the channel Perform SVD decomposition using the following formula: in, , , yes An orthogonal basis, base station to base station The service The equivalent channel for each user is Using the BD precoding algorithm Perform the following SVD decomposition: in, The former The column contains The right-hand singular phasor is defined The former Listed as Thus, the precoding matrix is ​​obtained. , , diagonal elements For matrix The singular value; at the receiving end, the singular value... The first base station User usage matrix We then perform weighted reception, and the weighted received signal is: The first term on the right side of the above equation represents the useful signal, the second term represents inter-cell interference, and the third term represents noise. From this, we obtain the base station... The service The first user's The signal-to-interference-plus-noise ratio (SIR) of each stream is: in Indicates noise power. Describes the 2-norm of a vector. Indicates base station to base station The service Interference matrix for each user Representation matrix The Row elements, for base stations All flows to the base station The service The first user's Interference power vector of the stream; base station Power distribution is achieved through power-driven water injection; therefore... For symbols Indicates taking 0 and If the larger value is: in, Indicates base station Maximum available power of base station The service The rate for each user is: set up Indicates base station The service If the queue length is [number] users, then the actual user rate is: This indicates that the user's actual rate does not exceed the user queue length; The user queue update rules are as follows: in, Indicates base station The service The user in the first The number of data packets arriving in each time slot. The magnitudes of the intervals follow a Poisson distribution, and the arrival intervals follow an exponential distribution. The metrics for measuring QoS in multi-cell, multi-user environments are derived by combining the key performance indicators of 3GPP for visible user service quality with metrics for measuring latency. The key performance indicators of 3GPP for display user service quality The calculation expression is as follows: in, This is the size of the successfully transmitted data packet. It is the time from the start of transmission to when the user queue is empty; The 3GPP metrics for measuring latency The calculation expression is as follows: in, It is the number of samples. It is the time from the time the data is processed until it is first scheduled. The metric for measuring QoS across multiple cells and multiple users is calculated using the following expression: in, Indicates up to the number Time-slot base station The service The total transmission time and waiting time for each user, i.e., up to the [number]th user. Time-slot base station The service The time during which each user queue is not empty; The optimization target model established by measuring the QoS of multiple cells and multiple users to maximize the system experience rate includes: The optimization goal is to maximize the system's user experience speed, and the calculation expression is as follows: Maximizing system experience speed is equivalent to minimizing total transmission latency: Without considering the future arrival of data, let the current system be the... The initial length of the user queue that has completed transmission is [number]. , No. The user who finished transmitting was in the first... The instantaneous rate of each time slot is , No. The user who has finished transmitting starts from the first... The time slot to the first The average rate of each time slot is , No. The user who finished transmitting was in the first... After each user's transmission is completed, the remaining transmission time is: The time taken for the first user to complete the transmission is... The time taken for the second user to complete the transmission was ;No. The time taken for each user to complete the transmission is Therefore, the total transmission delay of the entire system is: in, Therefore, the target model to be optimized is: In the step of solving the optimization objective model to obtain the optimal strategy for multi-cell distributed scheduling, if each base station schedules only a single user in each time slot, then each base station sets up a pair of DQN networks, including the following steps: Action settings: Actions are the user indexes scheduled by each base station. The action is set as follows: in, For base stations In the User indexes scheduled in time slots; State settings: The state includes channel state information and queue length information for all users, and the virtual transmission time of all users is used as the state input to the network and base station. The status is set as follows: in, Indicates base station The service The historical average transmission rate of each user is given by the formula, where the denominator of each element contains the channel state information of the corresponding user, and the numerator contains the queue length information. Reward settings: The instantaneous reward is the sum of two parts: the first part is the independent reward for each base station action, and the second part is the sum of the independent rewards for all base station actions; the optimization objective is: When the After the transmission of each time slot is completed, the optimization objective is: In the formula, the numerator is the number of times in the current time slot. The queue length of the users who have completed the transmission, with the denominator being the future... The first time slot The average transmission rate for each user who completes a transmission will be determined by future decisions. It is the first Priority factors for each user who has completed a transmission; Determine the first [item] based on the optimization objective. The first time slot The instantaneous reward of each base station is used to obtain the corresponding scheduling user's action; In the step of solving the optimization objective model to obtain the optimal strategy for multi-cell distributed scheduling, if each base station serves multiple users, each base station is configured with... For DQN networks, the following steps are included: Action settings: For the current time slot base station The number of users whose queues are not empty, each base station executes There are several DQN networks, each outputting a scheduling user index. Combining all network actions yields the corresponding set of users scheduled by the base station. All network outputs are: In the formula, For base stations The An action output by the network. For base stations The A network output index of scheduled users; State settings: The multi-user DQN network state includes channel state information, queue length information, and the index of the user previously scheduled by the corresponding base station for all users. The The state of each DQN network is: in, Indicates base station The service The historical average transmission rate of each user is given by the formula, where the denominator of each element contains the channel state information of the corresponding user, and the numerator contains the queue length information. Reward settings: Base station The The reward for each DQN network is: In the formula, For the first The first base station service Priority factors for each user For base stations The service Rate per user For base stations The service The historical average transmission rate of each user.

2. The multi-cell distributed scheduling method based on multi-agent reinforcement learning according to claim 1, characterized in that, The network reward is determined using a greedy method, and the action that minimizes the optimization objective is selected in each time slot. The denominator of the objective expression is optimized by approximating the historical average rate of users, as shown in the following expression: In the formula, For the first The first base station service Priority factors for each user; In the above formula As the first The first time slot The first base station scheduling Rewards for each user's unique action, the first The first time slot The instantaneous reward for each base station is: In the formula, For the first The first base station service Priority factors for each user For base stations The service Rate per user For base stations The service The historical average transmission rate of each user.

3. The multi-cell distributed scheduling method based on multi-agent reinforcement learning according to claim 2, characterized in that, The user's priority factor Obtained through the following methods: Calculate the estimated transmission time for all users whose buffers are not empty. ; Sort the estimated transmission times of all users whose buffers are not empty in descending order; The sorted user index is the user priority factor.

4. A multi-cell distributed scheduling system based on multi-agent reinforcement learning, characterized in that, include: The QoS indicator setting module is used to set indicators for measuring the QoS of multiple cells and multiple users based on the pre-established multi-cell cellular network system model of 3GPP. The optimization target model building module is used to build an optimization target model based on metrics that measure QoS across multiple cells and multiple users to maximize system experience rate. The optimal scheduling strategy solution module is used to solve the optimization target model to obtain the optimal strategy for multi-cell distributed scheduling. The multi-cell cellular network system model mentioned above includes There are several residential communities, each equipped with one [facility / equipment]. Base stations with one transmitting antenna and Each equipment One user receiving antenna, Each user shares the same frequency band with users in other cells. If the base station serves only a single user in each time slot, the base station uses SVD precoding; if the base station can serve multiple users in each time slot, the base station uses block diagonalization BD precoding. The first base station The received signal for each user is: in, Indicates base station to base station The service Channel state information for each user Indicates base station Send to the Precoding matrix of each user Indicates base station Send to the Symbol vectors for each user, Indicates base station The service The noise at each user location follows a mean of 0 and a variance of . The complex Gaussian distribution; , Indicates base station Distributed to the first The first user's The power of each stream; Indicates base station Except for the first The channel matrix for all service users other than the individual user is expressed as follows: in, For the current time slot base station The total number of scheduled users and the number of antennas satisfy the constraints. , For the channel Perform SVD decomposition using the following formula: in, , , yes An orthogonal basis, base station to base station The service The equivalent channel for each user is Using the BD precoding algorithm Perform the following SVD decomposition: in, The former The column contains The right-hand singular phasor is defined The former Listed as Thus, the precoding matrix is ​​obtained. , , diagonal elements For matrix The singular value; at the receiving end, the singular value... The first base station User usage matrix We then perform weighted reception, and the weighted received signal is: The first term on the right side of the above equation represents the useful signal, the second term represents inter-cell interference, and the third term represents noise. From this, we obtain the base station... The service The first user's The signal-to-interference-plus-noise ratio (SIR) of each stream is: in Indicates noise power. Describes the 2-norm of a vector. Indicates base station to base station The service Interference matrix for each user Representation matrix The Row elements, for base stations All flows to the base station The service The first user's Interference power vector of the stream; base station Power distribution is achieved through power-driven water injection; therefore... For symbols Indicates taking 0 and If the larger value is: in, Indicates base station Maximum available power of base station The service The rate for each user is: set up Indicates base station The service If the queue length is [number] users, then the actual user rate is: This indicates that the user's actual rate does not exceed the user queue length; The user queue update rules are as follows: in, Indicates base station The service The user in the first The number of data packets arriving in each time slot. The magnitudes of the intervals follow a Poisson distribution, and the arrival intervals follow an exponential distribution. The metrics for measuring QoS in multi-cell, multi-user environments are derived by combining the key performance indicators of 3GPP for visible user service quality with metrics for measuring latency. The key performance indicators of 3GPP for display user service quality The calculation expression is as follows: in, This is the size of the successfully transmitted data packet. It is the time from the start of transmission to when the user queue is empty; The 3GPP metrics for measuring latency The calculation expression is as follows: in, It is the number of samples. It is the time from the time the data is processed until it is first scheduled. The metric for measuring QoS across multiple cells and multiple users is calculated using the following expression: in, Indicates up to the number Time-slot base station The service The total transmission time and waiting time for each user, i.e., up to the [number]th user. Time-slot base station The service The time during which each user queue is not empty; The optimization target model established by measuring the QoS of multiple cells and multiple users to maximize the system experience rate includes: The optimization goal is to maximize the system's user experience speed, and the calculation expression is as follows: Maximizing system experience speed is equivalent to minimizing total transmission latency: Without considering the future arrival of data, let the current system be the... The initial length of the user queue that has completed transmission is [number]. , No. The user who finished transmitting was in the first... The instantaneous rate of each time slot is , No. The user who has finished transmitting starts from the first... The time slot to the first The average rate of each time slot is , No. The user who finished transmitting was in the first... After each user's transmission is completed, the remaining transmission time is: The time taken for the first user to complete the transmission is... The time taken for the second user to complete the transmission was ;No. The time taken for each user to complete the transmission is Therefore, the total transmission delay of the entire system is: in, Therefore, the target model to be optimized is: In the step of solving the optimization objective model to obtain the optimal strategy for multi-cell distributed scheduling, if each base station schedules only a single user in each time slot, then each base station sets up a pair of DQN networks, including the following steps: Action settings: Actions are the user indexes scheduled by each base station. The action is set as follows: in, For base stations In the User indexes scheduled in time slots; State settings: The state includes channel state information and queue length information for all users, and the virtual transmission time of all users is used as the state input to the network and base station. The status is set as follows: in, Indicates base station The service The historical average transmission rate of each user is given by the formula, where the denominator of each element contains the channel state information of the corresponding user, and the numerator contains the queue length information. Reward settings: The instantaneous reward is the sum of two parts: the first part is the independent reward for each base station action, and the second part is the sum of the independent rewards for all base station actions; the optimization objective is: When the After the transmission of each time slot is completed, the optimization objective is: In the formula, the numerator is the number of times in the current time slot. The queue length of the users who have completed the transmission, with the denominator being the future... The first time slot The average transmission rate for each user who completes a transmission will be determined by future decisions. It is the first Priority factors for each user who has completed a transmission; Determine the first [item] based on the optimization objective. The first time slot The instantaneous reward of each base station is used to obtain the corresponding scheduling user's action; In the step of solving the optimization objective model to obtain the optimal strategy for multi-cell distributed scheduling, if each base station serves multiple users, each base station is configured with... For DQN networks, the following steps are included: Action settings: For the current time slot base station The number of users whose queues are not empty, each base station executes There are several DQN networks, each outputting a scheduling user index. Combining all network actions yields the corresponding set of users scheduled by the base station. All network outputs are: In the formula, For base stations The An action output by the network. For base stations The A network output index of scheduled users; State settings: The multi-user DQN network state includes channel state information, queue length information, and the index of the user previously scheduled by the corresponding base station for all users. The The state of each DQN network is: in, Indicates base station The service The historical average transmission rate of each user is given by the formula, where the denominator of each element contains the channel state information of the corresponding user, and the numerator contains the queue length information. Reward settings: Base station The The reward for each DQN network is: In the formula, For the first The first base station service Priority factors for each user For base stations The service Rate per user For base stations The service The historical average transmission rate of each user.

5. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the multi-cell distributed scheduling method based on multi-agent reinforcement learning as described in any one of claims 1 to 3.