Energy-Efficient Uplink Resource Allocation Method Based on an Intelligent Framework under QoS Constraints

By adopting an intelligent framework in super-dense network combined with reinforcement learning and fractional planning, the problem of inter-cell interference management and resource allocation complexity in UDN is solved, and the resource allocation effect of high energy efficiency and high QoS is achieved.

CN116133142BActive Publication Date: 2025-06-24BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211610515.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-14
Publication Date
2025-06-24
Estimated Expiration
2042-12-14

AI Technical Summary

Technical Problem

In ultra-dense networks (UDN), the prior art is difficult to effectively manage inter-cell interference, resulting in reduced system energy efficiency, and the existing resource allocation methods are complex and difficult to scale, which cannot meet the users' Quality of Service (QoS) needs.

Method used

Using an intelligent framework-based approach, combined with reinforcement learning and fractional planning, a high-energy-efficient uplink resource allocation method under QoS constraints is designed. This method collects data through the wireless network module, performs interference modeling, and combines resource allocation modules to perform RB allocation and power allocation, and uses DDQN and fractional planning to interact with each other to optimize resource allocation.

Benefits of technology

It realizes that while ensuring user QoS, it reduces inter-cell interference, significantly improves network energy efficiency, and reduces the complexity of resource allocation, so that the method has better scalability and practical application value in super-intensive network scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116133142B_ABST
    Figure CN116133142B_ABST
Patent Text Reader

Abstract

The present invention discloses an energy-efficient uplink resource allocation method based on an intelligent framework under QoS constraints. The method includes the following steps: Step 1, energy efficiency definition; Step 2, energy efficiency optimization; Step 3, defining the effective energy efficiency of the system; Step 4, intelligent framework description; Step 5, Markov decision process modeling for RB allocation; Step 6, RB allocation based on DDQN; Step 7, power allocation based on fractional programming. The beneficial effects of the present invention are as follows: It can realize uplink user interference modeling based on the huge amount of wireless resource allocation data and wireless measurement data generated by network operation, and perform joint resource allocation based on the obtained interference model. Moreover, the present invention can perform resource allocation in real time according to network environment changes, reduce CCI and significantly improve network energy efficiency while ensuring the QoS of users, can reduce network fluctuations during training, and the obtained interference information reflects the large-scale fading situation of the network, eliminating the influence of small-scale fading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of wireless communication, energy efficiency, resource allocation, reinforcement learning, and fractional programming, and particularly relates to an energy-efficient uplink resource allocation method based on an intelligent framework under QoS constraints. Background Art

[0002] Currently, the number of connected devices and data traffic have increased exponentially. To meet the rapidly growing demand for transmission rates, the ultra-dense network (UDN) has become one of the key networking modes of 5G. Reusing the same spectrum resources among cells can greatly improve the spectrum efficiency and system capacity. However, the ultra-dense deployment of cells brings extremely high energy consumption and serious co-channel interference (CCI) problems, which will lead to a rapid increase in carbon emissions and a reduction in the energy efficiency (EE) performance of the system. In fact, the growth rate of carbon emissions in the information and communication technology (ICT) field is faster than the global total emissions and is twice that of the global emissions. Therefore, energy-efficient green communication has received wide attention.

[0003] Mining the interference relationship between users and studying the energy-efficient interference management and resource allocation (RA) method based on the obtained interference information can reduce the interference between cells and improve the system EE. Considering the user QoS (Quality of Service), the problem of jointly maximizing the system energy efficiency by resource block (RBs) and power allocation is a non-convex mixed-integer non-linear fractional programming problem. Reinforcement learning (RL) is model-free learning, and its agent obtains rewards by interacting with the environment to guide behavior. Therefore, RL is suitable for solving communication problems in complex interference environments. Since the energy efficiency is defined as the ratio of system throughput to power consumption, the fractional programming (FP) method can effectively solve the problem of maximizing energy efficiency. Considering interference identification, RL, and FP comprehensively, constructing an RA intelligent framework is an effective method to reduce complexity and obtain high solution accuracy at the same time. The reinforcement learning algorithm and fractional programming algorithm involved are briefly introduced below.

[0004] (1) Reinforcement learning algorithm:

[0005] RL is a branch of machine learning, consisting of an agent and an environment. The agent learns by interacting with the environment. The RL problem is modeled as a Markov Decision Process (MDP), represented by a quadruple <S, A, P, R>: S is the set of states, which describes the characteristics of the environment and the agent related to the research problem; A is the set of actions, including all actions that the agent can take; P is the state transition probability function; R is the reward function;

[0006] There are various RL algorithms. Q-learning (QL) is a common form, which constructs a Q-table to store the Q function. However, when the scale of the problem becomes larger, the state and action spaces grow rapidly, and storing the Q-table will incur a huge overhead. Deep Q-learning (DQL) emerged as a result. It fits the Q function by constructing a Deep Q-network (DQN): q(s, a; θ) ≈ q π (s, a), replacing the storage of the Q-table. However, DQL has the problem of high error, which can be solved by defining two DQNs and adopting the Double DQN (DDQN).

[0007] (2) Fractional programming algorithm:

[0008] The currently most common FP algorithm is the Dinkelbatch algorithm, which can transform the original non-convex fractional problem into a convex problem. Although it can obtain the same optimal solution as the original problem, the function value at the optimal solution has changed. Therefore, it is only applicable to problems with an outer fraction. For problems with an inner fraction, the quadratic transformation fractional programming method proposed in the existing technology [1] (K. Shen and W. Yu, "Fractional Programming for Communication Systems—Part I: Power Control and Beamforming," in IEEE Transactions on Signal Processing, vol. 66, no. 10, pp. 2616 - 2630, 15 May 15, 2018, doi: 10.1109 / TSP.2018.2812733.) can be used to solve it.

[0009] High - energy - efficient resource allocation is an important way to achieve green communication. High interference is a challenge faced by high - energy - efficient resource allocation. In the existing network, each cell base station manages and allocates wireless resources independently. To cope with inter - cell interference, the existing network conducts negotiation and signaling interaction between network elements and supplements it with enhanced technologies to make up for it to a certain extent. For example, information is exchanged through the X2 interface between base stations, and the inter - cell interference problem is solved by means of inter - cell interference coordination (ICIC) or enhanced ICIC (eICIC) technology; or by means of coordinated multiple points (CoMP) technology, interference is cooperatively processed, or interference is avoided, or interference is converted into useful signals between different base stations to provide higher data rates for users, thereby improving the utilization rate of the network. For ICIC and eICIC technologies, since they rely heavily on signaling exchange, the interference information they can transmit is extremely limited, resulting in poor granularity of transmitted interference information; and signaling transmission takes time, severely affecting timeliness. At the same time, a large number of adjacent cells in UDN will cause considerable signaling exchange overhead, affecting network performance. The CoMP technology, on the other hand, requires a large number of channel measurements and consumes a large amount of pilot resources; and it requires a large amount of computing resources to process and calculate signals, so this is not a suitable solution either. Soft frequency reuse (SFR) is also a common resource allocation scheme in the existing network. The spectrum in each cell is divided into two groups, the primary carrier and the secondary carrier. The primary carrier can be used for the entire cell and is orthogonal to each other in adjacent cells, and the secondary carrier is only used inside the cell. This is a common scheme that does not require accurate channel gain information, but due to reserving resources for edge users, the spectrum utilization rate is low. Some of the existing technologies still focus on the allocation of a single resource, and its optimization space is limited. Another part focuses on joint resource allocation, including user scheduling, RB allocation, and power allocation. Since energy efficiency is defined in fractional form, the energy - efficiency optimization problem is non - convex. The existing technologies first use fractional programming and Taylor expansion to transform the original problem into a convex problem, and then use convex optimization methods such as the Lagrangian method, the alternating direction method of multipliers, and the branch - and - bound method to solve it. The advantage of these analytical algorithms is strong interpretability and high accuracy, but they have high complexity, the solving process is cumbersome, and the scalability is not strong. In the ultra - dense network scenario with a large number of base stations and users, the disadvantages are particularly prominent. In contrast, machine - learning algorithms can adapt to the time - varying wireless communication environment and at the same time reduce the solving complexity, and have been widely used to solve communication problems. In the technologies related to EE maximization, various machine - learning methods such as supervised learning and reinforcement learning can already be seen, but some technologies do not consider the QoS requirements of users.

[0010] Regarding the problem of maximizing the system energy efficiency for joint resource allocation considering QoS, a clustering-based resource allocation method was proposed in the prior art [2] (K. Shen and W. Yu, "Fractional Programming for Communication Systems—Part I: Power Control and Beamforming," in IEEE Transactions on Signal Processing, vol. 66, no. 10, pp. 2616 - 2630, 15 May 15, 2018, doi: 10.1109 / TSP.2018.2812733.). The joint resource allocation is divided into two stages, but there is no interaction between the two stages, and the optimization ability is limited. In the prior art [3] (A. Kaur and K. Kumar, "Energy-Efficient Resource Allocation in Cognitive Radio Networks Under Cooperative Multi-Agent Model-Free Reinforcement Learning Schemes," in IEEE Transactions on Network and Service Management, vol. 17, no. 3, pp. 1337 - 1348, Sept. 2020, doi: 10.1109 / TNSM.2020.3000274), a resource allocation method based on multi-agent reinforcement learning was proposed. However, its state setting cannot accurately reflect the user QoS situation, it is difficult for the agent to make an optimal decision, and since it is based on QL, when the problem scale increases, its Q-table will increase rapidly, with high costs.

[0011] There are currently many technologies for energy efficiency optimization, but some technologies still focus on the allocation of a single resource, and their optimization space is limited. In the research related to joint resource allocation, some studies still focus on the single-cell scenario, and the optimization problem is relatively simple. The co-channel interference (CCI) in the multi-cell scenario will make the problem of maximizing EE more complex, and there will be two fractions, an outer fraction and an inner fraction, in the calculation formula of EE. One is the outer fraction corresponding to the ratio of throughput to power consumption according to the definition of EE, and the other is the inner fraction corresponding to the ratio of received power to interference noise according to the definition of SINR when calculating the throughput based on the Shannon formula. To reduce the complexity of the solution, joint resource allocation, such as joint resource block (RB) and power allocation, is divided into two stages. First, RB allocation is performed, and then power allocation is performed. The power allocation is completely based on the obtained RB allocation scheme, with little interaction between the two stages and limited optimization ability.

[0012] The advantage of the resource allocation scheme based on the parsing algorithm is strong interpretability and high precision. However, it often has high complexity, a cumbersome solution process, and poor scalability. In the ultra-dense network scenario with a large number of base stations and users, the disadvantages are particularly prominent.

[0013] The advantage of the resource allocation scheme based on the heuristic search algorithm is low complexity, but it has high randomness and low optimization accuracy. To improve the search accuracy, it is necessary to continuously increase the number of iterations.

[0014] The resource allocation scheme based on machine learning has low complexity and high optimization accuracy. In the resource allocation algorithm based on supervised learning, it is necessary to first use the parsing algorithm to obtain a large amount of training sets to train the neural network. Although it reduces the complexity of online computing, a large amount of computing is still required to obtain the training sets. In contrast, RL generates training data through continuous interaction between the agent and the environment, which can save computing time. As an important branch of RL, QL is a commonly used method for resource allocation. However, QL can only handle discrete action spaces. When performing power allocation, it is necessary to discretize the power, which reduces the optimization accuracy. DQL faces the same problem. To solve this problem, DDPG can be used to achieve continuous power allocation, but it increases the training difficulty.

[0015] A common problem with the above technologies is that they all assume that perfect channel state information (CSI) is known, which is difficult to achieve in the existing network. In the existing network, interference coordination and resource management schemes, such as ICIC and eICIC, are mostly based on signaling exchange to obtain interference information, and the information that can be transmitted is extremely limited. At the same time, a large number of adjacent cells in the UDN will cause considerable signaling exchange overhead, affecting network performance. The interference scheme based on cooperation (such as CoMP) requires a large number of channel measurements, consuming a large amount of pilot resources; and it requires a large amount of computing resources to process and calculate signals, with a high cost. Therefore, it is difficult for the resource management scheme in the existing network to achieve precise interference management and resource allocation, and it is also difficult to deploy the above technologies in the existing network.

[0016] In addition, there is a contradictory relationship between improving the QoS of cell-edge users and improving the system energy efficiency. The existing technologies consider the maximization of energy efficiency while ignoring the analysis of user QoS. Therefore, it is necessary to propose an index that can comprehensively consider both. Summary of the Invention

[0017] The purpose of the present invention is to provide an energy-efficient uplink resource allocation method based on an intelligent framework under QoS constraints that can overcome the above technical problems. The method of the present invention includes the following steps:

[0018] Step 1, Energy Efficiency Definition. Analyze the definition and calculation of energy efficiency in the UDN network. Let \(a(i)\) represent the serving base station of user \(u\) i \(\in U\). Then, through RB \(n\) k \(\in N\), the SINR of the signal received by base station \(a(i)\) from user \(u\) i

[0019] is expressed as:

[0020]

[0021] where is the transmission power of user \(u\) i on RB \(n\) k , is the channel gain from user \(u\) k to base station \(a(i)\) on RB \(n\) i , is the allocation indicator variable. When RB \(n\) k is allocated to user \(u\) i , then otherwise \(\sigma\) 2 is the additive Gaussian noise power. According to the Shannon formula, the transmission rate of user \(u\) i on RB \(n\) k is expressed as:

[0022]

[0023] where \(B\) is the bandwidth of each RB. The total transmission rate of user \(u\) i is expressed as:

[0024] The system energy efficiency is defined as the ratio of the system total transmission rate to the total power consumption, and is expressed as:

[0025]

[0026] where \(p\) c,i is the fixed circuit power consumption of user \(u\) i , and \(\mu\) is the reciprocal of the efficiency of the user power amplifier. It is assumed that all users are the same;

[0027] Step 2, Energy Efficiency Optimization. Regarding the minimum transmission rate constraint of users as QoS, jointly maximizing EE with RB and power allocation is expressed as:

[0028]

[0029]

[0030]

[0031]

[0032]

[0033]

[0034] Among them, is the RB allocation matrix, is the power allocation matrix, U j represents the set of users served by the same base station f j The formula (5) indicates that the maximum transmission power of user u i is

[0035] The formula (6) indicates that the user transmission power is non - negative. The formula (7) indicates that the RB allocation variable can only take 0 or 1. The formula (8) indicates the orthogonal allocation of RBs within the same cell, and an RB can only be allocated to one user within the same cell. The formula (9) is the QoS constraint, requiring the minimum transmission rate of the user to be

[0036] Step 3, Define the system effective energy efficiency (EEE). The new metric EEE for EE optimization is defined as the product of the average user QoS satisfaction rate and the system EE, expressed as:

[0037] η EEE = η EE * η satisfaction ……(10),

[0038] The user QoS satisfaction rate is defined as the ratio of the actual transmission rate of the user to the required minimum transmission rate. η satisfaction represents the average user QoS satisfaction rate, expressed as:

[0039]

[0040] EEE takes into account both the user QoS satisfaction rate and the system EE and is used to assist network training and evaluate the comprehensive performance of different resource allocation algorithms;

[0041] Step 4, Smart framework description. Design a smart framework for resource allocation. The smart framework consists of three modules: wireless network, performance prediction, and joint resource allocation. The functions of each part are as follows:

[0042] Wireless network module: During the operation of the wireless network, a lot of wireless data is generated, including resource allocation data and network measurement data. These data are collected and used for interference modeling in the performance prediction module. The resource allocation scheme obtained by the joint resource allocation module is implemented in the wireless network;

[0043] The performance prediction module obtains wireless data from the wireless network module, and uses the existing interference modeling scheme based on nonlinear regression to perform interference modeling, obtaining the information of the Interference to Signal Ratio (ISR) and the Noise to Signal Ratio (NSR), and further obtaining the information related to the channel gain ratio: and The obtained interference information is used to assist resource allocation. When training the RL agent, it evaluates the rewards of the current actions and predicts the rewards of different future actions, and is used as known information to solve the optimal allocation in power allocation;

[0044] The joint resource allocation module deploys a resource allocation algorithm, which includes two parts: the RB allocation based on reinforcement learning and the power allocation based on fractional programming, and the two interact continuously. When starting resource allocation, first allocate RBs one by one in the RB allocation part. Whenever an RB is allocated, power allocation is performed on all the allocated RBs. The power allocation module will obtain the current RB allocation status from the RB allocation module and pass back the power allocation scheme that maximizes the user QoS satisfaction rate. The RB allocation part allocates the next RB according to the satisfaction rate of the current user QoS and continues until all RBs are allocated. The result of power allocation is used to train the RL network. Finally, the obtained resource allocation scheme is deployed in the wireless network module;

[0045] Step 5, Markov decision process modeling of RB allocation:

[0046] Step 5.1, Allocate RBs, allocate RBs one by one and optimize the energy efficiency separately on each RB to reduce complexity. When allocating RBs, assume that the transmission power of the user is determined by the open-loop power control method, then RBn k The allocation problem is described as:

[0047]

[0048]

[0049]

[0050]

[0051] where is the energy efficiency on RBn k Formula (15) means that the QoS of the user is satisfied on the first k RBs. When allocating RBn k , RBn k′ , k′ < k have been allocated, that is, x k′ , k′ < k is known;

[0052] Step 5.2, model the RB allocation problem as a Markov decision process. In the RB allocation problem, set the central controller as the agent, which continuously interacts with the wireless environment. The elements in the quadruple <S, A, P, R> corresponding to the Markov decision process are defined as follows:

[0053] S: represents the set of states. The state of the agent at time t is set to be composed of and two parts, as shown in

[0054] the following formula:

[0055]

[0056] The first part of the state is the predicted reward value corresponding to different actions, expressed as:

[0057] |A| represents the number of actions. At each moment, not all actions can be executed. Then the reward value corresponding to action a is defined as:

[0058]

[0059] where, is the predicted reward value that can be obtained after executing action a. β is a negative value. Since QoS is taken into account, the second part of the state is defined as the QoS non - satisfaction rate of all users: ρ t,i represents the QoS non - satisfaction rate of user u i , that is, 1 minus the satisfaction rate of user u i , as shown in the following formula:

[0060]

[0061] where, R t,i is the transmission rate of user u i at time t, is the minimum transmission rate of user u i ;

[0062] A represents the set of actions of the agent, including all RB allocation actions. For the currently allocated RB, consider two actions. The first is that base station f j allocates the current RB to its served user u i for use, denoted as action The second is that base station f j does not allocate the current RB to any of its served users, denoted by action It is shown that the action set can be expressed as:

[0063]

[0064] When starting to allocate a certain RB, if the QoS of a user is not satisfied, randomly select one of them to allocate the current RB. At each moment, if a base station has previously selected the allocation action of the current RB, the related allocation actions will no longer be selected later. When some users among those associated with a certain base station satisfy the QoS, the base station will allocate the RB to the users whose QoS is not satisfied;

[0065] P: Represents the state transition probability function. In the current scenario, after the agent selects a certain execution action, it transfers to a definite state;

[0066] R: Represents the reward function. By observing the objectives and constraints of problem P1, the reward function reflects both the QoS situation of users and the system energy efficiency. The change in EEE before and after the action execution is used to evaluate the quality of the action selected at time t, as shown in the following formula:

[0067] As follows:

[0068] r t = η t,EEE -η t-1,EEE ……(20),

[0069] where η t,EEE represents the system EEE at time t. The interference model established by the performance prediction module will construct a virtual network to interact with the agent instead of the wireless network. According to the form of the interference information, when the performance

[0070] prediction module calculates and feeds back the reward, the SINR calculation formula is re-expressed as:

[0071]

[0072] Step 6, RB allocation based on DDQN. Use DDQN, which is a RL method, to solve the Markov decision process. DDQN contains two networks: the current Q network and the target Q network. The number of neurons in the input layer of the network is 3|F| + |U|, which is the state dimension, and the number of output neurons is 3|F|, equal to the number of actions. There is one hidden layer with 64 neurons. Before starting training, the parameters of the current Q network are randomly initialized, and the target Q network uses the same randomization parameters. The experience pool D is empty, the initial state is set as the RB not allocated to any user, the set of executable actions A′ is initialized as A, the training duration is set to 200 Transmission Time Intervals (TTIs). For each TTI, all resources will be reallocated. For each RB, at time t, the solution includes the following steps:

[0073] Step 6.1, select an action and observe the current state s t , when the RB has not been allocated to any user and there are still users not meeting the QoS, randomly select a non - satisfied user and allocate the current RB to it; otherwise, select an action a according to the ∈ - greedy policy t , as shown in the following formula:

[0074]

[0075] random(A′) represents randomly selecting an element from the set A′, c ∼ U[0,1], ∈ ∈ (0,1) is the probability of randomly selecting an action, and it satisfies where ∈ max and ∈ min represent the maximum and minimum values of ∈ respectively, and T is the total time;

[0076] Step 6.2, execute the action a selected in the previous step t , the agent's state transfers to s t+1 , and obtain the immediate reward r t+1 , for the base station related to the action a t , delete all actions related to it from A′;

[0077] Step 6.3, update the experience pool, and put the new experience sample (s t ,a t ,s t+1 ,r t+1 ) into the experience pool;

[0078] Step 6.4, update the policy. Every T1 moments, randomly select a sample set of Mini - batch size from the experience pool, and use the gradient descent method to update the current Q - network parameters. Every T2 moments, use the current Q - network parameters to update the target Q - network parameters;

[0079] Step 7, power allocation based on fractional programming. After allocating the k′ - th RB, re - perform power allocation based on the obtained allocation scheme of the k′ RBs, optimize the energy efficiency on each RB, and equally divide the minimum rate requirement among all RBs used by users to reduce the complexity. Set the maximum transmission power of each user on each RB to be the same, and represent the problem of maximizing the energy efficiency of power allocation on RBn k as:

[0080]

[0081]

[0082]

[0083]

[0084] Among them, is defined as is the maximum transmission power of user u i on RBn k Correspondingly, define is the minimum transmission rate of user u i on RBn k N i is the number of RBs allocated to user u i Using the above definitions and the interference identification information, can be re-expressed as:

[0085]

[0086] Then can be expressed as:

[0087]

[0088] There are fractions in formula (28) and formula (29). Formula (28) and formula (29) are non-convex. Then the optimization objective of problem (P2), that is, formula (24), is non-convex. The fraction in formula (28) satisfies that the numerator is concave and the denominator is convex. Using the fractional programming method, fix and introduce an auxiliary variable as shown in the following formula:

[0089]

[0090] Fix and re-express formula (28) as:

[0091]

[0092] Each user u i corresponds to an auxiliary variable All the auxiliary variables form a set Substitute the above formula into formula (29), then formula (29) will also satisfy the condition that the numerator is concave and the denominator is convex. Continue to use fractional programming, fix and Z k , and introduce an auxiliary variable y k : as shown in the following formula:

[0093]

[0094] Fix y k and Z k and transform formula (29) into:

[0095]

[0096] Similarly, the non-convex formula (27) is transformed into a convex form in the same way, and then problem (P2) is transformed into the following problem (P3), as shown in the following formula:

[0097]

[0098] s.t. Formula (25) Formula (26)

[0099]

[0100] The above problem is a convex problem and is solved using existing software packages such as CVX. Solving problem (P2) requires solving problem (P3). The algorithm for solving problem (P3) includes the following steps:

[0101] Step 7.1, Initialization is set to any feasible value, the iteration count variable m = 1 is initialized, and the maximum number of iterations M is set;

[0102] Step 7.2, Update Z according to formula (30) k ;

[0103] Step 7.3, Update y according to formula (32) k ;

[0104] Step 7.4, Fix Z k and y k , and update

[0105] by solving the convex optimization problem (P3)

[0106] Step 7.5, Update m: m = m + 1;

[0107] The method of the present invention has the following beneficial effects:

[0108] 1. The method of the present invention can realize uplink user interference modeling based on a large amount of wireless resource allocation data and wireless measurement data generated by network operation, and perform joint resource allocation based on the obtained interference model; the method of the present invention can perform resource allocation in real time according to network environment changes, reduce CCI and significantly improve network energy efficiency while ensuring the user's QoS (minimum user transmission rate).

[0109] 2. The method of the present invention divides resource allocation into two stages: RL-based RB allocation and FP-based power allocation. Different from the existing two-stage independent scheme, better optimization ability is obtained through continuous interaction between the two allocation stages. The power allocation result is continuously updated as the RB is allocated, and the power allocation result will affect the allocation of the next RB, and the two parts are continuously adjusted to achieve the goal of maximizing energy efficiency while meeting the user QoS.

[0110] 3. The method of the present invention takes offline resource allocation network training and online resource allocation as the process. Compared with the existing models that are trained offline based on the simulation environment, no online training is required when deployed in the real network. The user interference information is used to construct a virtual network identical to the real environment to assist the RL network training. After offline training, it can directly run online. In addition, it can reduce the network fluctuation during training, and the obtained interference information reflects the large-scale fading situation of the network, eliminating the influence of small-scale fading.

[0111] 4. A new metric EEE is proposed by the method of the present invention to better measure the algorithm performance and assist in training the RL network. It comprehensively considers the system energy efficiency and the average user QoS satisfaction rate. Compared with the scheme that only uses EE as the measurement metric, the method of the present invention has better practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0112] Figure 1 It is a schematic diagram of the 3GPP dual-strip UDN model of the method of the present invention;

[0113] Figure 2 It is a schematic diagram of the intelligent framework for energy-efficient resource allocation improved by the method of the present invention;

[0114] Figure 3 It is a flowchart of Algorithm 1 of the method of the present invention, i.e., DDQN-based RB allocation;

[0115] Figure 4 It is a flowchart of Algorithm 2 of the method of the present invention, i.e., FP-based power allocation;

[0116] Figure 5 It is a general flowchart of the joint resource allocation scheme of the method of the present invention;

[0117] Figure 6 It is a convergence curve graph of the method of the present invention;

[0118] Figure 7 It is a comparison graph of the algorithm performance from different network performance index perspectives of the method of the present invention;

[0119] Figure 8It is a performance comparison diagram of different algorithms when changing the QoS size of the method of the present invention;

[0120] Figure 9 It is a performance comparison diagram of different algorithms when changing the available RBs of the method of the present invention;

[0121] Figure 10 It is a performance comparison diagram of different algorithms when changing the number of cells of the method of the present invention;

[0122] Figure 11 It is a performance comparison diagram of different algorithms when changing the number of active users in a cell of the method of the present invention. Detailed implementation manners

[0123] The following describes the implementation manners of the present invention in detail with reference to the drawings. The method of the present invention includes the following steps:

[0124] Step 1, energy efficiency definition, as Figure 1 shown, in a small-scale wireless service hot spot area, to solve the huge demands of services and throughput, the operator will deploy a large number of network devices to form a UDN. There are two rows of rooms on both sides of the corridor, and each row has the same number of rooms. There is a base station in each room and the same number of users. Let F = {f1, f2,..., f |F|} represent the set of base stations,

[0125] U = {u1, u2,..., u |U|} represent the set of users, N = {n1, n2,..., n |N|} represent the set of RBs. Since the uplink adopts the Single-Carrier Frequency Division Multiple Access (SC-FDMA) technology, each base station allocates orthogonal RBs to the served users, and there is no intra-cell interference. Different base stations reuse the same RB resources, so there is serious CCI;

[0126] Analyze the definition and calculation of energy efficiency in the UDN network:

[0127] Set a(i) to represent the serving base station of user u i ∈U. Then, through RBn k ∈N, the SINR of the signal of user u i received by base station a(i) is expressed as:

[0128]

[0129] Among them, is the transmission power of user u i on RBn k ; For RBn k On it, user u i The channel gain to base station a(i), Is the allocation indication variable. When RBn k Is allocated to user u i , then Otherwise σ 2 Is the additive Gaussian noise power. According to Shannon's formula, the transmission rate of user u i On RBn k Is expressed as:

[0130]

[0131] Among them, B is the bandwidth of each RB. The total transmission rate of user u i Is expressed as:

[0132] The system energy efficiency is defined as the ratio of the system total transmission rate to the total power consumption and is expressed as:

[0133]

[0134] Among them, p c,i Is the fixed circuit power consumption of user u i , and μ is the reciprocal of the user power amplifier efficiency. It is assumed that all users are the same;

[0135] It should be noted that the above energy efficiency analysis does not depend on the UDN networking scenario as shown in Figure 1 . The high energy efficiency uplink resource allocation method based on the intelligent framework proposed in this application is applicable to any wireless communication networking model;

[0136] Step 2, energy efficiency optimization. Taking the user minimum transmission rate constraint as the QoS, jointly maximizing EE for RB and power allocation is expressed as:

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143] Among them, Is the RB allocation matrix, is the power allocation matrix, U j represents the set of users served by the same base station f j The formula (5) indicates that the maximum transmission power of user u i is

[0144] The formula (6) indicates that the user transmission power is non - negative. The formula (7) indicates that the RB allocation variable can only take 0 or 1. The formula (8) indicates the orthogonal allocation of RBs within the same cell, and one RB can only be allocated to one user within the same cell. The formula (9) is the QoS constraint, requiring the minimum transmission rate of the user to be

[0145] Step 3, Define the system effective energy efficiency (EEE), a new metric for EE optimization. The EEE is defined as the product of the average user QoS satisfaction rate and the system EE, expressed as:

[0146] η EEE = η EE *η satisfaction ……(10),

[0147] The user QoS satisfaction rate is defined as the ratio of the actual transmission rate of the user to the required minimum transmission rate. η satisfaction represents the average user QoS satisfaction rate, expressed as:

[0148]

[0149] EEE takes into account both the user QoS satisfaction rate and the system EE and is used to assist network training and evaluate the comprehensive performance of different resource allocation algorithms;

[0150] Step 4, Intelligent framework description. Design an intelligent framework for resource allocation. The framework is as Figure 2 shown. The framework consists of three modules: wireless network, performance prediction, and joint resource allocation. The functions of each part are as follows:

[0151] The wireless network module generates a lot of wireless data during the operation of the wireless network, including resource allocation data and network measurement data. These data are collected and used by the performance prediction module for interference modeling. The resource allocation scheme obtained by the joint resource allocation module is implemented in the wireless network;

[0152] The performance prediction module obtains wireless data from the wireless network module and uses the existing interference modeling scheme based on non - linear regression to perform interference modeling, obtaining information on the interference - to - signal ratio (ISR) and noise - to - signal ratio (NSR), and further obtaining information related to the channel gain ratio: and The obtained interference information is used to assist in resource allocation. When training the RL agent, it evaluates the rewards of the current action and predicts the rewards of different future actions, and is used as known information to solve the optimal allocation in power allocation;

[0153] The resource allocation algorithm is deployed in the joint resource allocation module, which includes two parts: RB allocation based on reinforcement learning and power allocation based on fractional programming, and the two interact continuously. When starting resource allocation, first allocate RBs one by one in the RB allocation part. Whenever an RB is allocated, power allocation is performed on all the allocated RBs. The power allocation module will obtain the current RB allocation status from the RB allocation module and pass back the power allocation scheme that maximizes the user QoS satisfaction rate. The RB allocation part allocates the next RB according to the satisfaction rate of the current user QoS and continues until all RBs are allocated. The result of power allocation is used to train the RL network. Finally, the obtained resource allocation scheme is deployed in the wireless network module;

[0154] The subsequent parts related to specific examples and performance evaluations are all based on Figure 1 As shown, taking the ultra-dense network and the OFDMA wireless system based on the 3GPP dual-strip model as a typical system scenario for specific elaboration;

[0155] Step 5, Markov decision process modeling for RB allocation:

[0156] Step 5.1, RB allocation problem description: Allocate RBs one by one and optimize the energy efficiency separately on each RB to reduce complexity. When allocating RBs, assume that the transmission power of the user is determined by the open-loop power control method. Then the RBn k allocation problem is described as:

[0157]

[0158]

[0159]

[0160]

[0161] where is the energy efficiency on RBn k Equation (15) represents meeting the user's QoS on the first k RBs. When allocating RBn k , RBn k′ , k′ < k have been allocated, that is, x k′ , k′ < k is known;

[0162] Step 5.2, model the RB allocation problem as a Markov decision process. In the RB allocation problem, set the central controller as the agent and continuously interact with the wireless environment. The elements in the quadruple <S, A, P, R> corresponding to the Markov decision process are defined as follows:

[0163] S: represents the state set. The state of the agent at time t is set to be composed of and two parts, as

[0164] shown in the following formula:

[0165]

[0166] The first part of the state is the predicted reward value corresponding to different actions, expressed as:

[0167] |A| represents the number of actions. At each moment, not all actions can be executed. Then the reward value corresponding to action a is defined as:

[0168]

[0169] where, is the predicted value of the reward that can be obtained after executing action a, and β is a negative value. Due to considering QoS, the second part of the state is defined as the QoS non - satisfaction rate of all users: ρ t,i represents the QoS non - satisfaction rate of user u i , that is, 1 minus the satisfaction rate of user u i , as shown in the following formula:

[0170]

[0171] where, R t,i is the transmission rate of user u i at time t, is the minimum transmission rate of user u i ;

[0172] A represents the action set of the agent, including all RB allocation actions. For the currently allocated RB, consider two actions. The first is that base station f j allocates the current RB to its served user u i for use, expressed as action The second is that base station f j does not allocate the current RB to any of its served users, with action

[0173] represented as It is shown that the action set A of the agent can be expressed as:

[0174]

[0175] When starting to allocate a certain RB, if the QoS of a user is not satisfied, first randomly select one of them to allocate the current RB. At each moment, if a certain base station has previously selected the allocation action of the current RB, the related allocation actions will no longer be selected later. When among the users associated with a certain base station

[0176] Some users meet the QoS, then the base station allocates the RB to the users whose QoS is not satisfied;

[0177] P: Represents the state transition probability function. In the current scenario, the agent selects a certain execution action

[0178] And then transfers to a definite state;

[0179] R: Represents the reward function. Observing the objectives and constraints of problem P1, the reward function reflects both the QoS situation of users and the system energy efficiency. The change in EEE before and after the action execution is used to evaluate the quality of the action selected at time t

[0180] As shown in the following formula:

[0181] r t = η t,EEE -η t-1,EEE ……(20),

[0182] where η t,EEE Represents the system EEE at time t. The interference model established by the performance prediction module will construct a virtual network to interact with the agent instead of the wireless network. According to the form of the interference information, when the performance

[0183] Prediction module calculates and feedbacks the reward, the SINR calculation formula is re-expressed as:

[0184]

[0185] Step 6, RB allocation based on DDQN. DDQN, a RL method, is used to solve the Markov decision process. DDQN contains two networks: the current Q-network and the target Q-network. The number of neurons in the input layer of the network is 3|F| + |U|, which is the state dimension, and the number of output neurons is 3|F|, equal to the number of actions. There is one hidden layer with 64 neurons. Before starting training, the parameters of the current Q-network are randomly initialized, and the target Q-network uses the same randomized parameters. The experience pool D is empty, and the initial state is set to the case where no RB is allocated to any user. The set of executable actions A′ is initialized as A. Set the training duration to 200 Transmission Time Intervals (TTIs). All resources will be reallocated for each TTI. For each RB, at time t, the solution includes the following steps, as Figure 3 shown:

[0186] Step 6.1, select an action and observe the current state s t . When no RB has been allocated to any user and there are still users not meeting the QoS, randomly select a non - satisfied user and allocate the current RB to it. Otherwise, select an action a according to the ε - greedy policy t , as shown in the following formula:

[0187]

[0188] random(A′) represents randomly selecting an element from the set A′, c ∼ U[0,1], and ε ∈ (0,1) is the probability of randomly selecting an action, satisfying where ε max 、ε min represent the maximum and minimum values of ε respectively, and T is the total time;

[0189] Step 6.2, execute the action a selected in the previous step t . The agent's state transfers to s t+1 , and an immediate reward r is obtained t+1 . For the base station related to the action a t , all actions related to it are removed from A′;

[0190] Step 6.3, update the experience pool. Put the new experience sample (s t ,a t ,s t+1 ,r t+1 ) into the experience pool;

[0191] Step 6.4, Update the policy. Every T1 moments, randomly select a sample set of Mini-batch size from the experience pool, and use the gradient descent method to update the current Q-network parameters. Every T2 moments, use the current Q-network parameters to update the target Q-network parameters;

[0192] Step 7, Power allocation based on fractional programming. After allocating the k'-th RB, re-perform power allocation based on the obtained allocation scheme of the k' RBs, optimize the energy efficiency on each RB, and equally divide the minimum rate requirement among all the RBs used by the users to reduce the complexity. Set the maximum transmission power of each user on each RB to be the same, and the power allocation on RBn k The problem of maximizing the energy efficiency P2 for the power allocation on it is expressed as:

[0193]

[0194]

[0195]

[0196]

[0197] where is defined as is the maximum transmission power of user u i on RBn k Correspondingly, define as the minimum transmission rate of user u i on RBn k N i is the number of RBs allocated to user u i Using the above definitions and the interference identification information, can be re-expressed as:

[0198]

[0199] Then can be expressed as:

[0200]

[0201] There are fractions in Equation (28) and Equation (29), and Equation (28) and Equation (29) are non-convex. Then the optimization objective of problem (P2), that is, Equation (24), is non-convex. The fraction in Equation (28) satisfies that the numerator is concave and the denominator is convex. Using the fractional programming method, fix and introduce an auxiliary variable as shown in the following equation:

[0202]

[0203] Fix Rewrite formula (28) as:

[0204]

[0205] Each user u i corresponds to an auxiliary variable All auxiliary variables form a set Substitute the above formula into formula (29), then formula (29) will also satisfy the condition that the numerator is concave and the denominator is convex. Continue to use fractional programming and fix with Z k , introduce an auxiliary variable y k : as shown in the following formula:

[0206]

[0207] Fix y k with Z k and transform formula (29) into:

[0208]

[0209] Similarly, the non-convex formula (27) is transformed into a convex form in the same way, then problem (P2) is transformed into the following problem (P3), as shown in the following formula:

[0210]

[0211] s.t. formula (25) formula (26)

[0212]

[0213] The above problem is a convex problem and is solved using existing software packages such as CVX. Solving problem (P2) requires solving problem (P3). The algorithm for solving problem (P3) includes the following steps, as Figure 4 shown:

[0214] Step 7.1, initialize to any feasible value, initialize the iteration count variable m = 1, and the maximum number of iterations M;

[0215] Step 7.2, update Z according to formula (30) k ;

[0216] Step 7.3, update y according to formula (32) k ;

[0217] Step 7.4, fix Z k , y k, update by solving the convex optimization problem (P3)

[0218] Step 7.5, update m: m = m + 1;

[0219] Step 7.6, determine whether the optimization objective formula (34) converges, and determine whether m satisfies m > M. If either of them is satisfied, end the iteration; otherwise, return to Step 7.2.

[0220] The performance evaluation and comparison of the method of the present invention are as follows. The performance of the algorithm is introduced from the aspects of the interference management ability, EEE, and average QoS satisfaction rate of the method of the present invention (Intelligent-Framework-Based Energy-Efficient Uplink Resource Allocation Method, hereinafter referred to as the "IFB" algorithm). The existing technologies [2], [3], and [5] (M. Qian, W. Hardjawana, Y. Li, B. Vucetic, X. Yang and J. Shi, "Adaptive Soft Frequency Reuse Scheme for Wireless Cellular Networks," in IEEE Transactions on Vehicular Technology, vol. 64, no. 1, pp. 118-131, Jan. 2015, doi: 10.1109 / TVT.2014.2321187.) are used. The cluster-based resource management algorithm (Cluster-based Energy-Efficient Resource Management Scheme, hereinafter referred to as the "CB" algorithm), the multi-agent reinforcement learning-based resource allocation algorithm (Multi-agent Reinforcement Learning Scheme, hereinafter referred to as the "MARL" algorithm), and the SFR algorithm are used as comparison schemes to reflect the performance superiority of the method of the present invention compared with the closest and best-performing existing scheme. At the same time, some system parameters are changed to verify the robustness of the method of the present invention.

[0221] Simulation parameter settings: Table 1 below summarizes the parameter table used in the simulation.

[0222] Table 1

[0223] Parameter Value Number of active users per cell 2-6 Number of cells per row 1-4 User QoS (minimum rate requirement) 0 - 3.2 Mbit / s System bandwidth 5 MHz Bandwidth per RB 180 kHz Number of RBs 25 Maximum transmission power of user 23 dBm (200 mW) Fixed circuit power consumption of user 20 dBm (100 mW) Path loss model (38.46 + 20lgd) dB Number of Mini - batch samples 32 Discount factor γ 0.5 Training duration 200 TTI

[0224] Set the number of users per cell to 2, the number of cells per row to 4, and the user QoS to the minimum transmission rates of 1 Mbit / s and 3 Mbit / s respectively. Other parameters are the same as those in Table 1. The convergence performance of the method described in the present invention is as follows:

[0225] The deep Q-network is the core of the method described in the present invention, and it is very important to verify its convergence. Set the training duration to 200 TTIs. In each TTI, the training network generates a new allocation scheme, and different minimum user transmission rates are set to observe the convergence under different QoSs. The convergence curves of different metrics such as EE, average user QoS satisfaction rate, and EEE are plotted respectively, as Figure 6 shown.

[0226] Figure 6 In (a), the wide shaded area represents the system EE and EEE curves, which coincide. The middle is the average user QoS satisfaction rate curve. Figure 6 In (b), from top to bottom are the system EE, average user QoS satisfaction rate, and EEE curves respectively. The upper and lower boundaries of all shaded areas represent the upper and lower bounds of 10 runs of the method described in the present invention, and the solid line in the middle represents the mean of 10 runs. It can be seen that after 150 TTIs of training, the system EE, average user QoS satisfaction rate, and EEE all reach a stable level, which verifies the convergence of the method described in the present invention. When the QoS is low, EEE is the same as EE because the QoS is easy to satisfy at this time, and the average user QoS satisfaction rate is always close to 1. When the QoS is high, it is difficult to satisfy all user QoSs, but the average user QoS satisfaction rate gradually increases with training, which also verifies the learning ability of the deep Q-network.

[0227] Set the number of users per cell to 2, the number of cells per row to 4, and the user QoS to the minimum transmission rate of 1 Mbit / s. Other parameters are the same as those in Table 1. The multi-faceted performance evaluation of the method described in the present invention is as follows:

[0228] Using multiple network performance metrics can more comprehensively analyze the performance of the method described in the present invention. Here, the average user QoS satisfaction rate, EEE, system throughput, interference management ability, cumulative distribution of user SINR on different RBs, and cumulative distribution of user throughput of different algorithms are compared.

[0229] A wireless communication network is an interference-limited system. Interference can be regarded as a kind of resource. Allocating the same RB to users with less mutual interference can effectively reduce CCI and improve EE. In the UDN scenario, the severe interference situation enables interference management to obtain a greater EE optimization space. The interference efficiency (IE) is used to evaluate the interference management ability of different algorithms. The definition of IE is:

[0230]

[0231] Among them, represents the total system throughput, represents the total system interference.

[0232] The comparison of the above indicators for different algorithms is as Figure 7 shown:

[0233] Figure 7 (a) shows the average user QoS satisfaction rate of different algorithms. It can be seen that the method described in the present invention satisfies the QoS of all users. The reason is that in the state setting of the DDQN algorithm, the degree of QoS dissatisfaction of users is considered, which will affect the RB allocation. Users with QoS dissatisfaction will have a higher allocation priority. In addition, when each RB starts to be allocated, a user who does not satisfy QoS will be randomly selected to allocate the current RB. The average user QoS satisfaction rate of the SFR algorithm is better than that of the CB and MARL algorithms because it reserves resources for edge users with poor channel quality. To better analyze the user QoS satisfaction situation, Figure 7 (f) shows the Cumulative Distribution Function (CDF) curve of user throughput. Compared with the CB and MARL algorithms, it can be seen that the proposed IFB algorithm and SFR algorithm have lower maximum user throughput, but higher minimum user throughput. The reason is that the method described in the present invention and the SFR algorithm sacrifice part of the performance of central users with good channel quality to satisfy the QoS of edge users. The method described in the present invention has better performance in this regard and sacrifices less performance of central users. Figure 7 (b) gives the comparison of the EEE performance of different algorithms. It can be seen that the IFB algorithm obtains the highest EEE, reaching 126.57%, 114.19%, and 116.39% of the SFR, CB, and MARL algorithms respectively. This verifies that the method described in the present invention can better optimize the EE performance while ensuring a high QoS satisfaction rate of users. Figure 7 (c) gives the comparison of system throughput. It can be seen that the method described in the present invention still obtains the highest value, which is 117.35%, 103.25%, and 113.46% of the SFR, CB, and MARL algorithms respectively. The performance improvement is not as high as that of EEE because high QoS satisfaction rate and high throughput are contradictory. The method described in the present invention can optimize system throughput and energy efficiency because it has good interference management ability, which is Figure 7 verified in Figure 7 (d) shows the IE value normalized with SFR as the benchmark. It can be seen that the method described in the present invention has the highest IE. Due to the reduction of interference, the SINR of users is correspondingly improved, as Figure 7as shown in (e).

[0234] Set the number of users per cell to 2 and the number of cells per row to 4. Other parameters are the same as those in Table 1. The method described in the present invention changes the user QoS to compare the performance of different algorithms:

[0235] To verify the applicability of the method described in the present invention, the QoS is continuously adjusted and the performance of different algorithms is compared. The results are as follows Figure 8 as shown.

[0236] Starting from 0, the QoS is gradually increased to 3.2 Mbit / s at a speed of 0.4 Mbit / s. The QoS of all users is the same. 0 represents no QoS requirement. Figure 8 (a) shows the average user QoS satisfaction rate. It can be seen that as the QoS increases, the average user QoS satisfaction rate of all algorithms gradually decreases. However, the IFB algorithm has the highest QoS satisfaction rate. Figure 8(b) shows the EEE performance. It can be seen that the EEE also gradually decreases as the QoS increases. The reason is that more resources need to be allocated to edge users to obtain a high QoS satisfaction rate, and the resources allocated to central users will be reduced accordingly, resulting in a decrease in the overall throughput. This is verified in Figure 8 (c). It can be seen that the throughput of central users decreases faster than the throughput increase of edge users.

[0237] Set the number of users per cell to 2 and the number of cells per row to 4. The user QoS takes the minimum rate of 1 Mbit / s. Other parameters are the same as those in Table 1. The method described in the present invention changes the available RBs to compare the performance of different algorithms:

[0238] To verify the applicability of the method described in the present invention, the available RBs are changed and the performance of different algorithms is compared. The results are as Figure 9 shown. Keeping the QoS at 1 Mbit / s, the available RBs are gradually increased from 7 to 25 at an interval of 3, and the changes in the performance of different algorithms are observed: Figure 9(a) shows the average user QoS satisfaction rate. It can be seen that as the available resources gradually increase, the average user QoS satisfaction rates of all algorithms gradually increase, and the proposed IFB algorithm always has the highest QoS satisfaction rate. In contrast, even when the available resources reach the maximum, it is difficult for the CB and MARL algorithms to meet the user QoS requirements. When the number of RBs exceeds 16, the average user QoS satisfaction rate of the CB algorithm no longer changes because it divides the resource allocation into two independent parts without interaction between them. When all the RB allocations are completed, power allocation is performed based on the obtained RB allocation scheme. As a result, when performing RB allocation based on the initial power, using 16 RBs has already met the QoS of the edge users, and the additional RB resources are all allocated to the central users. However, after power allocation, the QoS of the edge users is not satisfied. The reason why the MARL algorithm cannot meet the user QoS is that the state it sets is whether the user QoS is satisfied rather than the specific satisfaction rate of the user. The information obtained by the state is insufficient for resource allocation. The IFB algorithm solves these problems. The RB allocation and power allocation interact continuously. At the same time, for each allocated RB, the agent will track the user QoS satisfaction situation to achieve the purpose of meeting the user QoS with as few resources as possible. Figure 9 (b) shows the change of EEE. It can be seen that the EEE gradually increases as the available resources increase because more resources can increase the transmission rate. At the same time, it can be seen that the method described in the present invention can always obtain the maximum EEE.

[0239] Set the number of users per cell to 2, the user QoS takes the minimum rate of 1 Mbit / s, and other parameters are the same as in Table 1. The method described in the present invention changes the number of cells to compare the performance of different algorithms:

[0240] To verify the performance change of the method described in the present invention when the network scale changes, the number of cells is changed, and the results are as follows Figure 10 As shown, when keeping QoS = 1 Mbit / s and the number of RBs is 25, the number of cells in each row is changed, that is, the number of rooms. Since there are four rows of rooms, the total number of cells changes at intervals of 4. Figure 10 (a) shows the average user QoS satisfaction rate. It can be seen that the satisfaction rates of the proposed IFB algorithm and the SFR algorithm are always 1, while other algorithms can only be satisfied when the network scale is small. This shows that the method described in the present invention and the SFR algorithm can both adapt to different network scales in meeting the user QoS. From Figure 10(b) It can be seen that the ability of the SFR algorithm to optimize EE is limited. In addition, the EEE of all algorithms gradually decreases as the network scale increases because the network interference gradually increases and users need to increase the power to meet the QoS. When the network scale is small, the performance of the MARL algorithm is close to that of the proposed IFB algorithm, indicating that the MARL algorithm has good analysis ability for small-scale networks. When the network scale gradually becomes larger, the performance advantage of the method described in the present invention becomes gradually obvious.

[0241] Set the number of cells in each row to 4, and the minimum rate of user QoS is taken as 1 Mbit / s. Other parameters are the same as those in Table 1. The method described in the present invention changes the number of active users in the cell to compare the performance of different algorithms:

[0242] To verify the performance change of the method described in the present invention when the network load changes, the number of active users in the cell is changed, and the results are as Figure 11 shown. It can be seen from Figure 11 (a) that the average user QoS satisfaction rate of all algorithms decreases as the number of active users in each cell increases because the wireless resources become gradually tense. Among all algorithms, the performance of the MARL algorithm drops the fastest because its state and action space increase rapidly as the number of users increases, and it is difficult for QL to construct and update the Q table. In contrast, the IFB algorithm is based on DRL. It uses a deep Q network to fit the Q function and has better learning ability, so it can obtain a higher user QoS satisfaction rate. It can be seen from Figure 11 (b) that the network EEE decreases as the number of active users in each cell increases because the power consumption increases rapidly as the number of users increases, but the system throughput grows slowly. Compared with other algorithms, the IFB algorithm can obtain the highest EEE and verifies its better adaptability to different network loads.

[0243] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any change or replacement that can be easily thought of by those skilled in the art within the scope disclosed by the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. An energy-efficient uplink resource allocation method based on an intelligent framework under QoS constraints, characterized in that It includes the following steps: Step 1, Energy efficiency definition; Analyze the definition and calculation of energy efficiency in the UDN network; Let \(a(i)\) denote the serving base station of user \(u\) i \(\in U\). Then, through \(RB_n\) k \(\in N\), the SINR of the signal of user \(u\) received by base station \(a(i)\) i is expressed as: Among them, is the transmission power of user u i on RBn k ; is the channel gain from user u k on RBn i to base station a(i); is the allocation indication variable. When RBn k is allocated to user u i , then otherwise σ 2 is the additive Gaussian noise power. According to the Shannon formula, the transmission rate of user u i on RBn k is expressed as: where B is the bandwidth of each RB, and the total transmission rate of user u i is expressed as: The system energy efficiency is defined as the ratio of the total system transmission rate to the total power consumption, expressed as: Among them, p c,i is the fixed circuit power consumption of user u i , and μ is the reciprocal of the efficiency of the user power amplifier. It is assumed that all users are the same; Step 2, Energy efficiency optimization; Taking the minimum user transmission rate constraint as the QoS, jointly maximizing EE with RB and power allocation is expressed as: Among them, is the RB allocation matrix, is the power allocation matrix, U j represents the set of users served by the same base station f j Formula (5) indicates that the maximum transmission power of user u i is Equation (6) indicates that the user transmission power is non - negative. Equation (7) indicates that the RB allocation variable can only take 0 or 1. Equation (8) indicates the orthogonal allocation of RBs within the same cell, and one RB can only be allocated to one user within the same cell. Equation (9) is the QoS constraint, requiring the minimum user transmission rate to be Step 3, Define the system effective energy efficiency; The new metric EEE for EE optimization is defined as the product of the average user QoS satisfaction rate and the system EE, expressed as: η EEE = η EE * η satisfaction ……(10), The user QoS satisfaction rate is defined as the ratio of the actual transmission rate of the user to the required minimum transmission rate, η satisfaction represents the average user QoS satisfaction rate, expressed as: EEE takes into account both the user QoS satisfaction rate and the system EE and is used to assist network training and evaluate the comprehensive performance of different resource allocation algorithms; Step 4, Intelligent framework description; Design an intelligent framework for resource allocation. The intelligent framework consists of three modules: wireless network, performance prediction, and joint resource allocation. The functions of each part are as follows: Wireless network module: During the operation of the wireless network, a lot of wireless data is generated, including resource allocation data and network measurement data. These data are collected and used for interference modeling in the performance prediction module. The resource allocation scheme obtained by the joint resource allocation module is implemented in the wireless network; The performance prediction module obtains wireless data from the wireless network module, performs interference modeling based on the interference modeling scheme of nonlinear regression, obtains the interference-to-signal ratio and noise-to-signal ratio information, and further obtains the information related to the channel gain ratio: and The obtained interference information is used to assist resource allocation. When training the RL agent, it evaluates the rewards of the current action and predicts the rewards of different future actions, and is used as known information to solve the optimal allocation in power allocation; The resource allocation algorithm is deployed in the joint resource allocation module, including two parts: RB allocation based on reinforcement learning and power allocation based on fractional programming, and the two interact continuously. When starting resource allocation, first allocate RBs one by one in the RB allocation part. Whenever an RB is allocated, power allocation is performed on all the allocated RBs. The power allocation module will obtain the current RB allocation status from the RB allocation module and pass back the power allocation scheme that maximizes the user QoS satisfaction rate. The RB allocation part allocates the next RB according to the satisfaction rate of the current user QoS and continues until all RBs are allocated. The result of power allocation is used to train the RL network. Finally, the obtained resource allocation scheme is deployed in the wireless network module; Step 5, Markov decision process modeling for RB allocation; Step 6, RB allocation based on DDQN; Use DDQN, which is a RL method, to solve the Markov decision process. DDQN contains two networks: the current Q network and the target Q network. The number of neurons in the input layer of the network is 3|F|+|U|, which is the state dimension. The number of output neurons is 3|F|, equal to the number of actions. There is one hidden layer with 64 neurons. Before starting training, the parameters of the current Q network are randomly initialized, and the target Q network uses the same randomized parameters. The experience pool D is empty. The initial state is set to that no RB is allocated to any user. The set of executable actions A′ is initialized to A. Set the training duration to 200 transmission time intervals. All resources will be reallocated every TTI. For each RB, at time t, the solution includes the following steps: Step 6.1, select an action and observe the current state s t , when the RB has not been allocated to any user and there are still users not meeting the QoS, randomly select a user not meeting the requirements and allocate the current RB to it; otherwise, select the action a according to the ε-greedy policy t , as shown in the following formula: random(A′) represents randomly selecting an element from the set A′, θ is the network parameter, c ∼ U[0, 1], and ∈ ∈ (0, 1) is the probability of randomly selecting an action, satisfying where ∈ max , ∈ min represent the maximum and minimum values of ∈ respectively, and T is the total time; Step 6.2, execute the selected action a in the previous step t , the agent state transfers to s t+1 , and obtain the immediate reward r t+1 , for the base station associated with the action a t , delete all actions associated with it from A'; Step 6.3, update the experience pool and put the new experience samples (s t ,a t ,s t+1 ,r t+1 ) into the experience pool; Step 6.4, Update the policy. Every T1 moments, randomly select a sample set of Mini-batch size from the experience pool and update the parameters of the current Q network using the gradient descent method. Every T2 moments, update the parameters of the target Q network using the parameters of the current Q network; Step 7, Power allocation based on fractional programming; After each k′th RB is allocated, power will be reallocated based on the obtained k′th RB allocation scheme to optimize the energy efficiency on each RB and equally divide the minimum rate requirement into all RBs used by the user to reduce complexity. The maximum transmission power of each user on each RB is set to be the same. k The power allocation maximization energy efficiency problem P2 on is expressed as: Among them, is defined as is the maximum transmission power of user u i on RB n k Correspondingly, define is the minimum transmission rate of user u i on RB n k The lowest transmission rate on, N i is the number of RBs allocated to user u i Using the above definitions and the interference identification information, can be re-expressed as: Then Can be expressed as: There are fractions in Equation (28) and Equation (29). Since Equation (28) and Equation (29) are non-convex, the optimization objective of Problem (P2), that is, Equation (24), is non-convex. The fraction in Equation (28) satisfies that the numerator is concave and the denominator is convex. Using the fractional programming method, fix and introduce an auxiliary variable as shown in the following equation: Fixed Rewrite Equation (28) as follows: Each user u i corresponds to an auxiliary variable All the auxiliary variables form a set Substitute the above formula into formula (29), then formula (29) will also satisfy the condition that the numerator is concave and the denominator is convex. Continue to use fractional programming and fix and Z k , introduce an auxiliary variable y k , as shown in the following formula: Fix y k With Z k , transform formula (29) into: Similarly, the non-convex formula (27) is converted into a convex form in the same way, and then problem (P2) is transformed into the following problem (P3), as shown in the following formula: s.t. Formula (25) Formula (26) The above problem is a convex problem and is solved using a software package. Solving problem (P2) requires solving problem (P3); Step 5 includes the following steps: Step 5.1, RB allocation problem description: Allocate RBs one by one and optimize the energy efficiency individually on each RB to reduce the complexity. When allocating RBs, assume that the transmit power of the user is determined by the open-loop power control method. Then, for RB n k The allocation problem is described as: Among them, is the energy efficiency on RBn k The formula (15) represents that the QoS of users is satisfied on the first k RBs. When allocating RB n k , RB n k′ , k′ < k has been allocated, that is, x k′ , k′ < k is known; Step 5.2, model the RB allocation problem as a Markov decision process. In the RB allocation problem, set the central controller as the agent and continuously interact with the wireless environment. The elements in the quadruple <S, A, P, R> corresponding to the Markov decision process are defined as follows: S: represents the set of states, and the state of the agent at time t is set to be composed of and two parts, as shown in the following formula: The state of the first part Reward prediction values corresponding to different actions, expressed as: |A| represents the number of actions. At each moment, not all actions can be executed, and the reward value corresponding to action a is defined as: Among them, is the predicted value of the reward that can be obtained after the execution of action a. β is a negative value. Due to taking into account QoS, the second part of the state is defined as the QoS non - satisfaction rate of all users: ρ t,i represents the QoS non - satisfaction rate of user u i , that is, 1 minus the satisfaction rate of user u i , as shown in the following formula: where R t,i is the transmission rate of user u i at time t, and i is the minimum transmission rate of user u Let $\mathcal{A}$ denote the set of actions of the agent, including all RB allocation actions. For the currently allocated RB, two actions are considered. The first is that base station $f$ j allocates the current RB to its served user $u$, i denoted as action The second is that base station $f$ j does not allocate the current RB to any of its served users, denoted as action The set of actions can be expressed as: When starting to allocate a certain RB, if the QoS of a user is not satisfied, first randomly select one of them to allocate the current RB. At each moment, if a certain base station has selected the allocation action of the current RB before, the related allocation actions will no longer be selected later. When the QoS of some users associated with a certain base station is satisfied, the base station allocates the RB to the users whose QoS is not satisfied; P: represents the state transition probability function. In the current scenario, after the agent selects a certain execution action, it transfers to a definite state; R: represents the reward function. Observing the objectives and constraints of problem P1, the reward function reflects both the QoS situation of users and the system energy efficiency. The change in EEE before and after the action execution is used to evaluate the quality of the action selected at time t, as shown in the following formula: r t = η t,EEE - η t-1,EEE ……(20), Among them, η t,EEE represents the system EEE at time t. The interference model established by the performance prediction module will construct a virtual network to interact with the agent instead of the wireless network. According to the form of the interference information, when the performance prediction module calculates and feedbacks the reward, the SINR calculation formula will be re-expressed as: In step 7, the algorithm corresponding to solving problem (P3) includes the following steps: Step 7.1, Initialization Set it to any feasible value, initialize the iteration count variable m = 1, and the maximum number of iterations M; Step 7.2, update Z according to formula (30) k ; Step 7.3, update y according to formula (32) k ; Step 7.4, fix Z k , y k , update by solving the convex optimization problem (P3) Step 7.5, update m: m = m + 1; Step 7.6, determine whether the optimization objective formula (34) converges, and determine whether m satisfies m > M. If either of them is satisfied, end the iteration; otherwise, return to step 7.2.

Citation Information

Patent Citations

  • High-energy-efficiency resource optimizing method based on fractional programming and penalty function method

    CN103428767A

  • Method, device and system for service scheduling and service transfer rate controlling

    CN104160770A