A heterogeneous network resource energy efficiency optimization method based on deep reinforcement learning
An improved DDPG algorithm based on deep reinforcement learning solves the problems of resource allocation and energy efficiency in heterogeneous networks with large state spaces, achieving adaptive transmission power and subcarrier allocation, thereby improving system performance and energy efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV OF INFORMATION SCI & TECH
- Filing Date
- 2023-05-09
- Publication Date
- 2026-04-14
AI Technical Summary
Traditional reinforcement learning methods are inefficient when dealing with the large state spaces of heterogeneous networks, and cannot effectively optimize resource allocation and energy efficiency, resulting in severe network interference and high energy consumption.
An improved Deep Deterministic Policy Gradient (DDPG) algorithm based on deep reinforcement learning is constructed. It uses a multi-policy network Actor and a single-value network Critic to optimize transmission power and subcarrier allocation through Markov models and reward functions, thereby achieving adaptive resource allocation.
It improves the resource allocation efficiency of heterogeneous networks in dynamic environments, reduces network interference and energy consumption, and enhances system throughput and energy efficiency.
Smart Images

Figure CN116567667B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication technology, specifically but not limited to a method for optimizing the energy efficiency of heterogeneous network resources based on deep reinforcement learning. Background Technology
[0002] With the official commercialization of 5G, the development of wireless communication has entered a new stage. According to Ericsson's forecast, by 2022, the number of IoT devices will reach 29 billion, and by 2024, mobile data traffic will grow at a rate of 35% annually. The increasing demand for communication in society puts enormous pressure on current wireless networks and also places higher demands on communication technologies. The emergence of heterogeneous networks alleviates this pressure. Heterogeneous networks are a network architecture technology that can expand network coverage, improve spectrum utilization efficiency, and increase system capacity. To meet the needs of wireless communication, heterogeneous networking technology, under the premise of traditional cellular network coverage, adds multiple types of small base stations to cover specific areas, eliminating blind spots and coverage hotspots, reducing the distance between terminal devices and base stations, and enabling more devices to obtain better communication quality when accessing the network. Heterogeneous networks can deploy multiple small-coverage micro base stations or femtocells within a macro base station to improve spectrum utilization efficiency and network coverage. Specifically, micro base stations and femtocells can reuse and share the same spectrum with macro base stations, improving spectrum efficiency. Therefore, heterogeneous networks not only increase network capacity, but also meet the growing communication needs of users in future wireless networks, and reduce deployment costs.
[0003] However, the dense, random deployment of small base stations can lead to severe interference and high energy consumption. To reduce network interference, ensure Quality of Service (QoS) for users, and improve network energy efficiency, a framework for resource allocation and energy efficiency optimization is needed for heterogeneous networks. However, considering the real-world environment, users are mostly dynamic, and the vast state space of wireless networks, including location information, channel gain, and power, traditional reinforcement learning methods are not suitable. Traditional Q-learning methods in reinforcement learning suffer from enormous state spaces, resulting in huge Q-value tables for storage. Both searching and storing these tables consume significant time and space, greatly reducing the algorithm's convergence speed.
[0004] In view of this, a new method is needed to solve at least some of the above problems. Summary of the Invention
[0005] To address one or more problems in existing technologies, this invention proposes a resource energy efficiency optimization method for heterogeneous networks based on deep reinforcement learning. This method solves the problem that traditional algorithms cannot handle large state spaces and addresses the correlation that exists before and after each parameter update in the Actor-Critic neural network, thereby enhancing robustness.
[0006] The technical solution to achieve the purpose of this invention is as follows:
[0007] A method for optimizing the energy efficiency of heterogeneous network resources based on deep reinforcement learning, comprising:
[0008] S1. Establish a heterogeneous network model, initialize the communication environment and set the simulation environment area, including base station layout, number of base stations, number of user equipment and number of subcarriers. Among them, user equipment and base stations are associated based on the maximum signal-to-interference-plus-noise ratio (SINR) principle. The base station uses orthogonal frequency division multiple access to allocate resources to relevant user equipment.
[0009] S2. Determine the optimization objectives based on the signal-to-noise ratio of user equipment, network capacity, and energy efficiency;
[0010] S3. Introduce a Markov model to determine the agent, state space, action space, and reward function;
[0011] S4. Construct an improved Deep Deterministic Policy Gradient Algorithm (DDPG). The improved DDPG algorithm uses a multi-policy network (Actor) and a single-value network (Critic) for training and outputting transmission power and subcarrier allocation. The input of the Actor network is the current state of the agent, and the output is the subcarrier allocation policy and the transmit power on the subcarrier. The input of the Critic network is the agent's action and state, and the output is the action loss and the learned weight parameters.
[0012] S5. Set the number of training rounds and the number of training steps per round for the agent. Each agent continuously interacts with the set environment by improving the DDPG algorithm, optimizing and updating the network parameters, and obtaining the optimal resource allocation scheme.
[0013] Furthermore, the heterogeneous network resource energy efficiency optimization method based on deep reinforcement learning of the present invention includes a communication environment comprising a macro base station, N femto base stations and M user equipments, with K subcarriers. The M user equipments and N femto base stations are covered by the macro base station, wherein the N femto base stations follow a Poisson distribution and the M user equipments are uniformly and randomly distributed.
[0014] Furthermore, in the heterogeneous network resource energy efficiency optimization method based on deep reinforcement learning of the present invention, S2 determines the optimization objective and constraints, including:
[0015] S2-1. Determine the interference signal received by the user and calculate the signal-to-noise ratio information of the user equipment;
[0016] S2-2. Use Gaussian approximation to handle interference noise and calculate the network capacity and energy efficiency.
[0017] S2-3. The optimization objective is to ensure that the signal-to-noise ratio of the user equipment is greater than the minimum service quality requirement and to maximize energy efficiency.
[0018] Furthermore, in the heterogeneous network resource energy efficiency optimization method based on deep reinforcement learning of the present invention, the calculation of the user's signal-to-noise ratio information in S2-1 specifically includes:
[0019] S2-1-1. Assume that each user equipment can select at most one base station at any time. When the i-th user equipment selects and connects to the l-th base station, then: when l = n, a i,l (t) = 1; when l ≠ n, a i,l (t)=0, where n={1,...,N}, a i,l (t) represents the connection relationship between base station l and user equipment i at time t, i∈M, l∈N, N is the number of femtobase stations, and M is the number of user equipments;
[0020] S2-1-2, On the k-th subcarrier, the signal-to-noise ratio of user equipment i served by the l-th base station. for:
[0021]
[0022] Where k∈K, K is the number of subcarriers, a i,l This represents the connection coefficient between base station l and user equipment i. and They represent the lth and lth respectively. ′ The channel gain between a base station and a user on the k-th subcarrier, σ 2 Represented as Gaussian white noise, and They represent the lth and lth respectively. ′ The transmit power of a base station on the kth subcarrier.
[0023] Furthermore, in the resource energy efficiency optimization method for heterogeneous networks based on deep reinforcement learning of the present invention, the calculation of network capacity and energy efficiency in S2-2 specifically includes:
[0024] S2-2-1. On the k-th subcarrier, the capacity achieved by the macro base station and its associated user equipment. for:
[0025]
[0026] in, This represents the channel gain between macro base station h and user equipment i. This represents the transmit power of macro base station h on the k-th subcarrier. This represents the channel gain between the femtobase n and the user equipment i. σ represents the transmit power of the femtocell base station n on the k-th subcarrier. 2 This is represented as Gaussian white noise, where N is the number of femtocell base stations;
[0027] S2-2-2, Capacity achieved by the femtocell base station and its associated user equipment on the k-th subcarrier. for:
[0028]
[0029] in, This represents the channel gain between the femtobase n and the user equipment i. This represents the transmit power of the femtocell base station n on the k-th subcarrier;
[0030] S2-2-3, Capacity C in a network where macro base stations and femto base stations coexist. sum for:
[0031]
[0032] Where N is the number of femtocell base stations;
[0033] S2-2-4, Network energy efficiency η EE for:
[0034]
[0035]
[0036] Among them, P sum P represents the power consumption of all base stations per unit time in the network model. n Let P be the transmit power of the femtocell base station n. h P represents the transmit power of a macro base station. c This represents the power consumption of each circuit in both macro base stations and femto base stations.
[0037] Furthermore, in the resource energy efficiency optimization method for heterogeneous networks based on deep reinforcement learning of the present invention, the optimization objective and constraints in S2-3 specifically include:
[0038] The optimization objective is: argmaxη EE
[0039] The constraints include:
[0040] (a)
[0041] (b)
[0042] (c)
[0043] (d)a i,l (t)∈{0,1}
[0044] (e)P c =C
[0045] Where, η EE Indicates the network's energy efficiency. This represents the transmit power of the femtocell base station n on the k-th subcarrier. This represents the transmit power of the femtocell base station n on the k-th subcarrier. γ represents the signal-to-noise ratio of user equipment i served by the l-th base station on the k-th subcarrier. min Indicates the minimum quality of service requirement, a i,l (t) represents the connection relationship between base station l and user equipment i at time t, P c Let C be the power consumption of each circuit in the macro base station and the femto base station, and C is a constant.
[0046] Furthermore, in the heterogeneous network resource energy efficiency optimization method based on deep reinforcement learning of the present invention, S3 specifically includes determining the agent, state space, action space, and reward function, including:
[0047] 1) Treat the femtobase n as agents. Each agent updates its strategy independently. Each agent collects information from its own area and explores the network environment. Each agent selects its own subcarrier and transmit power. 1≤n≤N;
[0048] 2) State space S n,k (t) is defined as: S n,k (t)={M n (t),P n (t),I k (t),G n,k (t),a i,l (t)}, where M n (t) represents the number of users at the femtocell base station at time t; P n (t) represents the power of the femtocell base station at time t; I k (t)∈{0,1} represents the interference level from the macro base station on the k-th subcarrier at time t. Assume the minimum capacity requirement of the macro base station based on quality of service performance is α. h ,when Interference Level I kWhen (t) = 0, Interference Level I k (t) = 1; G n,k (t) represents the channel information of the femtobase n and the users at time t on the k-th subcarrier; a i,l (t) represents the connection relationship between the base station and the user at time t;
[0049] 3) The action space A is defined as: A = {k} n ,p n,k (t)}, where k n p represents the k-th subcarrier of the n-th base station, where k ∈ K; n,k (t) represents the power value on the k-th subcarrier of the n-th femtobase at time t, which is autonomously adjusted through algorithm learning;
[0050] 4) Reward function based on optimization objective Defined as the user's energy efficiency, that is:
[0051]
[0052] Here, β is a constant less than 0.
[0053] Furthermore, the improved DDPG algorithm in the heterogeneous network resource energy efficiency optimization method based on deep reinforcement learning of the present invention specifically includes:
[0054] Sampling phase:
[0055] The agent interacts with the communication environment, inputting its current state s(t) into the original Actor network μ(.|θ). μ ), the original Actor network μ(.|θ μ Choose action a(t) according to strategy μ: a(t) = μ(s(t)|θ μ )+N0, where N0 is noise;
[0056] After the agent performs action a(t), it receives environmental reward r(t), enters the next state s(t+1), obtains experience samples {s(t), a(t), r(t), s(t+1)} and stores them in the experience pool D until the storage amount reaches the threshold of the experience pool D.
[0057] Training phase: N empirical sample data are randomly sampled from the experience pool D to serve as the original Actor network. One training data of the original Critic network is denoted as {s′(t), a′(t), r′(t), s′(t+1)}.
[0058] Calculate the loss function Loss of the original Critic network, minimize the loss function using the gradient method, and update the Critic network parameters θ using backpropagation with the Adam optimizer. Q The loss function Loss is:
[0059]
[0060] Among them, y i =r i +γQ′(s i+1 ,μ′(s i+1 |θ μ′ )|θ Q′ ), where γ is a discount factor; the agent's objective function is defined as: J(θi) = E[Q μ (s,μ(s))], Maximize the objective function and update the Actor network parameters θ using the Adam optimizer. μ ;
[0061] The old target network parameters and the new corresponding network parameters are weighted and averaged to softly update the target Actor network and the target Critic network.
[0062]
[0063] Where τ is the discount factor.
[0064] Compared with the prior art, the present invention, employing the above technical solution, has the following technical effects:
[0065] 1. The resource energy efficiency optimization method for heterogeneous networks based on deep reinforcement learning of the present invention is more practically significant because it enables the acquisition of information through interaction with the environment and the attainment of long-term maximum benefits through continuous deep reinforcement learning in communication scenarios where the actual environment changes rapidly.
[0066] 2. The resource energy efficiency optimization method for heterogeneous networks based on deep reinforcement learning in this invention improves upon the traditional deep deterministic policy gradient algorithm. It uses a multi-Actor network and a single Critic network architecture to allocate transmission power and subcarriers. During the training phase, the Actor network needs to observe locally and generate actions based on the learned policy. The Critic network simultaneously provides feedback on the Actor network's policy based on global information. After the model training is completed, each agent performs distributed execution to obtain its own action output, which improves training speed and stability and enhances system performance.
[0067] 3. The heterogeneous network resource energy efficiency optimization method based on deep reinforcement learning of the present invention can improve the overall throughput of the system, enhance energy efficiency and reduce power consumption in heterogeneous wireless networks while ensuring the quality of service for users. Attached Figure Description
[0068] The accompanying drawings are provided to further illustrate the invention and, together with the description, serve to explain embodiments of the invention, but do not constitute a limitation thereof. In the drawings:
[0069] Figure 1 A schematic diagram of a communication system model for heterogeneous wireless networks is shown.
[0070] Figure 2 The Actor network and Critic network architectures of the improved DDPG algorithm of this invention are shown.
[0071] Figure 3 The overall flowchart of the improved DDPG algorithm of the present invention is shown.
[0072] Figure 4 A flowchart illustrating the resource allocation scheme completed by the training model of the present invention is shown.
[0073] Figure 5 A flowchart illustrating the training process of an example of the improved DDPG algorithm of the present invention is shown. Detailed Implementation
[0074] To further understand the present invention, preferred embodiments of the present invention are described below in conjunction with examples. However, it should be understood that these descriptions are only for further illustrating the features and advantages of the present invention, and not for limiting the scope of the claims of the present invention.
[0075] The description in this section pertains only to typical embodiments, and the present invention is not limited to the scope of the embodiments described. Combinations of different embodiments, substitution of some technical features in different embodiments, and substitution of similar or identical prior art with some technical features in the embodiments are also within the scope of the description and protection of the present invention.
[0076] To leverage heterogeneous networking technology by increasing the number of various types of small base stations, thereby shortening the distance between terminal devices and base stations and effectively improving system capacity to meet wireless communication demands, this invention proposes a resource energy efficiency optimization method for heterogeneous networks based on deep reinforcement learning. Furthermore, to address interference, improve energy efficiency and resource allocation, and handle large state spaces for more efficient wireless communication, this invention primarily proposes such a method. In the context of cellular heterogeneous networks, while ensuring user service quality, this method utilizes deep reinforcement learning to autonomously select subcarriers and transmit power, thereby improving overall system throughput, enhancing energy efficiency, and reducing power consumption.
[0077] The primary scenario considered is the downlink of a heterogeneous wireless network. This heterogeneous wireless network consists of one macro base station and multiple femto base stations. The macro base station is located at the center of the geographical area and contains user equipment, while the femto base stations are covered by the macro base station. The femto base stations follow a Poisson distribution, and users are uniformly and randomly distributed within the area, such as... Figure 1 As shown. Specifically, each femtobase can connect a maximum of 8 users to meet the minimum communication requirements of users. Users exceeding this limit will be associated with a macrobase. The base station uses an Orthogonal Frequency Division Multiple Access (OFDMA) scheme to allocate resources to its associated users. It is assumed that each base station can use all available resources. Each user can configure multiple subcarriers, and each subcarrier can serve at most one user in a time slot. The received signal of the user equipment in the downlink includes interference from the base station and thermal noise. It is assumed that the channel uses a specific fading model, namely a Rayleigh fading channel. This invention uses a deep reinforcement learning method to achieve autonomous selection of subcarriers and transmit power, maximizing the system's energy efficiency while meeting user QoS requirements.
[0078] The solution adopted in this invention is as follows: Figure 4 As shown, it includes the following steps:
[0079] Establish a network model, initialize the communication environment, and set the base station layout, the number of each base station, and the number of user devices.
[0080] Based on the signal-to-noise ratio of user equipment, network capacity, and energy efficiency, the optimization objectives and constraints are determined to ensure the quality of communication services.
[0081] Since the state space describing this problem is huge, a deep reinforcement learning approach is used to transform the optimization problem into a Markov decision process, which determines the agent, state space, action space, reward function, objective function, and loss function, and constructs an improved DDPG algorithm to solve the problem.
[0082] By continuously interacting with the set environment through the improved DDPG algorithm, the intelligent agent updates the parameters to optimize the network, enabling the agent to obtain the optimal strategy and achieve the goal of autonomously allocating the best resources.
[0083] Example 1
[0084] The communication environment in this example mainly consists of a macro base station (MBS) located at the center of a geographical area, with M user equipment (UE) and N femtocell base stations (FBS) covered by the macro base station. The femtocell base stations follow a Poisson distribution, and the users are uniformly and randomly distributed within the area. The communication system model is as follows: Figure 1 As shown. To make the objectives and advantages of the present invention clearer, the specific technical solutions will be further described below.
[0085] Step 1: Initialize the number of cellular users to M and the number of subcarriers to K. Users move randomly within the area at a speed of V. t m and the angle of movement If a user leaves the area, the user will reappear at the other end. The system will randomly assign these users with a probability of 0.1. The association between users and base stations is based on the principle of maximum signal-to-interference plus noise ratio (SINR).
[0086] Step 2.1: Determine the interference signal received by the user and calculate the user's signal-to-noise ratio information.
[0087] In this communication environment, because femtobase stations (FBSs) are deployed within the coverage area of macrobase stations (MBS), the interference originates from the interference generated by the femtobase stations on user equipment. Specifically, as follows... Figure 1 As shown.
[0088] Define binary variable a i,l (t), i∈M, l∈N, represents the connection relationship between the base station and the user at time t. When the i-th user equipment selects and connects to the l-th base station, a i,l (t) = 1, l = n and a i,l (t) = 0, l ≠ n, where n = {1, ..., N}. Assume that each user equipment can select at most one base station at any given time.
[0089] The signal-to-noise ratio (SINR) of a user served by the l-th base station on the k-th (k∈K) subcarrier is expressed as:
[0090]
[0091] Considering that all user equipment wants to meet the minimum Quality of Service (QoS) requirement γ min At the same time, it obtains maximum transmission capacity from its selected base station. Therefore, the SINR of the user equipment should not be less than the minimum quality of service requirement γ. min Where a i,l This represents the connection coefficient between the base station and the user. and They represent the lth and lth respectively. ′ The channel gain between a base station and a user on the k-th subcarrier, σ 2 It is represented as Gaussian white noise. and They represent the lth and lth respectively. ′ The transmit power of a base station on the kth subcarrier.
[0092] The system performance can be evaluated using capacity and energy efficiency formulas, where Gaussian approximation can be used to handle interference. The capacity achieved by the macro base station and its associated users on the k-th subcarrier can be expressed as:
[0093]
[0094] in, This represents the channel gain between the macro base station and the user on the k-th subcarrier. This represents the transmit power of the macro base station on the k-th subcarrier. This represents the channel gain between the femtobase and the user on the k-th subcarrier. σ represents the transmit power of a femtocell base station on the k-th subcarrier. 2 It is represented as Gaussian white noise.
[0095] On the k-th subcarrier, the capacity achieved by the femtocell base station and its associated users can be expressed as:
[0096]
[0097] in, This represents the channel gain between the femtobase and the user on the k-th subcarrier. σ represents the transmit power of a femtocell base station on the k-th subcarrier. 2 It is represented as Gaussian white noise.
[0098] Therefore, the capacity of a network in which macro base stations and femto base stations coexist can be expressed as:
[0099]
[0100] Energy efficiency is generally defined as the ratio of total throughput to total power consumption per unit time. In this invention, energy efficiency is expressed as:
[0101]
[0102] Among them, P sum The power consumption of all base stations per unit time in the system model can be expressed as:
[0103]
[0104] Among them, P c This represents the power consumption of each circuit in both macro base stations and femto base stations.
[0105] Step 2.2: Determine the optimization objective.
[0106] The objective of this invention is to ensure that the user's SINR is greater than γ. min At the same time, maximize energy efficiency η EE .
[0107] The optimization objective can be described as: argmaxη EE
[0108] The following restrictions apply: (a)
[0109] (b)
[0110] (c)
[0111] (d)a i,l (t)∈{0,1}
[0112] (e)P c =C
[0113] (a) and (b) indicate that the base station ensures that the transmit power allocated to the user meets the minimum receive power. (c) indicates the user's SINR requirement. (d) indicates that each user can be associated with at most one base station. (e) indicates that the power consumption of each circuit of the macro base station and femto base station is a constant.
[0114] Step 3: Construct a reinforcement learning model, introducing a Markov model to define the agent, state space, action space, and reward function. Train the model using a deep reinforcement learning algorithm to assign the optimal policy to each agent.
[0115] Agent: The femtobase station n (1≤n≤N) is used as an agent. Each agent updates its policy independently. Each agent can collect information and explore the network environment from its own area. Each agent can choose subcarriers and transmit power on its own.
[0116] State space: defined as S n,k (t)={M n (t), P n (t), I k (t), G n,k (t), a i,l (t)}.
[0117] Among them, M n (t) represents the number of users at the femtocell base station at time t; P n (t) represents the power of the femtocell base station at time t; I k (t)∈{0,1} represents the interference level from the macro base station on the k-th subcarrier at time t. Assume the minimum capacity requirement of the macro base station based on quality of service performance is α. h ,when Interference Level I k When (t) = 0, Interference Level I k (t) = 1; G n,k (t) represents the channel information of the femtocell base station n and the users at time t on the k-th subcarrier. i,l (t) represents the connection relationship between the base station and the user at time t.
[0118] Action Space: The agent has k (k∈K) subcarriers to choose from. An action is defined as the transmit power and subcarrier selected by the agent, A={k n p n,k (t)}. Where, k n p represents the k-th subcarrier of the n-th base station; n,k (t) represents the power value on the k-th subcarrier of the n-th femtobase at time t, which will be autonomously adjusted through algorithm learning.
[0119] Reward Function: When the agent performs an action and the constraints (a), (b), and (c) are met, it will receive a reward value. The reward function is defined based on the data rate and energy efficiency goals. According to the optimization objective, we define the reward as the user's energy efficiency:
[0120]
[0121] Here, β is a constant less than 0. As the training and learning process progresses, each agent will adjust itself in the direction of maximizing the reward.
[0122] Step 4: Employ deep reinforcement learning algorithms to enable the agent to learn autonomously. Reinforcement learning algorithms include policy-based methods, such as Policy Gradient (PG) and Actor-Critic (AC) algorithms; and value-based methods, such as Q-Learning (DQN). While these traditional algorithms are simple and easy to implement, they cannot handle large state spaces in practical applications, significantly reducing convergence speed and even leading to training instability. Therefore, this invention uses an improved version of the DDPG algorithm to address the shortcomings of the above algorithms. It uses a convolutional neural network to simulate the policy function and Q-function, and trains it using deep learning methods. This extends DQN to continuous action spaces or high-dimensional discrete values, incorporates the experience-based recycling method from DQN, and improves the fixed-target method in DQN by using soft target update. Random noise is also added to interact with the environment, increasing the system's robustness. Furthermore, this invention improves upon the previous one by using a multi-Actor and single-Critic architecture to allocate transmission power and subcarriers. During the training phase, the Actor needs to observe locally and generate actions based on the learned policy. At the same time, the Critic provides feedback on the Actor's policy based on global information. After the model training is completed, each agent performs distributed execution to obtain its own action output, which improves the training speed and stability and enhances the system performance.
[0123] The DDPG algorithm is based on the Actor-Critic framework and includes four networks: the original Actor network μ(.|θ) μ ) and the original Critic network Q(.|θ Q In addition, each network has its corresponding target network, the target Actor network μ′(.|θ μ′ ) and the target Critic network Q′(.|θ Q′ The Actor component can observe the network state s(t) at time t and take an action a(t). The agent then transitions to the next new state s(t+1) and receives an immediate reward r(t) after performing the action. The Critic is used to evaluate the quality of the actions generated by the Actor. The improved DDPG algorithm used in this invention includes multi-Actor network and single-Critic network structures, such as... Figure 2 As shown, the learning of an intelligent agent can be broadly divided into a sampling phase and a training and learning phase.
[0124] During the sampling phase, each agent continuously interacts with the environment, inputting its current state s(t) into the original Actor network. Based on the original Actor network's policy μ, an action a(t) is selected. Policy μ is a stochastic process based on the current original Actor policy and random U0 noise. The value of action a(t) is sampled from this stochastic process, the action is executed, and a reward r(t) is returned to the environment. Simultaneously, the environment enters the next state s(t+1). Critic is used to evaluate the quality of the action. The DDPG algorithm uses the experience retrieval method from DQN, storing the experience samples transition{s(t),a(t),r(t),s(t+1)} obtained from interactions with the environment into the experience pool D. Note that because the experience samples generated when the Actor interacts with the environment are highly correlated in time, these data sequences cannot be directly used for training, as this would lead to overfitting of the neural network and difficulty in convergence. In the DDPG algorithm, the Actor stores the experience sample data into the experience pool until the storage volume exceeds the set threshold of the experience pool. N experience sample data are randomly sampled from the experience pool D. The sampled data can be considered uncorrelated and will not produce overfitting.
[0125] During the training phase, the gradient of the original Critic network is calculated using the mean squared error (MSE). The loss function of the original Critic network is defined as follows:
[0126]
[0127] Among them, y i It can be seen as a "tag":
[0128] y i =r i +γQ′(s i+1 ,μ′(s i+1 |θ μ′ )|θ Q′ ) (Formula 8)
[0129] y i The calculation uses the target policy network μ′ and the target Q network Q′, which makes the learning process of the Q network parameters more stable and easier to converge. γ is a discount factor. The loss function is minimized by the gradient method, and the Critic network parameters θ are updated by backpropagation. Q .
[0130] Define the objective function J(θi) = E[Q] for the agent. μ (s,μ(s))]
[0131]
[0132] The Actor network updates θ by maximizing the cumulative expected return. μ .
[0133] The target network is updated using a soft update method, also known as Exponential Moving Average (EMA). This involves introducing a learning rate τ, taking a weighted average of the old target network parameters and the new corresponding network parameters, and then assigning this average to the target network. The main algorithm flow is as follows: Figure 3 As shown.
[0134] The specific process of the improved DDPG algorithm is shown in the table below.
[0135]
[0136]
[0137] Example 2
[0138] Step 1: Initialize the Communication Environment. The communication environment is simulated as a heterogeneous network architecture, including one macro base station and five femto base stations. The simulated environment area is a 600m × 600m rectangular area. The macro base station is located at the center of the area, with a coverage radius of 300m, and the femto base stations have a coverage radius of 30m. Sixty users move at a speed of 36km / h in random directions. If a user leaves the area, they will reappear at the other end, and these users are randomly reassigned with a probability of 0.1. The association between users and base stations is based on the maximum SINR principle.
[0139] Step 2: Set the maximum transmit power of the macro base station to 46dBm, the maximum transmit power of the femtobase to 30dBm, and the minimum transmit power to 20dBm. The number of subcarriers is 64, the minimum SINR for users is -6dB, the discount factor τ is 0.001, the discount factor γ is 0.9, the path loss from the cellular user to the macro base station is 34 + 40lg(d[km]), the bandwidth is 10MHz, and the Gaussian white noise σ 2 = -114dBm, noise power density N0 is -174dBm / Hz, and dropput rate is 0.8.
[0140] Step 3: Initialize network parameters. The improved DDPG algorithm's Actor network model consists of four one-dimensional convolutional layers and two fully connected layers. The input is the current agent's state, and the output includes the subcarrier allocation strategy and the transmit power on the subcarrier. The Critic network model consists of four two-dimensional convolutional layers and two fully connected layers. The input is the agent's action and state, and the output is the action loss and the learned weight parameters. The experience pool capacity is set to 5000, and the batch size for updates is set to 32.
[0141] Step 4: Set the number of training rounds for the agent to Episode = 10000, and the number of training steps per round to Step = 100. Every 50 steps, use the Adam optimizer to optimize the neural network parameters. Record the rewards obtained during the agent's training process. The agent continuously optimizes its strategy based on the proposed algorithm, ultimately obtaining the optimal resource allocation scheme. Finally, apply the trained model to real-world scenarios, allowing users to independently choose the optimal resource allocation scheme to improve energy efficiency. The main training process is as follows: Figure 5 As shown.
[0142] The description and application of the present invention herein are illustrative and not intended to limit the scope of the invention to the embodiments described above. The effects or advantages described in the specification may not be apparent in actual experimental cases due to uncertainties in specific conditions or other factors, and such descriptions are not intended to limit the scope of the invention. Variations and modifications to the embodiments disclosed herein are possible, and various substitutions and equivalents of the components in the embodiments are well known to those skilled in the art. It should be understood by those skilled in the art that the invention can be implemented in other forms, structures, arrangements, proportions, and with other components, materials, and parts without departing from the spirit or essential characteristics of the invention. Other variations and modifications can be made to the embodiments disclosed herein without departing from the scope and spirit of the invention.
Claims
1. A method for optimizing the energy efficiency of heterogeneous network resources based on deep reinforcement learning, characterized in that, include: S1. Establish a heterogeneous network model, initialize the communication environment and set the simulation environment area, including base station layout, number of base stations, number of user equipment and number of subcarriers. Among them, user equipment and base stations are associated based on the maximum signal-to-interference-plus-noise ratio (SINR) principle. The base station uses orthogonal frequency division multiple access to allocate resources to relevant user equipment. S2. Based on the signal-to-noise ratio of the user equipment Network capacity and energy efficiency η EE Determine the optimization objective; S3. Introduce a Markov model to determine the agent, state space, action space, and reward function; S4. Construct an improved Deep Deterministic Policy Gradient Algorithm (DDPG). The improved DDPG algorithm uses a multi-policy network (Actor network) and a single-value network (Critic network) for training and outputting transmission power and subcarrier allocation. The input of the Actor network is the current state of the agent, and the output is the subcarrier allocation policy and the transmit power on the subcarrier. The input of the Critic network is the agent's action and state, and the output is the action loss and the learned weight parameters. S5. Set the number of training rounds and the number of training steps per round for the agent. Each agent continuously interacts with the set environment by improving the DDPG algorithm, optimizing and updating network parameters, and obtaining the optimal heterogeneous network resource allocation scheme.
2. The method for optimizing the energy efficiency of heterogeneous network resources based on deep reinforcement learning according to claim 1, characterized in that, The communication environment includes one macro base station, N femtobase stations and M user equipments, with K subcarriers. The macro base station covers the M user equipments and N femtobase stations. The N femtobase stations follow a Poisson distribution, and the M user equipments are uniformly and randomly distributed.
3. The method for optimizing the energy efficiency of heterogeneous network resources based on deep reinforcement learning according to claim 1, characterized in that, S2 defines the optimization objective and constraints, including: S2-1. Determine the interference signal received by the user equipment and calculate the signal-to-noise ratio information of the user equipment; S2-2. Use Gaussian approximation to handle interference noise and calculate the network capacity and energy efficiency. S2-3. The optimization objective is to ensure that the signal-to-noise ratio of the user equipment is greater than the minimum service quality requirement and to maximize energy efficiency.
4. The method for optimizing the energy efficiency of heterogeneous network resources based on deep reinforcement learning according to claim 1, characterized in that, The calculation of the user's signal-to-noise ratio information in S2-1 specifically includes: S2-1-1. Assume that each user equipment can select at most one base station at any time. When the i-th user equipment selects and connects to the l-th base station, then: when l = n, a i,l (t) = 1; when l ≠ n, a i,l (t)=0, where n={1,...,N}, a i,l (t) represents the connection relationship between base station l and user equipment i at time t, i∈M, l∈N, N is the number of femtobase stations, and M is the number of user equipments; S2-1-2, On the k-th subcarrier, the signal-to-noise ratio of user equipment i served by the l-th base station. for: Where k∈K, K is the number of subcarriers, a i,l This represents the connection coefficient between base station l and user equipment i. and They represent the lth and lth respectively. ′ The channel gain between a base station and a user on the k-th subcarrier, σ 2 Represented as Gaussian white noise, and They represent the lth and lth respectively. ′ The transmit power of a base station on the kth subcarrier.
5. The method for optimizing the energy efficiency of heterogeneous network resources based on deep reinforcement learning according to claim 1, characterized in that, The capacity and energy efficiency of the computational network in S2-2 specifically include: S2-2-1. On the k-th subcarrier, the capacity achieved by the macro base station and its associated user equipment. for: in, This represents the channel gain between macro base station h and user equipment i. This represents the transmit power of macro base station h on the k-th subcarrier. This represents the channel gain between the femtobase n and the user equipment i. σ represents the transmit power of the femtocell base station n on the k-th subcarrier. 2 This is represented as Gaussian white noise, where N is the number of femtocell base stations; S2-2-2, Capacity achieved by the femtocell base station and its associated user equipment on the k-th subcarrier. for: in, This represents the channel gain between the femtobase n and the user equipment i. This represents the transmit power of the femtocell base station n on the k-th subcarrier; S2-2-3, Capacity C in a network where macro base stations and femto base stations coexist. sum for: Where N is the number of femtocell base stations; S2-2-4, Network energy efficiency η EE for: Among them, P sum P represents the power consumption of all base stations per unit time in the network model. n Let P be the transmit power of the femtocell base station n. h P represents the transmit power of a macro base station. c This represents the power consumption of each circuit in both macro base stations and femto base stations.
6. The method for optimizing the energy efficiency of heterogeneous network resources based on deep reinforcement learning according to claim 1, characterized in that, The optimization objective and constraints in S2-3 specifically include: The optimization objective is: argmaxη EE The constraints include: (d)a i,l (t)∈{0,1} (e)P c =C Where, η EE Indicates the network's energy efficiency; This represents the transmit power of the femtocell base station n on the k-th subcarrier. Let be the minimum and maximum transmit power of the femtocell base station n on the k-th subcarrier, respectively; This represents the transmit power of macro base station h on the k-th subcarrier. These are the minimum and maximum transmit power of macro base station h on the k-th subcarrier, respectively; γ represents the signal-to-noise ratio of user equipment i served by the l-th base station on the k-th subcarrier. min Indicates the minimum quality of service requirement; a i,l (t) represents the connection relationship between base station l and user equipment i at time t; P c Let C be the power consumption of each circuit in the macro base station and the femto base station, and C is a constant.
7. The method for optimizing the energy efficiency of heterogeneous network resources based on deep reinforcement learning according to claim 1, characterized in that, In S3, defining the agent, state space, action space, and reward function specifically includes: 1) The femtobase station n is used as an agent. Each agent updates its strategy independently. Each agent collects information from its own area and explores the network environment. Each agent selects its own subcarrier and transmit power. 1≤n≤N; 2) State space S n,k (t) is defined as: S n,k (t)={M n (t),P n (t),I k (t),G n,k (t),a i,l (t)}, where M n (t) represents the number of users at the femtocell base station at time t; P n (t) represents the power of the femtocell base station at time t; I k (t)∈{0,1} represents the interference level from the macro base station on the k-th subcarrier at time t. Assume the minimum capacity requirement of the macro base station based on quality of service performance is α. h ,when Interference Level I k When (t) = 0, Interference Level I k (t) = 1; G n,k (t) represents the channel information of the femtobase n and the users at time t on the k-th subcarrier; a i,l (t) represents the connection relationship between the base station and the user at time t; 3) The action space A is defined as: A = {k} n ,p n,k (t)}, where k n p represents the k-th subcarrier of the n-th base station, where k ∈ K; n,k (t) represents the power value on the k-th subcarrier of the n-th femtobase at time t, which is autonomously adjusted through algorithm learning; 4) Reward function based on optimization objective Defined as the user's energy efficiency, that is: Here, β is a constant less than 0.
8. The method for optimizing the energy efficiency of heterogeneous network resources based on deep reinforcement learning according to claim 1, characterized in that, The improved DDPG algorithm specifically includes: Sampling phase: The agent interacts with the communication environment, inputting its current state s(t) into the original Actor network μ(.|θ). μ ), the original Actor network μ(.|θ μ Choose action a(t) according to strategy μ: a(t) = μ(s(t)|θ μ )+N0, where N0 is noise; After the agent performs action a(t), it receives environmental reward r(t) and enters the next state s(t+1), obtaining experience samples {s(t),a(t),r(t),s(t+1)} and storing them in the experience pool D until the storage amount reaches the threshold of the experience pool D. Training phase: N empirical samples are randomly sampled from the experience pool D to form the original Actor network, and the original Critic network Q(.|θ) is formed. Q A training dataset for t is denoted as {s′(t), a′(t), r′(t), s′(t+1)}; Calculate the original Critic network Q(.|θ) Q The loss function Loss is calculated, minimized using the gradient method, and the Critic network parameters θ are updated using backpropagation with the Adam optimizer. Q The loss function Loss is: Among them, y i =r i +γQ′(s i+1 ,μ′(s i+1 |θ μ′ )|θ Q′ ), where γ is a discount factor; Define the agent's objective function as: J(θi)=E[Q μ (s,μ(s))], Maximize the objective function and update the Actor network parameters θ using the Adam optimizer. μ ; The old target network parameters and the new corresponding network parameters are weighted and averaged to softly update the target Actor network and the target Critic network. Where τ is the discount factor.
Citation Information
Patent Citations
Femtocell heterogeneous network power adaptive optimization method based on deep reinforcement learning
CN113795049A
Multi-agent heterogeneous network resource optimization method based on Actor-Critic algorithm
CN114585004A