Lightweight energy efficiency optimization method based on deep reinforcement learning in heterogeneous network
By using a lightweight cloud-edge collaboration framework and deep reinforcement learning, edge base stations can independently control power, solving the problems of high computational complexity and information exchange overhead in heterogeneous networks and achieving efficient energy efficiency optimization.
Patent Information
- Application Number
- CN202511919218.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-03-06
AI Technical Summary
In heterogeneous networks, existing energy efficiency optimization methods require global instantaneous information exchange, resulting in high computational complexity and information exchange overhead, making it difficult to achieve efficient energy efficiency control.
A lightweight cloud-edge collaboration framework is adopted, in which each edge base station independently controls power. By utilizing local information and historical data, an actor-commentator neural network is trained through deep reinforcement learning to reduce information exchange and optimize global energy efficiency.
It achieves real-time and efficient power control for each base station, reduces computational complexity and information exchange overhead, and achieves global energy efficiency performance similar to traditional methods.
Smart Images

Figure CN121619619A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of wireless communication and reinforcement learning technology, specifically relating to a lightweight energy efficiency optimization method based on deep reinforcement learning. Background Technology
[0002] The increasing number of smart devices has led to a significant increase in data throughput demands, placing a heavy burden on traditional cellular network systems. To address this issue, a new type of network called heterogeneous networking has emerged as an effective solution for enhancing network capacity and coverage by deploying different types of base stations within the service area. [1] However, the ultra-dense deployment of heterogeneous base stations will further increase network energy consumption, potentially reducing energy efficiency. Energy efficiency is a key metric in next-generation wireless communication systems, defined as the ratio of total throughput to total power consumption, measured in bits per joule. Maximizing energy consumption in wireless networks is crucial for reducing carbon emissions and creating sustainable communication systems. [2] .
[0003] Joint power control for maximizing energy efficiency has been mathematically proven to be a challenging nondeterministic polynomial problem, making it difficult to obtain the optimal solution. [3] To address this problem, various energy-saving algorithms have been proposed. Existing energy-efficiency power control methods can be broadly categorized into three types: iterative optimization-based methods, deep learning-based methods, and deep reinforcement learning-based methods. Iterative optimization-based methods include the following: For example, a Sequential Fractional Programming (SFP) algorithm for joint power optimization has been proposed by combining fractional programming and sequential optimization techniques. This algorithm solves the problem of maximizing global energy efficiency and achieves near-optimal performance. Another algorithm based on branch and bound can be used to solve common high-efficiency power control problems and obtains a globally optimal solution. [5] In addition, there is an optimization framework for solving fractional polynomial problems, which can be applied to the energy efficiency maximization problem in cellular networks. [6] However, these traditional power control optimization methods require collecting instantaneous channel state information from all base stations, which is challenging in real-world scenarios. Furthermore, they typically operate iteratively, resulting in high computational complexity and potentially making the optimization results outdated.
[0004] In recent years, deep learning and deep reinforcement learning-based methods have gradually demonstrated significant advantages in the field of wireless communication. Some studies have successfully employed these methods to solve the energy efficiency maximization problem. Specifically, this includes using unsupervised methods to train deep neural networks for power optimization, achieving higher energy efficiency fairness performance compared to traditional optimization techniques. Furthermore, centralized, distributed, and transfer learning-based solutions based on deep reinforcement learning have been proposed to investigate the energy efficiency maximization problem in 5G cognitive heterogeneous networks. Among these solutions, multi-agent distributed solutions achieve optimal energy efficiency performance due to effective coordination among agents.
[0005] However, deep learning-based algorithms are susceptible to inconsistencies between offline training and online test datasets in typical heterogeneous networks, especially when the heterogeneous network environment changes rapidly. Deep reinforcement learning-based optimization methods can achieve high energy efficiency, but require extensive information exchange between the core network and different base stations. Summary of the Invention
[0006] To address the aforementioned issues, this invention develops an intelligent power control algorithm that enables each base station to independently control its transmission power using only local information, thereby improving overall energy efficiency.
[0007] This invention considers a typical heterogeneous network, which includes a macro base station and micro base stations, such as Figure 1 As shown. Each base station serves a corresponding user, with all base stations and users using a single antenna. This invention will... The user of the service is represented as a user. ,in And index 0 corresponds to the macro base station and its served users. Note that macro base stations and micro base stations share the spectrum band used for synchronizing downlink transmission.
[0008] This invention models the wireless channel between the base station and the user as consisting of two parts: large-scale fading (path loss and shadowing) and small-scale blocky Rayleigh fading. Path loss is determined by the distance between the base station and the user, while shadowing fading is caused by physical congestion of the base station-user link. This invention represents large-scale fading as the sum of path loss and shadowing, expressed as: ,in and They are from the base station To users These are path loss attenuation and shadow attenuation. Small-scale block Rayleigh fading is caused by multipath transmission in the base station-user link and is a rapidly changing random variable. Specifically, from the base station... To users Block Rayleigh fading is represented as Furthermore, it follows a cyclic symmetric complex Gaussian distribution with zero mean and unit variance, i.e. .
[0009] Then, the channel gain can be expressed as In the user The signal-to-interference-plus-noise ratio measured at the point can be expressed as:
[0010] (1)
[0011] in, For time slots Time base station The transmission power, This refers to the noise power spectral density present at the user's location. Therefore, in the time slot... From the base station To users The downlink rate can be modeled as:
[0012] (2)
[0013] in This refers to the spectrum bandwidth. Considering that total power consumption consists of the base station's transmit power and the power consumed by the hardware circuitry within the base station, therefore, time slots... The global energy efficiency of a time-dependent heterogeneous network can be expressed as:
[0014] (3)
[0015] in It is a base station The reciprocal of the power amplifier efficiency, It is the total power of the circuit.
[0016] As mentioned above, the problem of maximizing global energy efficiency can be formulated as a joint power optimization problem, i.e.,
[0017] (4)
[0018] in, It is a base station Maximum transmit power constraints. Considering that typical heterogeneous networks deploy different base stations to serve users in different areas, these base stations may have different limits on their maximum transmit power.
[0019] Framework for high-efficiency power control methods:
[0020] As can be seen from (4), the global energy efficiency optimization problem in heterogeneous networks is highly related to the network's global instantaneous information, including downlink data rates and energy loss information of all base stations. However, constructing this global instantaneous information requires a large amount of information exchange between different base stations, which will bring significant information exchange overhead. Therefore, this invention is designed to enable lightweight collaboration between the core network (cloud) and different base stations (edge). Specifically, edge base stations do not need to exchange local instantaneous information with each other, while the cloud only needs to collect historical local energy efficiency-related data from each edge base station and calculate historical global rewards, and then use global reward feedback to improve the power control strategy of local base stations.
[0021] Based on the above discussion, this invention proposes a lightweight cloud-edge collaboration framework, such as... Figure 2 As shown. In this framework, each edge base station is connected to the cloud via a two-way wired or wireless link with latency.
[0022] At the edge, this invention establishes an independent actor-commentator structure and two experience buffers for each base station: a local experience replay buffer and a local-global experience replay buffer. Each edge base station will utilize its local actor neural network to determine an appropriate downlink transmit power and store the corresponding local experience and timestamp in its local experience replay buffer. Then, it only needs to upload the local energy efficiency-related data to the cloud.
[0023] In the cloud, this invention establishes a set of queues (receive queue and send queue) for exchanging historical data with each edge base station. Furthermore, this invention also designs a global calculation module that schedules the received data in the receive queue and calculates the historical global reward. The calculated historical global reward, along with its corresponding timestamp, is placed in each send queue, which then sends it to each edge base station in a first-in, first-out manner.
[0024] After receiving data from the cloud, each edge base station retrieves local experiences with the same timestamp from its local experience replay cache and combines them with global reward feedback to form a local-global experience. This local-global experience is then stored in the local-global experience replay cache. In this way, each edge base station can sample mini-batch local-global experiences in each update cycle and then periodically train the actor deep neural network and the critic deep neural network until each deep neural network converges. At this point, each actor deep neural network can learn a better local power control strategy to improve global energy efficiency performance.
[0025] Independent Agent-Judge Algorithm Design
[0026] 1) Edge Network:
[0027] Edge State Design: Actor-Based Deep Neural Networks The state consists of historical local information from the previous time slot, local instantaneous information from the current time slot, and local auxiliary information, namely:
[0028] (5)
[0029] Historical local information includes channel gain. Edge base station Transmission power The received interference The obtained signal-to-interference-plus-noise ratio and downlink achievable speed Then, the local instantaneous information includes channel gain. Interference received at the beginning of the current time slot before configuring the new transmit power. In addition, local auxiliary information is for edge base stations. The reciprocal of the power amplifier efficiency In order to obtain time slots Local instantaneous information at any given moment, base station In the time slot The beginning of users Transmission power is The orthogonal pilot sequence is used to obtain the channel gain locally. and interference .
[0030] Edge motion design: in time slots Moment, Actor Deep Neural Network The action is designed for downlink transmit power, i.e. ;
[0031] Edge Experience: In Time Slots At the end of the time, the edge base station It can acquire local experience and construct it into a set, including state-action pairs from the previous time slot and the state in the current time slot, that is,
[0032] (6)
[0033] Upon receiving historical global rewards from the cloud, each edge base station can construct a local-global experience by combining its local experience with the historical global rewards using the corresponding timestamps. If the global reward is represented as... And the transmission delay between the edge and the cloud is expressed as Then edge base station In the time slot The local-global experience at a given moment can be represented as
[0034] (7)
[0035] Edge Actor and Judge Deep Neural Network Design: Actor Deep Neural Network The network structure design is as follows Figure 3 As shown. Specifically, it is a fully connected deep neural network, with the input layer having a state corresponding to the design state. The number of eight ports, the output is mapped to 0 and The feasible transmit power between. Figure 4 The critique deep neural network was designed in the middle. The network structure includes a state module, an action module, and a value evaluation module. The state module receives local states. As input, the action module receives local actions. As input, the value assessment module takes the outputs of the state and action modules as input. Then, the value assessment module outputs a long-term... The value is used to evaluate the performance of global energy efficiency.
[0036] 2) Cloud network:
[0037] As discussed in the framework, the cloud is designed to calculate historical global rewards by collecting local energy efficiency-related data from each base station. To achieve this process, the present invention establishes a cloud-based... The system consists of a queue group and a global computing module. Each queue group (one receive queue and one transmit queue) is associated with each edge base station and operates on a first-in, first-out (FIFO) basis.
[0038] Global Reward Design: At a given time slot, the global reward is designed to evaluate the global energy efficiency performance achievable by all base stations. To ensure effectiveness and simplicity, this invention designs time slots... The global reward at time step is
[0039] (8)
[0040] Then, according to (4), in the time slot At any time, edge base stations Local energy efficiency data is designed to be data rate and energy consumption information, including transmission power. The reciprocal of the power amplifier efficiency And local circuit power. In this way, each edge base station can upload the acquired local energy efficiency data and timestamps to the cloud at the end of each time slot.
[0041] D Deep Neural Network Training Process
[0042] This invention extends the Deep Deterministic Policy Gradient (DDPG) algorithm. [9] A multi-agent independent actor-judge power control algorithm is introduced to train an edge deep neural network. Based on the edge design, this invention designs deep neural networks for each actor and judge as follows: and ,in and These are the weights of the actor deep neural network and the judge deep neural network, respectively. To ensure the stability of the training process of the judge deep neural network and the actor deep neural network, this invention creates a target judge deep neural network for each judge deep neural network, which is... This represents, and for each actor, a deep neural network Established target actor deep neural network ,Depend on It should be noted that the weights of the target commentator deep neural network and the target actor deep neural network are initialized using the corresponding weights of the commentator deep neural network and the actor deep neural network.
[0043] Initially, all edge base stations determine a random transmit power for downlink transmission at the start of each time slot. Then, by continuously exchanging local energy efficiency data and historical global rewards with the cloud, edge base stations are able to build and store local-global experience in their local-global experience replay caches. Once each local-global experience replay cache has stored at least... Based on experience, edge base stations will randomly sample small batches of local-global experience to train the commentator deep neural network and the actor deep neural network. Below, the present invention provides the training processes for the commentator deep neural network, the actor deep neural network, and the target deep neural network, respectively.
[0044] 1) Training the judge's deep neural network:
[0045] For edge base stations Represent its local-global experience as Then, the judges use a deep neural network. target value It can be represented as
[0046] (9)
[0047] in, It is a discount factor. It is a target actor neural network The output action. Then, the loss function of each commenter's deep neural network can be expressed as the mean squared error between the predicted Q-value and the target's long-term global energy efficiency, i.e.,
[0048] (10)
[0049] Therefore, each commentator's deep neural network can backpropagate its loss and update its weights using a gradient descent-based method. By repeating the above process, each commentator's deep neural network can gradually optimize its weights in the direction of maximizing global energy efficiency performance.
[0050] 2) Training the actor deep neural network:
[0051] The training of the actor-based deep neural network will utilize gradients obtained from the critic-based deep neural network; specifically, the actor-based deep neural network... The weight updates need to maximize the commentator's deep neural network. The long-term Q value, i.e.
[0052] (11)
[0053] in It is an actor-based deep neural network learning rate, Is it the expected Q-value pair? The partial derivative is expressed as:
[0054] (12)
[0055] To balance the utilization of learned policies and the exploration of better policies during the training phase of the actor-based deep neural network, this invention adds Gaussian noise to the output of the actor policy, i.e.,
[0056] (13)
[0057] in, It is zero-mean motion noise, and its variance is... A trade-off between strategy exploitation and exploration was determined. Intuitively, higher variance can accelerate the learning of deep neural networks in the early training phase, while lower variance can help stabilize performance in the later training phase. Here, this invention designs noise variance. With initial value and minimum value It is initialized to As time slots increase at a fixed rate Exponential decay, and when it falls below The time will remain as .
[0058] 3) Training the target deep neural network:
[0059] This invention employs a soft update method to update the target deep neural network to stabilize the training process, that is,
[0060] (14)
[0061] (15)
[0062] in and These are the soft update rates of the target evaluator deep neural network and the target actor deep neural network, respectively.
[0063] The beneficial effects of this invention are as follows: Compared with traditional power control optimization methods, the proposed framework allows each edge base station to determine the appropriate transmit power using only local information. This ensures real-time power control and reduces the computation caused by centralized optimization. Secondly, the proposed framework enables lightweight collaboration between the cloud and different edge base stations, thereby minimizing the information exchange overhead between the cloud and edge base stations. Finally, this invention significantly reduces time complexity while maintaining global energy efficiency performance similar to traditional iterative optimization methods. Attached Figure Description
[0064] Figure 1 This is a schematic diagram of a downlink heterogeneous network structure.
[0065] Figure 2 This is a schematic diagram of the lightweight cloud-edge collaboration framework of the present invention.
[0066] Figure 3 This is a schematic diagram of the deep neural network structure for actors.
[0067] Figure 4 A schematic diagram of a deep neural network structure for evaluation.
[0068] Figure 5 This is a schematic diagram comparing the global energy efficiency performance during the training phase.
[0069] Figure 6 This is a schematic diagram comparing the global energy efficiency performance during the testing phase. Detailed Implementation
[0070] The practicality of this invention is illustrated below with simulation results and accompanying figures. First, the system model settings and hyperparameters used in the simulation are provided. Then, simulation results are presented to evaluate the performance of the proposed algorithm. The simulation results compare the proposed algorithm with the SFP algorithm, stochastic power algorithm, and maximum power algorithm in terms of global energy efficiency and time complexity.
[0071] Table 1 shows the channel parameters and algorithm hyperparameters used in the simulation of this invention. This invention considers a two-layer heterogeneous network scenario with one macro base station and four micro base stations. The macro base station is deployed in the first layer, located at coordinates (0, 0). Micro base stations 1 to 4 are distributed in the second layer, located at coordinates (500, 0), (0, 500), (-500, 0), and (0, -500), respectively. Each base station is responsible for signal coverage of a disk-shaped area. The maximum radius of the macro base station is 1000m, and the maximum radius of the micro base stations is 200m. The minimum coverage radius for all base stations is set to 10m. It is worth noting that the macro base station is subject to a maximum power constraint of 30dBm, and the micro base stations are subject to a maximum power constraint of 23dBm.
[0072] Table 1 Simulation hyperparameters
[0073]
[0074] Specifically, this invention divides the simulation into a training phase and a testing phase. During the training phase, the proposed algorithm trains a deep neural network for 30,000 time slots. During the testing phase, the proposed algorithm tests for 2,000 time slots, directly configuring the base station's transmission power using the well-trained actor deep neural network without any further parameter updates. To ensure the robustness of the simulation results, this invention conducted 10 independent trials using different seed settings. During each trial, the user's location was randomly generated within their coverage area.
[0075] Figure 5 The results of the algorithm during the training phase are provided. It can be seen that after approximately 1000 time slots, the proposed algorithm rapidly outperforms both the random power and maximum power algorithms. After this point, the algorithm converges rapidly and achieves global energy efficiency performance close to that of the SFP algorithm at the end of the training phase. Figure 6 The performance results during the testing phase are presented. It can be seen that the global energy efficiency of the proposed algorithm is much higher than that of the random algorithm and the maximum power algorithm, while achieving approximately 99.08% of the performance of the SFP algorithm.
[0076] It is important to note that while the SFP algorithm utilizes global instantaneous channel state information for centralized power optimization at each time slot, the proposed algorithm enables each base station to determine the appropriate transmit power using only local information. These results highlight the advantages of the proposed algorithm in optimizing global energy efficiency performance in heterogeneous networks.
[0077] Table 2 presents the training parameters of the deep neural networks in the proposed algorithm, and compares the average time complexity of the proposed algorithm with that of the SFP algorithm. Note that the (target) actor and (target) commentator deep neural networks only require local information as input, which reduces the complexity of the input data. Consequently, these deep neural networks can be trained efficiently with relatively fewer training parameters. Furthermore, Table 2 shows that training each deep neural network takes approximately 2 ms, and training each target deep neural network takes less than 3 ms. During testing, the proposed algorithm requires approximately 1.020 ms per actor deep neural network to calculate transmit power for global energy efficiency, while the SFP algorithm requires approximately 53.042 ms to optimize transmit power. This result demonstrates the significant advantage of the proposed algorithm in terms of time complexity. This is expected, as each edge base station can efficiently utilize its local actor deep neural network to directly calculate transmit power, whereas traditional methods rely on iterative calculations, which are very time-consuming.
[0078] Table 2. Comparison of Algorithm Training Parameters and Time Complexity
[0079]
Claims
1. A lightweight energy efficiency optimization method based on deep reinforcement learning in heterogeneous networks, wherein the heterogeneous network includes a macro base station and... Each micro base station serves one corresponding user, and all base stations and users use a single antenna. The user of the service is represented as a user. ,in Furthermore, index 0 corresponds to the macro base station and its served users, and the macro base station and micro base station share the spectrum band used for synchronous downlink transmission; characterized in that, The method comprises: The target model is established by setting the target as jointly optimizing the transmission power of the macro base station and each micro base station and maximizing the global energy efficiency of the whole heterogeneous network: , wherein is a time slot a time base station transmit power, is a micro base station maximum transmit power constraint; Based on the target model, the base station jointly optimizes the transmission power by using deep reinforcement learning, and a lightweight cloud-edge collaboration framework is designed, specifically: Considering that the global energy efficiency optimization problem in the heterogeneous network is highly related to the global instantaneous information of the network, including the downlink data rate and energy loss information of all base stations, the cloud core network and the edge network of different base stations are collaboratively lightweight; Specifically, the edge base station does not need to exchange local instantaneous information with each other, and the cloud only needs to collect historical local energy efficiency related data from each edge base station and calculate the historical global reward, and then use the global reward feedback to improve the power control strategy of the local base station; In the edge, an independent actor-critic structure and two experience buffers, i.e. local experience replay buffer and local-global experience replay buffer, are established for each base station; Each edge base station will use its local actor neural network to determine a suitable downlink transmission power, and store the corresponding local experience and timestamp in its local experience replay buffer, and then upload the local energy efficiency related data to the cloud; In the cloud, a send-receive queue is built to exchange historical data with each edge base station; In addition, a global calculation module is also included, which will schedule the received data in the receive queue, calculate the historical global reward, and the calculated historical global reward will be placed in each send queue together with the corresponding timestamp, and then the send queue will send it to each edge base station in a first-in, first-out manner; After receiving data from the cloud, each edge base station will retrieve local experience with the same timestamp from its local experience replay buffer, and combine it with the global reward feedback to form local-global experience, which will be stored in the local-global experience replay buffer; Thus, each edge base station samples a small batch of local-global experience in each update cycle, and then periodically trains the actor deep neural network and the critic deep neural network until each deep neural network converges.
2. The lightweight energy efficiency optimization method based on deep reinforcement learning in a heterogeneous network according to claim 1, characterized in that, The specific implementation method of the cloud core network and the edge network of different base stations is as follows: 1) Edge network: Edge state design: actor deep neural networks The state of the actor deep neural network is composed of the history local information in the previous time slot, the local instantaneous information in the current time slot and the local auxiliary information, that is: , Historical local information includes channel gain. Edge base station Transmission power The received interference The obtained signal-to-interference-plus-noise ratio and downlink achievable speed Local instantaneous information includes channel gain. Interference received at the beginning of the current time slot before configuring the new transmit power. Local auxiliary information is for edge base stations. The reciprocal of the power amplifier efficiency In order to obtain time slots Local instantaneous information at any given moment, base station In the time slot The beginning of users Transmission power is The orthogonal pilot sequence is used to obtain the channel gain locally. and interference ; Edge action design: in time slot Moment, actor deep neural network The action design of the edge is the downlink transmit power, that is ; Edge experience: in a time slot At the end of the time instant, the edge base station The local experience can be acquired and built as a set including the state-action pairs in the previous time slot and the state in the current time slot, i.e.: , Upon receiving the historical global rewards from the cloud, each edge base station is able to construct the local-global experience by combining the corresponding local experience with the historical global rewards using the corresponding timestamps; if the global reward is denoted as and the transmission delay between the edge and the cloud is denoted as then the edge base station The local-global experience at time slot is represented as: ; 2) Cloud network: The cloud is designed to calculate the historical global reward by collecting local energy efficiency related data from each base station. To achieve this process, a send-receive queue and a global calculation module are established in the cloud, where each set of send-receive queue is associated with each edge base station and operates data in a first-in, first-out manner. Global reward design: At a given time slot, the global reward is designed to evaluate the global energy efficiency performance that all base stations can achieve, time slot The global reward at a time slot is: , Then, according to the target model, at the end of each time slot the local energy efficiency related data of the edge base station is designed as data rate and energy consumption information, including transmit power , the inverse of power amplifier efficiency and local circuit power; so that each edge base station uploads the obtained local energy efficiency related data and time stamp to the cloud at the end of each time slot.