TD3-based new energy enabling cellular-free large-scale MIMO system intelligent AP management method
By using the TD3 algorithm to construct energy-channel state information indicators in a new energy-enabled non-cellular massive MIMO system, and jointly optimizing AP sleep and clustering, the problem of unbalanced AP energy state is solved, and the power grid energy efficiency and communication performance are improved.
Patent Information
- Application Number
- CN202511189886.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-12-02
AI Technical Summary
In new energy-enabled non-cellular massive MIMO systems, existing technologies have failed to effectively utilize the energy state differences of APs, resulting in some APs having excess energy while others have insufficient energy. Furthermore, AP sleep strategies may affect communication performance, and there is a lack of unified AP management methods to improve grid energy efficiency.
The TD3 algorithm is used to construct energy-channel state information indicators, jointly optimize AP sleep and AP clustering, learn network state through MDP process, and dynamically adjust AP management strategy to maximize grid energy efficiency.
It improves the grid energy efficiency in new energy-enabled scenarios, especially performing well in energy-rich networks, dynamically adapting to changes in user distribution and channel status, and balancing energy consumption and communication performance.
Smart Images

Figure CN121056898A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of new energy-enabled non-cellular massive MIMO communication, specifically involving a smart access point (AP) management method for new energy-enabled non-cellular massive MIMO systems based on Twin Delayed Deep Deterministic policy gradient (TD3). Background Technology
[0002] The performance gains of cellular-free massive MIMO systems stem from centralized processing and joint transmission. However, not every access point (AP) needs to participate in providing communication services. The uneven spatiotemporal distribution of user equipment (UEs) means that some APs may only provide minimal performance gains during certain periods. Although cellular-free massive MIMO places baseband processing in the Central Processing Unit (CPU), reducing the power consumption of a single AP compared to cellular base stations, the large number of deployed APs, distributed antennas, and massive connections in the network result in high transmit power and complex communication processing. This invention addresses the AP deployment problem in cellular-free massive MIMO systems by improving network energy efficiency (EE) through strategies such as reducing the number of APs and rationally selecting deployment locations. However, given a fixed AP deployment, energy harvesting (EH) efficiency and UE distribution constantly change at different times, altering the AP's energy storage state and channel state information (CSI). Therefore, AP management can be implemented based on real-time network conditions. Specifically, AP hibernation and AP clustering control can improve the EE of the communication network.
[0003] AP hibernation technology, based on network metrics such as user distribution, propagation loss, and channel gain, enables non-cellular massive MIMO networks to achieve higher EE (Efficiency-Effective) by hibernating APs with high energy consumption or low communication performance gain. However, AP hibernation impacts network communication performance. For a given UE, AP hibernation is equivalent to removing a hibernating AP from the set of APs serving that UE, leading to reduced received signal power, degraded signal quality, and even failure to meet the UE's communication needs. Therefore, to maintain high Spectral Efficiency (SE) performance, AP clustering technology can be used to compensate for the decrease in user SE caused by hibernation and reduce network energy consumption by controlling the AP groups serving each UE in real time. AP clustering optimizes SE and EE metrics based on CSI (Communication-Sensitive Interface), user density, and service fairness, achieving a balance between system performance, complexity, and energy consumption. Therefore, joint optimization of AP hibernation and AP clustering through AP management can better leverage the collaborative capabilities and energy-saving advantages of APs in non-cellular massive MIMO. Furthermore, since AP sleep and AP clustering optimization problems are often non-convex, the DRL algorithm can escape local extrema and find a better global strategy, and is therefore often used to solve such complex problems.
[0004] However, current research focuses on AP management technology in renewable energy-enabled scenarios. Due to the uneven energy consumption and arrival rates of each AP, some APs face high energy consumption, yet EH devices struggle to collect sufficient energy, still requiring significant grid energy consumption. Conversely, some APs with high EH efficiency store generated energy in batteries, even causing energy overflow, hindering the effective utilization of renewable energy. While research on AP hibernation focuses on reducing network energy consumption or improving energy efficiency (EE), it doesn't consider the energy status of each AP. For APs with high EH efficiency, even with significant energy consumption, this can be offset by renewable energy collected by EH devices, eliminating the need for hibernation. For APs with low EH efficiency, hibernation should be determined by considering the overall network impact of their services and the level of energy consumption. Furthermore, simply hibernating APs may degrade communication performance; therefore, it's necessary to dynamically adjust the AP set serving UEs based on hibernation status to improve grid energy efficiency. Thus, AP management technology in renewable energy-enabled scenarios cannot solely focus on CSI (Communication Performance Index) and other information beneficial for improving communication performance; it also needs to consider whether the AP's energy status is suitable for maintaining existing connections or establishing new connections with other UEs. Summary of the Invention
[0005] Purpose of the invention: The purpose of this invention is to provide an AP management method for non-cellular massive MIMO empowered by new energy sources. It proposes an energy-channel state information index adapted to new energy empowerment scenarios and designs an optimization strategy based on the TD3 algorithm that combines AP hibernation and AP clustering to maximize the grid energy efficiency of the network.
[0006] Technical solution: To achieve the above-mentioned objectives, the present invention adopts the following technical solution:
[0007] A smart AP management method for a new energy-enabled non-cellular massive MIMO system based on TD3 includes the following steps:
[0008] With the goal of maximizing grid energy efficiency, this paper addresses the AP management optimization problem in a non-cellular massive MIMO system under a new energy-enabled scenario, which involves dynamically adjusting AP sleep and AP clustering strategies. The connection strength between the AP and UE is measured using energy-channel state information indicators in the new energy-enabled scenario, and AP clustering control is performed by comparing the indicators with a set threshold.
[0009] The joint optimization problem of AP hibernation and AP clustering is modeled as an MDP process, setting up a state space, action space, and reward function; the TD3 algorithm is used to learn the network's energy state and communication state to optimize the AP management strategy.
[0010] Furthermore, the energy-channel state information index model for the new energy-enabled scenario is as follows:
[0011]
[0012] Among them, v m The weights β used to control the channel state and energy state mk β represents the large-scale fading between the m-th AP and the k-th UE. max and β min b represents the maximum and minimum large-scale fading between the AP and UE, respectively. m P represents the battery energy storage of the m-th AP. max This represents the maximum energy consumption of an AP per time slot; determined by a threshold set for each AP. In comparison, determine whether to make a connection, i.e.:
[0013]
[0014] Among them, c mk This indicates whether the m-th AP serves the k-th UE. This represents the threshold used to determine the overall performance of the channel state and AP energy.
[0015] Furthermore, the constructed AP management optimization problem is as follows:
[0016]
[0017]
[0018]
[0019]
[0020]
[0021] in, v m (τ) EE grid (τ), b m (τ) represents whether the m-th AP in the τ-th time slot is activated, the weights of the control channel state and energy state, the threshold for judging the comprehensive performance of the channel state and AP energy, grid energy efficiency, and battery energy storage, respectively. max This indicates the maximum battery energy storage.
[0022] Furthermore, the TD3 algorithm's states include the available renewable energy and energy consumption status of each AP, the connection status between the AP and the UE, and the user's SE information; actions include AP sleep control and the connection status between the AP and the UE; rewards are based on the target setting of maximizing grid energy efficiency.
[0023] The state s is in the form of:
[0024] s={b1,b2,...,b M ,c 11 ,c 12 ,...,c MK SE1, SE2, ..., SE K}
[0025] Among them, SE k Let M represent the spectral efficiency of the k-th UE, M represent the number of APs, and K represent the number of UEs.
[0026] Action a is of the following form:
[0027]
[0028] in, Used to control the sleep and activation states of each AP in the network, v = {v1, v2, ..., v M}and Used to calculate the ECSI index and determine the AP clustering status.
[0029] During policy generation, the tanh function is used to normalize the output range to [-1, 1]. For AP sleep control, a comparison with 0 is used to determine whether the AP is in sleep mode; a value less than 0 indicates AP sleep, otherwise AP is active. Other parameters related to AP clustering in the action are expressed using formulas. Linearly mapped to [0,1].
[0030] The reward r is in the form of:
[0031]
[0032] Where B represents bandwidth, and the spectral efficiency of the k-th UE and the grid power consumption of the m-th AP are related to the AP sleep strategy. It is related to the AP clustering strategy C, where η is a positive number.
[0033] The present invention also discloses a computer system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the intelligent AP management method for the new energy-enabled non-cellular massive MIMO system based on TD3.
[0034] The present invention also discloses a computer program product, including a computer program, which, when executed by a processor, implements the steps of the intelligent AP management method for a new energy-enabled non-cellular massive MIMO system based on TD3.
[0035] Beneficial Effects: Compared with existing technologies, this invention considers the impact of AP sleep and AP clustering on downlink communication links, derives a power grid energy efficiency model under AP management, constructs an AP management optimization problem, and proposes an energy-channel state information index adapted to renewable energy-enabled scenarios to measure the connection strength between APs and UEs. AP clustering control is then performed by comparing this index with a set threshold. Furthermore, this invention considers the energy state of each AP and uses the TD3 algorithm to learn the network's energy and communication states, thereby optimizing AP management strategies and improving power grid energy efficiency for non-cellular massive MIMO. This method demonstrates excellent performance in improving power grid energy efficiency, especially in networks with abundant renewable energy sources. It can also dynamically adapt to changes in user distribution, channel state, and renewable energy collection efficiency within the network, exhibiting good stability and adaptability in balancing energy consumption and communication performance. This provides an efficient AP management solution for renewable energy-enabled non-cellular massive MIMO systems. Attached Figure Description
[0036] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention.
[0037] Figure 2A model diagram of a non-cellular massive MIMO system that empowers new energy sources.
[0038] Figure 3 This is a schematic diagram of AP hibernation and AP clustering.
[0039] Figure 4 This is the convergence graph of the TD3 algorithm.
[0040] Figure 5 This is the CDF diagram of power grid energy efficiency.
[0041] Figure 6 This is a graph showing the average grid energy efficiency as a function of EH efficiency.
[0042] Figure 7 This is a graph showing the average grid energy efficiency as a function of the number of access points (APs). Detailed Implementation
[0043] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0044] like Figure 1 As shown in the figure, an embodiment of the present invention discloses an intelligent AP management method for a new energy-enabled non-cellular massive MIMO system based on TD3, comprising: constructing an AP management optimization problem for a new energy-enabled non-cellular massive MIMO system under a new energy-enabled scenario, with the goal of maximizing grid energy efficiency; wherein the connection strength between AP and UE is measured by the energy-channel state information index of the new energy-enabled scenario, and AP clustering control is performed by comparing it with a set threshold; the joint optimization problem of AP sleep and AP clustering is modeled as an MDP process, setting the state space, action space and reward function; and the TD3 algorithm is used to learn the energy state and communication state of the network to optimize the AP management strategy.
[0045] The following describes the detailed implementation of the embodiments of the present invention in conjunction with a specific system model.
[0046] Figure 2 The paper presents a model of a new energy-enabled non-cellular massive MIMO system. The system consists of M access points (APs) and K user units (UEs), where each AP has N antennas and each UE has a single antenna. These APs are connected to a CPU via communication links to process the information received by each AP. Simultaneously, an energy storage device (EH) and a battery are deployed at each AP to store energy. The APs preferentially use the energy collected and stored in the battery by the EH device; grid energy is used only as a supplement when the battery is depleted.
[0047] In the AP sleep problem, each AP is either active or sleep. To acquire information such as network channel state and energy state, it is assumed that all APs participate in uplink pilot training. Based on the acquired network knowledge, the CPU makes an overall decision. Simultaneously, APs can be clustered using certain strategies to control the connection state between APs and UEs. The AP sleep and AP clustering are illustrated below. Figure 3 .
[0048] In cellular-free massive MIMO that allows AP hibernation and AP clustering, assuming the use of c represents the sleep status of the m-th AP. mk To indicate whether the m-th AP serves the k-th UE, we have:
[0049]
[0050]
[0051] The connection status between the UE and the AP will affect the downlink. Using maximum ratio transmission, let's assume the downlink signal sent to the UE is q. k And satisfy The signal vector sent by the m-th AP can be:
[0052]
[0053] Where, ρ d To determine the maximum downlink transmit power normalized using noise power, η mk Let m be the power control coefficients assigned by the m-th AP to the k-th UE, and let m be the channel between the m-th AP and the k-th UE estimated by MMSE. It contains N independent and identically distributed Gaussian components, where the mean square value of any component is γ. mk .
[0054] Considering the AP's sleep state, the signal received by the k-th UE is:
[0055]
[0056] Where, n k It is the additive white Gaussian noise of the signal received by the k-th UE, which follows the... distributed, For the set of all active APs, g mk Let C be the channel coefficient between the m-th AP and the k-th UE. Assume C is the control factor for the connection between the AP and the UE. mk The M×K dimensional matrix represents the signal-to-noise ratio (SINR) of the k-th UE in downlink communication. k It can be represented as:
[0057]
[0058] in, β represents the set of UEs that use the same pilot signal as the k-th UE. mk This represents the large-scale fading coefficient between the m-th AP and the k-th UE.
[0059] When an AP is in sleep mode, an uplink training phase is still required to obtain the network-wide CSI and transmit the AP's energy status. Simultaneously, the operation of the fronthaul link, internal circuitry, and components needs to be maintained. In a system where APs are in sleep mode, the power consumption of the m-th AP is remodeled as follows:
[0060]
[0061] in, This represents the fixed power consumption of the m-th AP. For an active AP, the non-fixed power consumption is... Affected by the connection status between the AP and the UE, namely:
[0062]
[0063] Where, α m Denotes the power amplifier coefficient, and α m ∈(0,1],P noise Indicates noise power. For the flow-related forward power factor, SE k =log2(1+SINR) k ) represents the spectral efficiency of the k-th UE.
[0064] In the τth time slot, the available new energy source can be used to store the available battery energy b in the current time slot. m (τ). Since priority is given to using new energy sources to maximize their utilization rate, the grid energy consumption of the m-th AP can be calculated as follows:
[0065]
[0066] Based on this, the battery status of the m-th AP is updated as follows:
[0067]
[0068] in, This refers to the new energy source collected by the m-th AP in the τ-th time slot.
[0069] Based on the above system model, the AP management optimization problem in the new energy empowerment scenario constructed by the embodiments of the present invention is described as follows:
[0070] Due to the introduction of AP hibernation and AP clustering, SE kand Both are related to AP sleep and AP clustering strategies, thus enabling the calculation of EE. grid At this time, the power grid energy efficiency (EE) grid This can be expressed as:
[0071]
[0072] By dynamically adjusting the AP's sleep and AP clustering strategies, energy efficiency per time slot is maximized. The optimization problem can be formulated as follows:
[0073]
[0074]
[0075]
[0076] Constraint (11c) is the energy storage limit for each AP, specifying that the battery storage cannot exceed the maximum energy storage limit b. max .
[0077] Existing AP clustering schemes determine the connection status between UE and AP based on information such as CSI and user density. This invention, considering the impact of the energy status of each AP on grid energy efficiency in renewable energy-enabled scenarios, proposes a new index, ECSI, to measure CSI and AP energy information. Its expression is modeled as follows:
[0078]
[0079] Among them, v m ∈[0,1], used to control the weights of the channel state and energy state, β max and β min These represent the maximum and minimum large-scale fading between the AP and UE, respectively, in dB and P. max This represents the maximum energy consumption of the AP per time slot. The significance of setting the ECSI (Energy Management System Index) lies in assessing the combined capabilities of channel conditions and AP energy-saving services, and in establishing thresholds for each AP. In comparison, determine whether to make a connection, i.e.:
[0080]
[0081] in This represents the threshold used to determine the overall performance of the channel state and AP energy.
[0082] In equation (12) s max β min and P maxThe introduction of this parameter is to map the ECSI index to the range of [0,1] when the available renewable energy is insufficient to support the maximum energy consumption, so as to facilitate the judgment. However, it should be noted that when the available renewable energy is greater than the maximum energy consumption, the renewable energy storage of the AP at that location cannot be exhausted, and the second term in equation (12) increases, leading to an increase in ECSI. mk It may be greater than 1. In this case, the m-th AP and the k-th UE will definitely be connected, which meets the expectation of prioritizing the use of energy-rich APs to provide SE gain to the network, thereby increasing the efficiency of new energy utilization.
[0083] At this point, the parameter c related to AP clustering in the optimization variables... mk (τ) can be derived from v m (τ) and The replacement requires modifying the constraints, and the optimization problem can be rewritten as follows:
[0084]
[0085]
[0086]
[0087]
[0088]
[0089] Among them, due to the introduction of ECSI, the optimization variables used for AP clustering have been reduced from MK to 2M.
[0090] Due to the large number of connections in large-scale non-cellular networks, the optimization variables for jointly handling AP sleep and AP clustering problems are enormous, and the empowerment of new energy sources exacerbates this situation. Twin Delayed Deep Deterministic Policy Gradient (TD3) builds upon the DDPG algorithm by delaying network parameter updates, making the training process more stable. This invention proposes a method using the TD3 algorithm to solve the joint AP sleep and AP clustering optimization problem. By optimizing neural network parameters to learn the characteristics of the communication network, an efficient AP management method is obtained, improving grid energy efficiency.
[0091] The TD3 algorithm, also an offline reinforcement learning algorithm based on an "actor-critic" structure, is a type of DRL algorithm. It introduces three key improvements over DDPG: using a dual Q-network to reduce overestimation of Q-values; delaying action network updates to reduce policy oscillations; and adding noise to the target Q-value through smoothing operations on the target action, thereby improving robustness. Similar to DDPG, TD3 aims to find the optimal policy such that, given policy π... φGiven the condition of maximizing the cumulative discount reward, its objective function is:
[0092]
[0093] Among them, s t Indicates the state of time slot t, a t Let t be the action in time slot t, γ∈[0,1] be the discount factor, p(s0) be the initial state distribution, and T be the number of time steps in the finite time domain.
[0094] TD3 includes an action network responsible for generating optimal actions, two evaluation networks for evaluating the Q-values of state-action correctness to reduce overestimation bias, and two target evaluation networks and one target action network for stable training.
[0095] The goal of evaluating the network is to minimize the temporal difference (TD) error.
[0096]
[0097] Among them, Q θ (s t ,a t To evaluate the network's output, we represent the state-action pair (s). t ,a t Q-value estimation is used to predict the long-term cumulative reward of the current policy under a given state and action. This is an experience replay pool, storing experience tuples (s) t ,a t ,r t ,s t+1 ), target Q value y t Calculated by the target Q-network:
[0098]
[0099] Where θ1′ and θ2′ are the parameters of the target evaluation network. Compared to the DDPG algorithm, TD3 uses a dual evaluation network, calculating the TD error using the minimum of the two target Q values, thus avoiding overestimation of the Q value. Through the next action a... t+1 Adding Gaussian noise can smooth the target's movement, that is:
[0100] a t+1 =π φ′ (s t+1 )+ε, (18)
[0101] ε~N(0,σ 2(19) where φ′ is the target action network parameter and σ is the noise standard. By smoothing the target action, it is possible to prevent the policy from fitting to sharp peaks in the Q value. In order to prevent the policy from deviating too far from the actual policy due to the influence of noise, the clip operation is used to limit the range of noise ε, which can enhance the stability of the TD3 algorithm.
[0102] The goal of the action network is to maximize the Q-value estimated by the evaluation network.
[0103]
[0104] Gradient calculation of action networks only uses This avoids affecting gradient updates:
[0105]
[0106] It is important to note that TD3 uses a delayed update method to update the action network, that is, the action network is updated only after the evaluation network has been updated multiple times. This can reduce the learning of suboptimal or incorrect decisions due to the instability of the evaluation network.
[0107] Target network soft update:
[0108] φ′←τφ+(1-τ)φ′, (22)
[0109] θ′ i ←τθ i -(1-τ)θ i ', (twenty three)
[0110] Here, τ << 1 is a soft update parameter to ensure the stability of training.
[0111] In this embodiment, AP management is implemented based on the TD3 algorithm. AP management improves grid energy efficiency by dynamically adjusting AP sleep and AP clustering states. In cellular-free massive MIMO, an agent is set up at the CPU to input the current observation of the energy and information state in the network into the action network, which can obtain AP sleep and AP clustering strategies and apply them to change the network state in cellular-free massive MIMO. Through continuous interaction between the agent and the communication network environment, network parameters are updated, and network performance can be continuously improved. The joint optimization problem of AP sleep and AP clustering can be modeled as an MDP process, with the following settings for the state space, action space, and reward function:
[0112] State space: State s∈S describes the available renewable energy and energy consumption status of each AP at time t, the connection status between the AP and the UE, and the user's SE information.
[0113] s={b1,b2,...,b M ,c 11 ,c12 ,...,c MK SE1, SE2, ..., SE K}(twenty four)
[0114] Action space: Action This refers to the decision made by the agent at time t based on the network state to maximize the reward. In the AP sleep and cooperative management problem, AP sleep control and the connection state between the AP and the UE are treated as actions, taking the following form:
[0115]
[0116] in, Used to control the sleep and activation states of each AP in the network, v = {v1, v2, ..., v M}and Used to calculate the ECSI index and determine the AP clustering status.
[0117] Actions in the action space are continuous. During policy generation, the output range is normalized to [-1, 1] using the tanh function. AP sleep control is determined by comparing with 0; a value less than 0 indicates AP sleep, otherwise AP is active. Other parameters related to AP clustering in the actions are handled using formulas. Linearly mapped to [0,1].
[0118] Reward function: In state s, the agent performs action a according to the action network, and thus obtains a reward value r based on the feedback from the environment. t To achieve the optimization goal of maximizing grid energy efficiency, the reward design is as follows:
[0119]
[0120] The specific process of the AP management method based on TD3 is shown in Algorithm 1.
[0121]
[0122]
[0123] To verify the performance of the proposed algorithm in the new energy empowerment scenario, it was compared with the following algorithms: 1) AP sleep strategy based on DDPG without considering new energy conditions, and AP clustering was added; 2) random sleep and dynamic AP clustering (DCC) based on CSI; 3) random sleep and active AP provides services to all UEs, denoted as random in the figure.
[0124] Figure 4 The convergence curve of the TD3 algorithm is shown, with the number of APs M=40. It can be seen that the proposed TD3-based AP management method can converge to a high level in about 350 rounds, improving grid energy efficiency performance by approximately 39% compared to the DDPG algorithm.
[0125] Figure 5 The cumulative distribution function (CDF) of grid energy efficiency is shown, with M=40 access points (APs) and a total of 200 time slots operated. The results demonstrate that the proposed TD3-based AP management method achieves the best performance. This is because, in the scenario of renewable energy empowerment, the proposed algorithm selects and clusters APs based on their energy state, SE state, and connectivity, enabling more efficient use of green energy in AP management, thereby reducing grid energy consumption and achieving higher grid energy efficiency.
[0126] Figure 6 The graph illustrates the variation of average grid energy efficiency with EH efficiency, where EH efficiency is set from 0% to 150% in the aforementioned parameter settings, increasing arithmetically in increments of 25%, with the number of access points (APs) M=40. It can be seen that when EH devices can provide a higher level of renewable energy, the proposed AP management method based on TD3 exhibits better performance. This is because the proposed algorithm pays more attention to the network's energy distribution and implements strategies beneficial to grid energy efficiency based on battery energy status. Even when the overall renewable energy supply to the network is insufficient, the proposed algorithm still achieves results similar to DDPG and outperforms the other two benchmark methods.
[0127] Figure 7 The graph shows the average grid energy efficiency as a function of the number of access points (APs), where the total number of antennas is fixed at 480, and the number of APs is M = {20, 30, 40, 60, 80}, with each AP having the same number of antennas. As the number of APs increases, the grid energy efficiency gradually decreases. This is because each additional AP consumes energy, regardless of whether it is in sleep mode or not. Furthermore, the increased number of APs leads to increased grid energy consumption due to their higher energy consumption during activation. Moreover, the increased number of APs does not provide sufficient SE gain to match the energy consumption level under equal power distribution conditions, resulting in reduced grid energy efficiency.
[0128] This invention also discloses a computer system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the intelligent AP management method for a new energy-enabled non-cellular massive MIMO system based on TD3.
[0129] This invention also discloses a computer program product, including a computer program that, when executed by a processor, implements the steps of the intelligent AP management method for a new energy-enabled non-cellular massive MIMO system based on TD3.
[0130] The program code used to implement the method of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the steps of the method of the present invention to be performed. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a standalone software package, or entirely on a remote machine or server. All aspects not detailed in this invention are well-known to those skilled in the art.
Claims
1. A smart AP management method for a new energy-enabled non-cellular massive MIMO system based on TD3, characterized in that, Includes the following steps: With the goal of maximizing grid energy efficiency, this paper addresses the AP management optimization problem in a non-cellular massive MIMO system under a new energy-enabled scenario, which involves dynamically adjusting AP sleep and AP clustering strategies. The connection strength between the AP and UE is measured using energy-channel state information indicators in the new energy-enabled scenario, and AP clustering control is performed by comparing the indicators with a set threshold. The joint optimization problem of AP hibernation and AP clustering is modeled as an MDP process, setting up a state space, action space, and reward function; the TD3 algorithm is used to learn the network's energy state and communication state to optimize the AP management strategy.
2. The intelligent AP management method for new energy-enabled non-cellular massive MIMO systems based on TD3 according to claim 1, characterized in that, The energy-channel state information index model for the new energy empowerment scenario is as follows: Among them, v m The weights β used to control the channel state and energy state mk β represents the large-scale fading between the m-th AP and the k-th UE. max and β min b represents the maximum and minimum large-scale fading between the AP and UE, respectively. m P represents the battery energy storage of the m-th AP. max This represents the maximum energy consumption of an AP per time slot; determined by a threshold set for each AP. In comparison, determine whether to make a connection, i.e.: Among them, c mk This indicates whether the m-th AP serves the k-th UE. This represents the threshold used to determine the overall performance of the channel state and AP energy.
3. The intelligent AP management method for a new energy-enabled non-cellular massive MIMO system based on TD3 according to claim 2, characterized in that, The AP management optimization problem is as follows: in, v m (τ) EE grid (τ), b m (τ) represents whether the m-th AP in the τ-th time slot is activated, the weights of the control channel state and energy state, the threshold for judging the comprehensive performance of the channel state and AP energy, grid energy efficiency, and battery energy storage, respectively. max This indicates the maximum battery energy storage.
4. The intelligent AP management method for a new energy-enabled non-cellular massive MIMO system based on TD3 according to claim 1, characterized in that, The TD3 algorithm's status includes the available renewable energy and energy consumption status of each AP, the connection status between the AP and the UE, and the user's SE information; actions include AP sleep control and the connection status between the AP and the UE; rewards are based on the goal of maximizing grid energy efficiency.
5. The intelligent AP management method for a new energy-enabled non-cellular massive MIMO system based on TD3 according to claim 4, characterized in that, The state s is of the following form: s={b1,b2,...,b M ,c 11 ,c 12 ,...,c MK ,SE1,SE2,...,SE K } Among them, b m c represents the battery energy storage of the m-th AP. mk Indicates whether the m-th AP serves the k-th UE, SE k Let M represent the spectral efficiency of the k-th UE, M represent the number of APs, and K represent the number of UEs.
6. The intelligent AP management method for a new energy-enabled non-cellular massive MIMO system based on TD3 according to claim 4, characterized in that, Action a is of the following form: in, Used to control the sleep and activation states of each AP in the network, v = {v1, v2, ..., v M }and Used to calculate the ECSI index and determine the AP clustering status, where M represents the number of APs.
7. The intelligent AP management method for a new energy-enabled non-cellular massive MIMO system based on TD3 according to claim 6, characterized in that, During policy generation, the tanh function is used to normalize the output range to [-1, 1]. For AP sleep control, a comparison with 0 is used to determine whether the AP is in sleep mode; a value less than 0 indicates AP sleep, otherwise AP is active. Other parameters related to AP clustering in the action are expressed using formulas. x∈[-1,1] is linearly mapped to [0,1].
8. The intelligent AP management method for a new energy-enabled non-cellular massive MIMO system based on TD3 according to claim 4, characterized in that, The reward r is in the form of: Where B represents bandwidth, Representing the spectral efficiency of the k-th UE and the grid power consumption of the m-th AP, respectively, both are related to the AP sleep strategy. It is related to the AP clustering strategy C, where M represents the number of APs, K represents the number of UEs, and η is a positive number.
9. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by the processor, it implements the steps of the intelligent AP management method for a new energy-enabled non-cellular massive MIMO system based on TD3, according to any one of claims 1-8.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the intelligent AP management method for a new energy-enabled non-cellular massive MIMO system based on TD3 as described in any one of claims 1-8.