New energy enabling cellular-free large-scale MIMO intelligent AP deployment method based on K-means and DDPG

By combining K-means and DDPG methods, high EH efficiency AP candidate locations were screened, and AP deployment in cellular-free massive MIMO systems was optimized. This solved the problem of the disconnect between energy supply and communication performance, achieved a balance between energy consumption and quality of service in dense UE areas, and improved grid energy efficiency and spectrum efficiency.

CN121126366AActive Publication Date: 2025-12-12SOUTHEAST UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511223592.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-12-12
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

In existing non-cellular massive MIMO systems, energy supply and communication performance optimization are disconnected, failing to achieve a joint balance between energy consumption and quality of service in dense UE areas and low EH efficiency areas. Traditional DRL methods are difficult to converge effectively due to the high-dimensional action space, and the static environment assumption ignores user mobility and time-varying energy characteristics, resulting in limited energy efficiency in actual deployment.

Method used

We adopt a new energy-enabled, non-cellular, large-scale MIMO smart AP deployment method based on K-means and DDPG. We use a geographic-energy dual-dimensional clustering mechanism to screen out AP candidate locations with high energy efficiency, and combine the DDPG algorithm to optimize AP deployment. This constructs an optimization problem that maximizes long-term grid energy efficiency, reducing the training complexity of DDPG and improving learning efficiency.

Benefits of technology

It optimizes the EH efficiency and communication conditions of AP deployment locations in networks with uneven traffic distribution, reduces the training complexity of DDPG, improves grid energy efficiency and spectrum efficiency, adapts to non-uniform user distribution and energy collection differences, and breaks through the limitations of the traditional uniform distribution assumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121126366A_ABST
    Figure CN121126366A_ABST
Patent Text Reader

Abstract

The invention discloses a new energy enabling cellular-free large-scale MIMO intelligent AP deployment method based on K-means and DDPG, and the method comprises the steps: building an AP deployment optimization problem in a new energy enabling scene, selecting a specified number of APs from all AP candidate positions which are non-uniformly distributed, cooperating with UE in a service network, and carrying out the optimization of the AP deployment in a new energy enabling scene under a battery energy storage constraint condition, the optimization target of maximizing the average power grid energy efficiency is realized; clustering all AP candidate positions based on a K-means algorithm by adopting a geography-energy two-dimensional clustering mechanism, and selecting a plurality of AP candidate positions with the highest statistical EH efficiency in each group to form a new AP candidate position set so as to reduce the action dimension of the DDPG algorithm; and solving an AP deployment optimization problem based on a DDPG algorithm, wherein the action design is a selection state of a new AP candidate position screened by a K-means algorithm. The result shows that the provided AP deployment method has a great energy-saving advantage in a new energy enabling cellular-free large-scale MIMO scene with non-uniform flow distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of new energy-enabled non-cellular massive MIMO technology, specifically involving a new energy-enabled non-cellular massive MIMO intelligent access point (AP) deployment method based on K-means and deep deterministic policy gradient (DDPG). Background Technology

[0002] In existing research on non-cellular massive MIMO systems, AP deployment strategies primarily focus on communication performance optimization, such as improving spectral efficiency or expanding coverage. Traditional methods (such as greedy algorithms and minimization strategies) select AP locations through static optimization, but they generally rely on pure grid power supply models and fail to integrate the spatial imbalance of renewable energy. Although some studies attempt to adapt to non-uniform user distribution (such as dynamic AP selection or UAV trajectory optimization), they still ignore differences in energy harvesting efficiency, leading to increased grid energy consumption in high-service-density areas due to insufficient energy harvesting efficiency. Furthermore, although some work has introduced deep reinforcement learning, the problems of low training efficiency and convergence difficulties caused by the high-dimensional action space when directly dealing with large-scale potential deployment locations have not yet been resolved.

[0003] Existing technologies have significant shortcomings: First, the optimization of energy supply and communication performance is disconnected, failing to achieve a joint balance between energy consumption and quality of service in densely populated UE areas and areas with low EH efficiency; second, traditional DRL methods are difficult to converge effectively due to the explosion of action space dimensions (when the number of deployment locations is large), which restricts the feasibility of large-scale network deployment and the scalability of the algorithm; third, the static environment assumption ignores user mobility and time-varying energy characteristics (such as day / night / weather fluctuations), lacking the ability to dynamically optimize long-term grid energy efficiency, resulting in limited energy efficiency in actual deployment. Summary of the Invention

[0004] Purpose of the invention: To address the shortcomings of existing technologies, the purpose of this invention is to provide a method based on K-means and...

[0005] DDPG's new energy-enabled non-cellular large-scale MIMO smart AP deployment method solves the energy-saving problem in situations where different deployment locations in networks with uneven traffic distribution have different EH efficiency and communication conditions.

[0006] Technical solution: To achieve the above-mentioned objectives, the present invention adopts the following technical solution:

[0007] A new energy-enabled, large-scale MIMO smart AP deployment method based on K-means and DDPG includes the following steps:

[0008] The problem of AP deployment optimization in the context of new energy empowerment is constructed. A specified number of APs are selected from all candidate AP locations that are not uniformly distributed, and the UEs in the network are coordinated to achieve the optimization goal of maximizing the average grid energy efficiency under the constraint of battery energy storage.

[0009] A geographic-energy dual-dimensional clustering mechanism is adopted. All AP candidate locations are clustered based on the K-means algorithm. In each group, multiple AP candidate locations with the highest statistical EH efficiency are selected to form a new set of AP candidate locations, thereby reducing the action dimension of the DDPG algorithm.

[0010] The DDPG algorithm is used to solve the AP deployment optimization problem, where the action design is the selection state of the new AP candidate positions after being filtered by the K-means algorithm.

[0011] Furthermore, the optimization problem can be expressed as:

[0012]

[0013] in, Let T be the set of candidate AP locations, and T be the number of time slots. M represents the number of APs selected, EE grid (τ) represents the grid energy efficiency in the τ-th time slot, b m (τ) represents the battery energy storage that the m-th AP can use in the τ-th time slot, b max This represents the upper limit of battery energy storage.

[0014] Furthermore, the steps for initially screening candidate AP locations based on the K-means algorithm include:

[0015] M APs are randomly selected from all AP candidate locations as center locations, and each center corresponds to a group, where M is the number of APs specified in the optimization problem.

[0016] Calculate the distance of the remaining APs from the center and classify them into the nearest group;

[0017] The group center is continuously updated as the new AP deployment location, and the groups are regrouped according to the new center until the termination condition is met;

[0018] After the termination condition is met, select the L≥2 AP candidate positions with the highest statistical EH efficiency from each group to form a new set of candidate APs.

[0019] Furthermore, in the DDPG algorithm, the state space is the current AP deployment scheme s = {a1, a2, ..., a...} I}, where I represents the number of candidate positions for AP, and the action space is a = {a1, a2, ..., a...} I}; The reward is in Indicates AP combination as The SE of the k-th UE, where K is the number of UEs and B is the communication bandwidth. Let ξ represent the grid energy consumption of the m-th AP, and let ξ be a positive number to prevent the reward from growing rapidly due to the small grid energy consumption.

[0020] Furthermore, considering the spatial variability of renewable energy collection efficiency and the non-uniformity of user distribution, the grid energy consumption of the m-th AP in the τ-th time slot is expressed as: Where P m (τ) represents the total power consumption of the m-th AP, including downlink transmit power and fixed circuit losses, b m (τ) represents the time-varying characteristics of the battery state. in This refers to the new energy source collected by the m-th AP in the τ-th time slot.

[0021] Based on the same inventive concept, this invention also provides a new energy-enabled non-cellular large-scale MIMO intelligent AP deployment system based on K-means and DDPG, comprising:

[0022] The problem construction module is used to construct AP deployment optimization problems in new energy empowerment scenarios. It selects a specified number of APs from all AP candidate locations that are not uniformly distributed, and coordinates with UEs in the service network to achieve the optimization goal of maximizing average grid energy efficiency under battery energy storage constraints.

[0023] The clustering and dimensionality reduction module is used to adopt a geographic-energy dual-dimensional clustering mechanism, based on the K-means algorithm to cluster all AP candidate locations, and select the AP candidate locations with the highest statistical EH efficiency in each group to form a new set of AP candidate locations, so as to reduce the action dimension of the DDPG algorithm.

[0024] And the DDPG solution module, which is used to solve the AP deployment optimization problem based on the DDPG algorithm, wherein the action design is the selection state of the new AP candidate positions after being filtered by the K-means algorithm.

[0025] The present invention also provides a computer system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the new energy-enabled non-cellular large-scale MIMO smart AP deployment method based on K-means and DDPG.

[0026] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the described method for deploying new energy-enabled non-cellular large-scale MIMO smart APs based on K-means and DDPG.

[0027] Beneficial Effects: Compared with existing technologies, this invention addresses the issue of differences in EH efficiency and communication conditions at different deployment locations in networks with uneven traffic distribution. It studies a deployment strategy for non-cellular massive MIMO access points empowered by new energy sources, constructs an optimization problem to maximize long-term grid energy efficiency, and establishes a joint optimization framework for non-uniform services and differentiated energy sources. This modeling considers the spatial non-uniform distribution of user equipment (UEs) (e.g., dense and sparse areas), dynamic battery status, and spatial differences in renewable energy collection efficiency, integrating both into a unified optimization framework. This is closer to actual deployment scenarios and overcomes the limitations of the traditional uniform distribution assumption. For the high-dimensional action space problem, this invention proposes a K-means-based pre-screening method, employing a geographic-energy dual-dimensional clustering mechanism to select a set of candidate APs with both high energy efficiency (EH) and coverage capabilities from potential deployment locations. This preprocessing not only reduces the training complexity of DDPG but also preserves spatial diversity and energy efficiency advantages through clustering. The K-means clustering criterion in this invention not only depends on geographical location but also incorporates EH efficiency, thereby optimizing the candidate AP set in terms of both spatial coverage and energy supply. This is an extension of the traditional K-means clustering based solely on distance, and it can achieve good results for complex AP layouts in different non-cellular large-scale MIMO systems. Attached Figure Description

[0028] Figure 1 This is a flowchart of a method according to an embodiment of the present invention.

[0029] Figure 2 Deployment model diagram for AP.

[0030] Figure 3 This is a schematic diagram of the DDPG algorithm.

[0031] Figure 4 This is the convergence graph of the DDPG algorithm.

[0032] Figure 5 This is a graph showing how average grid energy efficiency changes with the number of deployed APs.

[0033] Figure 6 This is a graph showing the average grid energy efficiency as a function of the number of UEs.

[0034] Figure 7 This is a graph showing the average grid energy efficiency as a function of EH efficiency. Detailed Implementation

[0035] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0036] like Figure 1 As shown in the embodiments of the present invention, the new energy-enabled non-cellular large-scale MIMO smart AP deployment method based on K-means and DDPG first constructs an AP deployment optimization problem in a new energy-enabled scenario, selecting a specified number of APs from all non-uniformly distributed AP candidate locations to coordinate with UEs in the network, and achieving the optimization objective of maximizing average grid energy efficiency under battery energy storage constraints. Then, a geographic-energy dual-dimensional clustering mechanism is adopted, and all AP candidate locations are clustered based on the K-means algorithm. In each group, multiple AP candidate locations with the highest statistical EH efficiency are selected to form a new AP candidate location set to reduce the action dimension of the DDPG algorithm. Finally, the AP deployment optimization problem is solved based on the DDPG algorithm, where the action design is the selection state of the new AP candidate locations after K-means algorithm screening.

[0037] In this embodiment, a new energy-enabled massive MIMO network with uneven UE distribution is considered, such as... Figure 2 As shown, there are multiple locations in the network where APs can be deployed, called AP candidate locations. These AP candidate locations are randomly distributed throughout the network, and their efficiency (EH) may vary depending on surrounding buildings, weather, and other natural factors. It is assumed that the statistical information of the EH efficiency of each candidate location can be surveyed in advance.

[0038] The specific steps of the new energy-enabled large-scale MIMO smart AP deployment method based on K-means and DDPG are illustrated below:

[0039] Step S1: Construct the AP deployment optimization problem in the new energy empowerment scenario. Select a specified number of APs from all AP candidate locations that are not uniformly distributed, and coordinate with the UEs in the network to achieve the optimization goal of maximizing the average grid energy efficiency under the constraint of battery energy storage.

[0040] System Model: In renewable energy-enabled non-cellular massive MIMO scenarios, the energy consumed comes from energy-saving devices (EH devices) and the power grid. The renewable energy provided by EH devices is considered clean. Therefore, achieving energy saving in communication under renewable energy-enabled conditions does not necessarily require improving the network's energy efficiency (EE). Instead, it involves minimizing power grid energy consumption and achieving better energy efficiency (SE) performance, while taking EH into account, thereby improving the power grid's energy efficiency (EE). grid This refers to the SE performance per unit of grid energy consumption. Assuming... Let a be the set of candidate locations for AP, and let its element a be the set of candidate locations for AP. m Indicates whether the m-th position is selected:

[0041]

[0042] Channel modeling is as follows:

[0043]

[0044] Where β mk The large-scale fading coefficient depends on the probability-based line-of-sight (LoS) or non-line-of-sight (NLoS) transmission model. This represents small-scale fading. The pilot signal received by the m-th AP from the UE is represented as...

[0045]

[0046] Where ρ p τ represents the power of the pilot signal. d Indicates downlink data transmission delay. It is a dimension τ p A complex vector of size ×1, which is the pilot sequence assigned to the k-th UE and satisfies the normalization condition. Its design goal is to achieve maximum orthogonality among different users to reduce mutual interference. (Superscript: ·) H W represents the conjugate transpose operation. m It is a dimension N×τ p The complex noise matrix, where each element is an independent and identically distributed (i.i.d.) complex Gaussian random variable.

[0047] The channel estimated by MMSE is represented as follows:

[0048]

[0049] in

[0050] Use symbols Indicates from AP candidate position The M APs selected are combined as follows: The SE of the k-th UE is expressed as:

[0051]

[0052] Where, τ c Indicates the channel coherence time. Indicates the desired signal gain. This represents the gain due to beamforming uncertainty. ρ represents the inter-user interference gain. d η represents the downlink transmission power normalized to noise. mk This represents the power control coefficient from the m-th AP to the k-th UE.

[0053] The power grid energy efficiency is expressed as follows:

[0054]

[0055] in, Indicates AP combination as The SE of the k-th UE, where K is the number of UEs and B is the communication bandwidth. This represents the power grid energy consumption of the m-th AP.

[0056] In the τth time slot, since the energy stored in the battery is used preferentially, the grid energy consumption can be expressed as:

[0057]

[0058] Among them, P m (τ) represents the total energy consumption of the m-th AP, b m (τ) represents the battery energy storage that the m-th AP can use in time slot τ, indicating the time-varying characteristics of the battery state. Therefore, the battery capacity available for the m-th AP in the next time slot is:

[0059]

[0060] in, In this embodiment, the new energy collected by the m-th AP in the τ-th time slot is... It follows a uniform distribution, representing the amount of renewable energy collected.

[0061] In this embodiment, the total energy consumption of the AP Includes downlink transmit power and fixed circuit losses α m It refers to the power amplifier efficiency. The mean square value for channel estimation is given, where N0 represents the noise power and N represents the number of antennas configured in the AP.

[0062] Based on the above system model, the AP deployment problem in the new energy empowerment scenario is described as selecting M APs from candidate AP locations to collaboratively serve UEs in the network. Since the location of the APs affects the network's channel state, and the EH efficiency of APs at different locations also differs, it is necessary to select a suitable combination of APs to achieve the optimization objective of maximizing average grid energy efficiency over a long period. Therefore, the optimization problem can be constructed as follows:

[0063]

[0064]

[0065] In the above problem, equations (5a) and (5b) specify the number of APs to be deployed, and equation (5c) is the battery energy storage constraint, where b maxThis represents the upper limit of battery energy storage. It is important to note that the problem of maximizing long-term grid energy efficiency is modeled as maximizing grid energy efficiency over T coherent time intervals, and it is assumed that T coherent time intervals are sufficient to represent the long-term performance of the network.

[0066] Step S2: Solve the AP deployment optimization problem in the constructed new energy empowerment scenario based on K-means and DDPG.

[0067] The aforementioned problem is a combinatorial optimization problem, which is non-convex and difficult to solve using traditional algorithms, or prone to getting trapped in local optima. Due to the large number of connections in non-cellular massive MIMO, the AP deployment problem model is complex. Therefore, the Deep Deterministic Policy Gradient (DDPG) algorithm is used, which automatically learns the optimal policy through its reward mechanism to optimize AP deployment.

[0068] In the AP deployment problem, using the DDPG algorithm requires transforming the problem into a Markov decision process (MDP). The actions, designed as selection states (selected or not selected) for all AP candidate positions, have a large dimension, easily leading to low sample efficiency and exploration difficulties for the DDPG method, thus hindering strategy optimization. K-means clustering can divide a continuous or high-dimensional state or action space into several clusters, with high similarity among points within each cluster. Using clustering preprocessing, advantageous subsets can be selected from multiple subspaces, ensuring spatial coverage performance of AP candidate positions and reducing dimensionality to simplify the learning task. Therefore, this embodiment first uses the K-means algorithm to reduce the action dimensionality before using the DDPG algorithm to solve the AP deployment problem.

[0069] Step S2.1: The geographic-energy dual-dimensional clustering mechanism is adopted. All AP candidate locations are clustered based on the K-means algorithm. In each group, the AP candidate locations with the highest statistical EH efficiency are selected to form a new set of AP candidate locations, so as to reduce the action dimension of the DDPG algorithm.

[0070] Considering the need to improve grid energy efficiency, according to the grid energy efficiency calculation formula (2), it is necessary to improve the network speed and reduce grid energy consumption. It is desirable to eliminate APs with poor energy conditions and APs that are not frequently served by UEs in the existing AP candidate location set. Therefore, the newly selected locations for AP deployment need to have the following characteristics: First, they need to meet coverage requirements so that there is a sufficient number and coverage of AP candidate locations for selection during DDPG optimization, preventing SE deterioration in some areas due to UEs being far from all APs, thus affecting grid energy efficiency. Second, considering the uneven distribution of UEs, multiple AP candidate locations need to be retained in the same area to meet the needs of densely distributed UE areas; in addition, the supply of new energy sources needs to be considered, and it is desirable that the selected APs have high EH efficiency. Based on this, this embodiment will combine the advantages of the K-means algorithm in spatial clustering to complete the initial screening of AP candidate locations.

[0071] K-means is a distance-based unsupervised clustering algorithm. Its core objective is to divide data into multiple groups, ensuring high similarity among data points within the same group and low similarity between data points in different groups. Using the K-means algorithm, we can obtain spatially close AP (Actively Accessible Hierarchical) groups, and then select a certain number of APs with high EH (External Hierarchical Efficiency) from these groups. This allows us to identify APs with advantages in new energy sources while still ensuring the coverage of new AP candidate locations.

[0072] The objective function of the K-means algorithm is to minimize the sum of squared errors from the centroid to the sample within the group:

[0073]

[0074] Among them, Z i Let x represent the i-th group. i Let I represent the centroid of the i-th group, and let I represent the number of groups.

[0075] The K-means algorithm first randomly selects I initial centroids, and then, for a sample z, assigns it to the group containing the nearest centroid, satisfying:

[0076]

[0077] After the allocation is complete, update the centroids of the I groups, which are the means of the samples within each group.

[0078]

[0079] Stop clustering and output the results when the centroids no longer change.

[0080] The K-means algorithm can accelerate convergence and improve sample efficiency, but it is sensitive to the number of clusters and the initial centroids. If the number of clusters is too small or the initial centroids are improperly distributed, some potentially optimal regions may be merged with large clusters, causing the loss of key points in these regions during subsequent screening and leading to local optima. Therefore, by dividing the region into multiple groups and selecting multiple AP candidate locations within each group, a balance can be achieved between action dimension and sample efficiency.

[0081] Based on the fundamental principles of the K-means algorithm, the process for reducing action dimensionality is as follows: First, randomly select M APs from all candidate AP locations as center locations, with each center corresponding to a group. Then, calculate the distances of the remaining APs to these centers and classify them into the group closest to them, completing the first grouping. Afterward, continuously update the center of each group as the new AP deployment location, and regroup according to the new center until the termination condition is met. Once the termination condition is met, select the L AP candidate locations with the highest statistical EH efficiency from each group to form a new candidate AP set. This approach ensures that the newly selected AP candidate locations are spatially sufficient to cover non-cellular large-scale MIMO networks, while also exhibiting high EH efficiency. The specific process of the K-means-based AP preliminary screening method is shown in Algorithm 1.

[0082]

[0083] Step S2.2: Solve the AP deployment optimization problem based on the DDPG algorithm, where the action design is the selection state of the new AP candidate positions after being filtered by the K-means algorithm.

[0084] After performing action dimensionality reduction, the AP deployment problem is solved using the DDPG algorithm. For example... Figure 3 As shown, the DDPG algorithm is a reinforcement learning algorithm based on an "actor-critic" architecture. It combines the experience replay technique from Deep Q-Network (DQN) with the stable training techniques of the target network, and uses deterministic policy gradients for policy optimization. Unlike pure policy gradient methods that directly estimate policy gradients through Monte Carlo sampling, DDPG uses an evaluation network to provide a value benchmark for the action network (policy network), greatly reducing the variance of gradient estimation and making action network updates more accurate, independent of the sampling stability of the policy itself. Furthermore, the use of experience replay breaks the correlation between samples, allowing for the repeated use of previously generated data to improve data utilization efficiency, further enhancing the stability and generalization ability of the learning process.

[0085] The core idea of ​​the DDPG algorithm is to maximize the cumulative discount reward:

[0086]

[0087] Among them, s t Indicates the state of time slot t, a t For the action in time slot t, r(s) t ,a t ) is in state s t and action a t The reward function value is given below, where p(s0) is the initial state distribution, and φ is the action network π. φ The neural network parameters, γ∈[0,1] are discount factors.

[0088] To achieve the goals of the DDPG algorithm, four neural networks are used: an action network, an evaluation network, a target action network, and a target evaluation network. The action network generates corresponding actions based on the current state acquired from the environment; in the optimization problem, this determines the optimization variables. The evaluation network assesses the effectiveness of the current action network's policy and guides its updates. The target network is a copy of the main network used for stable training. The learning process of DRL is essentially the process of updating network parameters. DDPG uses an experience replay pool to store experience tuples (s) generated by the agent's interactions with the environment. t ,a t ,r t ,s t+1 These tuples are subsequently used for batch sampling and updates.

[0089] The evaluation network learns a Q-function to approximate the action-value function under the current policy. The Bellman equation is used to update the evaluation network, aiming to make the current Q-value close to the target value. The Q-value is defined as the expected reward of taking action a in state s and the expected cumulative reward obtained by acting according to the policy afterward. The Q-value is defined as:

[0090]

[0091] Target value y i This is generated through a target evaluation network to enhance training stability; the target value is:

[0092] y t =r t +γQ θ′ (s t+1 ,a t+1 ), (11)

[0093] Where θ′ represents the neural parameters of the target evaluation network, and a t+1 The target action is generated by the target action network:

[0094] a t+1 =π φ′ (st+1 ), (12)

[0095] Where φ′ represents the neural parameters of the target action network.

[0096] The evaluation network with neural network parameters θ updates by minimizing the mean squared error between the Q-value and the target value, and its loss function is:

[0097]

[0098] Further update θ using gradient descent:

[0099]

[0100] Where, λ Q To evaluate the learning rate of the network.

[0101] The goal of action networks is to learn policy π. φ To maximize the Q value, the objective function is:

[0102]

[0103] By evaluating which direction the network indicates yields higher returns, the gradient formula for a deterministic policy is used:

[0104]

[0105] The action network moves along this direction and is updated via gradient ascent:

[0106]

[0107] Where, λ π is the learning rate of the action network.

[0108] Using the current evaluation network to estimate future Q-values ​​can lead to instability or even divergence. Therefore, DDPG uses a slowly updated target network to enhance stability. The target network is updated using a soft update method.

[0109] φ′←τφ(1-τ)φ′, (18)

[0110] θ′←τθ+(1-τ)θ′, (19)

[0111] Where τ << 11 is the soft update coefficient, which makes the target network slowly move closer to the current network to avoid drastic changes.

[0112] Reinforcement learning frameworks are based on Multiplication Table (MDP) processes. A typical MDP process includes three key elements: state space, action space, and reward function. The AP deployment optimization problem can be established as an MDP process, with the following settings for the state space, action space, and reward function:

[0113] 1. State Space: In a real-world AP deployment problem, AP deployment cannot change according to the real-time state of different networks, but the energy and traffic characteristics of the network are fixed. Therefore, the state of the neural network is set to the current AP deployment scheme (i.e., the action in the previous time slot):

[0114] s={a1,a2,...,a I}, (20)

[0115] Where I represents the number of AP candidate locations.

[0116] 2. Action Space: Through the action network, the agent generates actions, i.e., the policy generated by DDPG, which includes the selection state for each AP candidate position:

[0117] a = {a1, a2, ..., a} I}. (twenty one)

[0118] 3. Reward Function: In the AP deployment problem, the optimization objective is to maximize EE (Effective Expiration). grid To prevent the rapid increase in rewards due to low grid energy consumption, which could negatively impact the training performance of the neural network, the calculation of EE... grid A very small positive number was added to the denominator, and the reward design is as follows:

[0119]

[0120] In optimizing AP deployment using the DDPG algorithm, the performance of the current action network generation strategy is verified through simulation, and the update of the DDPG neural network is guided. After the algorithm converges, the trained action network is used to generate the optimal deployment scheme, thereby generating a deployment strategy with high grid energy efficiency. The actual operation flow of the DDPG-based AP deployment method is as follows: first, preliminary screening is performed using Algorithm 1; then, the neural network is trained using the DDPG algorithm, where the action network will generate the selected AP candidate locations from the set... Select the AP to be deployed. The specific process of AP deployment based on DDPG is shown in Algorithm 2.

[0121]

[0122] In the embodiments, to verify the performance of the proposed DDPG-based solution, the following two benchmark schemes were compared: a random AP deployment method, with the result curve denoted as Random; and a K-means-based AP deployment method (E. Nayebi and B.D. Rao, “Access Point Location Design in Cell-Free Massive MIMO Systems”), which selects the AP with the highest EH efficiency in each group.

[0123] Figure 4 The convergence curve of average grid energy efficiency as a function of rounds for the AP deployment method based on DDPG is shown, and compared with the case without K-means action dimensionality reduction. Figure 3 After approximately 25 rounds, the proposed DDPG-based AP deployment method converged to a high level. The comparison shows that using K-means to reduce dimensionality leads to faster convergence and higher performance, while methods without initial screening, due to their large action dimensions, struggle to explore advantageous actions under the same network settings, thus hindering effective learning.

[0124] Figure 5 The graph shows the average grid energy efficiency as a function of the number of deployed access points (APs), where the number of APs is M = {20, 30, 40, 50, 60}. Figure 4 It can be seen that the average grid energy efficiency gradually decreases with the increase of APs. This is because when deploying a small number of APs, the K-means+DDPG method can select APs with high renewable energy collection efficiency. However, as the number of deployed APs increases, the locations of newly added APs are often associated with lower energy efficiency, leading to an increase in the overall grid energy consumption of the network. At the same time, as the number of APs increases, the improvement in UE's SE performance in the network is not as rapid as the increase in grid energy consumption, thus causing a decrease in grid energy efficiency.

[0125] Figure 6 The graph illustrates the variation of average grid energy efficiency with the number of user units (UEs), where the number of UEs K = {8, 10, 12, 14, 16, 18}. For any number of UEs, the proposed AP deployment method based on K-means+DDPG outperforms both the K-means and Random benchmark schemes. As the number of UEs increases, the proposed algorithm provides better performance and steadily improves, demonstrating the ability of the proposed DDPG-based AP deployment method to learn and adapt to new network characteristics.

[0126] Figure 7The graph shows the variation of average grid energy efficiency with EH efficiency, where EH efficiency ranges from an initial setting of 50% to 100%, increasing incrementally in 10% increments. The graph demonstrates that the proposed DDPG-based method outperforms both K-means and Random methods. Furthermore, it shows that as EH efficiency increases, the AP deployment methods based on K-means and DDPG gradually exhibit greater advantages. This is because in the scenario of non-cellular massive MIMO with uneven UE distribution and new energy-enabled deployment, the K-means-based method can deploy APs at locations with higher EH efficiency. The proposed K-means+DDPG method, through K-means grouping, initially selects a better set of AP deployment locations, reducing grid energy consumption. Simultaneously, by learning network features through deep neural networks, it achieves higher spectral efficiency, ultimately realizing high grid energy efficiency.

[0127] Based on the same inventive concept, this invention also discloses a new energy-enabled non-cellular large-scale MIMO intelligent AP deployment system based on K-means and DDPG, comprising: a problem construction module, used to construct an AP deployment optimization problem in a new energy-enabled scenario, selecting a specified number of APs from all non-uniformly distributed AP candidate locations, coordinating with UEs in the network, and achieving the optimization objective of maximizing average grid energy efficiency under battery energy storage constraints; a clustering and dimensionality reduction module, used to use a geographic-energy dual-dimensional clustering mechanism, clustering all AP candidate locations based on the K-means algorithm, and selecting multiple AP candidate locations with the highest statistical EH efficiency in each group to form a new set of AP candidate locations, thereby reducing the action dimension of the DDPG algorithm; and a DDPG solution module, used to solve the AP deployment optimization problem based on the DDPG algorithm, wherein the action design is the selection state of the new AP candidate locations after K-means algorithm screening.

[0128] This invention also discloses a computer system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the new energy-enabled non-cellular large-scale MIMO smart AP deployment method based on K-means and DDPG.

[0129] This invention also discloses a computer program product, including a computer program that, when executed by a processor, implements the steps of the new energy-enabled non-cellular large-scale MIMO smart AP deployment method based on K-means and DDPG.

[0130] The program code used to implement the method of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the steps of the method of the present invention to be performed. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a standalone software package, or entirely on a remote machine or server. All aspects not detailed in this invention are well-known to those skilled in the art.

Claims

1. A new energy enabled cell-free massive MIMO intelligent AP deployment method based on K-means and DDPG, characterized in that, The method comprises the following steps: An AP deployment optimization problem in a new energy empowerment scenario is constructed, a specified number of APs are selected from all AP candidate positions in a non-uniform distribution, UEs in a service network are cooperated, and an optimization target of maximizing average grid energy efficiency under a battery energy storage constraint is realized; A geographic-energy dual-dimensional clustering mechanism is adopted, K-means algorithm is used to cluster all AP candidate positions, and a plurality of AP candidate positions with the highest statistical EH efficiency are selected in each group to form a new AP candidate position set, so as to reduce the action dimension of the DDPG algorithm; The AP deployment optimization problem is solved based on the DDPG algorithm, and the action is designed as the selection state of the new AP candidate position selected by the K-means algorithm.

2. The K-means and DDPG-based new energy-enabled cell-free massive MIMO intelligent AP deployment method according to claim 1, characterized in that, The optimization problem is represented as: wherein, T is the number of time slots, M is the number of selected APs, EE grid (τ) is the grid energy efficiency in the τth time slot, m (τ) is the battery energy storage available to the mth AP in the τth time slot, max is the battery energy storage upper bound.

3. The K-means and DDPG-based new energy-enabled cell-free massive MIMO smart AP deployment method of claim 1, wherein The steps of preliminarily screening the AP candidate positions based on the K-means algorithm comprise: M APs are randomly selected as center positions in all AP candidate positions, each center corresponds to a group, and M is the number of APs specified in the optimization problem; The distances of the remaining APs to the centers are calculated and classified into the nearest groups; The centers of the groups are updated as new AP deployment positions, and the groups are re-grouped according to the new centers until the termination condition is reached; After the termination condition is reached, L≥2 AP candidate positions with the highest statistical EH efficiency are selected in each group to form a new candidate AP set.

4. The K-means and DDPG-based new energy-enabled cell-free massive MIMO smart AP deployment method of claim 1, wherein The state space in the DDPG algorithm is a current AP deployment scheme s = {a1, a2,..., a I}, I represents the number of AP candidate positions, the action space is a = {a1, a2,..., a I}, and the reward is , wherein represents the SE of the kth UE when the AP combination is , K is the number of UEs, B is the communication bandwidth, represents the grid energy consumption of the mth AP, and ξ is a positive number for preventing the reward from rapidly increasing due to small grid energy consumption.

5. The K-means and DDPG-based new energy-enabled cell-free massive MIMO smart AP deployment method according to claim 1, characterized in that, The grid energy efficiency considers the spatial difference of renewable energy collection efficiency and the non-uniformity of user distribution. The grid energy consumption of the mth AP at the τth time slot is represented as where P m (τ) is the total energy consumption of the mth AP, including the downlink transmission power and the fixed circuit loss, b m (τ) represents the time-varying characteristics of the battery state, where is the new energy collected by the mth AP at the τth time slot.

6. A new energy enabled cell-free massive MIMO smart AP deployment system based on K-means and DDPG, characterized in that, comprise: A problem construction module is configured to construct an AP deployment optimization problem in a new energy empowerment scenario, select a specified number of APs from all AP candidate positions in a non-uniform distribution, cooperate with UEs in a service network, and realize an optimization target of maximizing average grid energy efficiency under a battery energy storage constraint. A clustering dimension reduction module is configured to adopt a geographic-energy dual-dimensional clustering mechanism, cluster all AP candidate positions based on K-means algorithm, and select a plurality of AP candidate positions with the highest statistical EH efficiency in each group to form a new AP candidate position set, so as to reduce the action dimension of the DDPG algorithm. A DDPG solving module is configured to solve the AP deployment optimization problem based on the DDPG algorithm, and the action is designed as the selection state of the new AP candidate position selected by the K-means algorithm.

7. The K-means and DDPG-based new energy enabled cell-free massive MIMO smart AP deployment system of claim 6, wherein, The optimization problem is represented as: wherein, T is the number of time slots, M represents the number of selected APs, EE grid (τ) is the grid energy efficiency of the τth time slot, b m (τ) is the battery energy storage available to the mth AP in the τth time slot, b max is the upper limit of the battery energy storage.

8. The K-means and DDPG-based new energy enabled cell-free massive MIMO smart AP deployment system of claim 6, wherein, The state space in the DDPG algorithm is a current AP deployment scheme s = {a1, a2,..., a I}, I represents the number of AP candidate positions, the action space is a = {a1, a2,..., a I}, and the reward is , wherein represents the SE of the kth UE when the AP combination is , K is the number of UEs, B is the communication bandwidth, represents the grid energy consumption of the mth AP, and ξ is a positive number for preventing the reward from rapidly increasing due to small grid energy consumption.

9. A computer system comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The computer program is executed by the processor to realize the steps of the K-means and DDPG-based new energy empowerment cell-less large-scale MIMO intelligent AP deployment method according to any one of claims 1-5.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the K-means and DDPG-based new energy empowerment cell-less large-scale MIMO intelligent AP deployment method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Service function chain low-cost intelligent deployment method based on environmental perception

    CN111093203A

  • Ultra-dense networking-oriented dynamic resource allocation method

    CN113490219A

  • Multi-agent learning method for non-cellular network user scheduling and resource configuration

    CN117221925A

  • Multi-unmanned aerial vehicle communication system optimization control method based on internal curiosity mechanism

    CN118647032A

  • Resource allocation method and system of cellular-free large-scale system, and storage medium

    CN119545426A