Intelligent surface-assisted cellular-free network system power optimization method
By optimizing AP transmit power and RIS phase shift through a smart surface-assisted non-cellular network system, and combining fractional programming and deep reinforcement learning, the problems of high energy consumption and uneven resource allocation in non-cellular networks are solved, achieving efficient and green communication.
Patent Information
- Application Number
- CN202511909801.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-03-03
AI Technical Summary
Cellular non-cellular networks have high hardware costs and energy consumption when deployed on a large scale. Their fixed signal characteristics cannot be actively optimized, resulting in uneven resource allocation and degraded cell edge performance.
A smart surface-assisted cellular network system is adopted to optimize system power by optimizing AP transmit power, RIS phase shift, and AP selection, combined with a near-end policy optimization method using fractional programming and deep reinforcement learning.
It significantly reduces total system power consumption, improves received signal quality, ensures user service quality, achieves a balance between high capacity and high energy efficiency, and adapts to dynamic wireless environments.
Smart Images

Figure CN121604091A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication, and more particularly to a power optimization method for a smart surface-assisted non-cellular network system. Background Technology
[0002] With the continuous development of mobile communication technology, cellular networks, as a simple and efficient basic network architecture, have been widely used in many commercial communication scenarios. However, as users' basic communication needs continue to increase, cellular network architectures are gradually failing to achieve ideal performance indicators such as high frequency bands and low latency. Furthermore, these architectures also suffer from high energy consumption, uneven resource allocation, and performance degradation at cell edges. Especially in communication applications requiring large-scale deployment, cellular networks struggle to efficiently utilize communication resources. Based on this, researchers have recently proposed cell-free network architectures. This architecture involves densely distributing a large number of low-cost, low-power access points (APs) within the service area and connecting them to a central processor via fiber optic or wireless links to work collaboratively and serve all users within the network. Deploying cell-free networks can achieve higher reliability, spectral efficiency and system capacity, lower latency, stronger massive device connectivity capabilities, and a more uniform and stable service effect.
[0003] Deploying cellular-free networks requires a large number of access points (APs), each needing a complete radio frequency link. This increases hardware costs, deployment costs, and system power consumption. Furthermore, in typical communication scenarios, once APs are deployed, the signal transmission and reception characteristics are essentially fixed, making proactive optimization of the wireless environment impossible. The emergence of Reconfigurable Intelligent Surface (RIS) technology can effectively address these issues. A RIS is a plane composed of numerous passive metamaterial units. It can intelligently encode and control each metamaterial unit, thereby reconstructing the phase, amplitude, polarization, and other characteristics of the incident electromagnetic wave. Using RIS can enhance received signal quality, improve signal coverage, and provide APs with a more ideal channel environment. Summary of the Invention
[0004] Purpose of the invention: The technical problem to be solved by the present invention is to provide a power optimization method for a smart surface-assisted non-cellular network system, which addresses the shortcomings of the existing technology.
[0005] To address the aforementioned technical problems, this invention discloses a power optimization method for a smart surface-assisted non-cellular network system, comprising the following steps:
[0006] Step 1: The sum of AP transmit power and AP operating power in a smart surface-assisted non-cellular network system is used as the optimization objective; AP selection, AP transmit power, and phase shift of each RIS unit are used as optimization variables; and the power optimization problem is established by combining the AP transmit power constraint, the RIS phase shift constraint, and the data capacity constraint of each fronthaul link.
[0007] Step 2: Optimize the beamforming matrix of the AP with fixed AP selection and phase shift of each RIS unit, and then use a convex optimization method based on fractional programming (FP) to solve for the AP's transmit power.
[0008] Step 3: Fix the beamforming matrix of AP, and use the proximal policy optimization (PPO) method in deep reinforcement learning (DRL) to solve for the AP selection and the phase shift variables of each RIS unit;
[0009] Step 4: After the results of Step 3 stabilize, the above algorithm process is combined to finally realize an online and efficient intelligent surface-assisted power optimization method for cellular network-free systems.
[0010] In this invention, uppercase bold letters represent matrices, lowercase bold letters represent vectors, and lowercase letters represent scalars. Representation matrix The List, Representation matrix The Column, number Row element. Representing vectors The Each element. The dimension is The complex field, Let A represent a vector consisting of the diagonal elements of A. express An identity matrix of dimension 1, with superscript , They represent the transpose and conjugate transpose of a matrix, respectively, with superscripts... These represent the optimal solution and the value at the t-th iteration, respectively. (Backslash) This indicates that something is excluded from the set.
[0011] Further, step 1 includes: setting up a non-cellular network with There are 1 AP, and the number of antennas on each AP is 1. Meanwhile, the total number of users in the system is By default, all user equipment uses a single antenna. The system includes a RIS (Resource Identifier System) that can be used to assist communication; the number of components on the RIS is... The AP set, user set, and RIS set are denoted as follows: From AP To RIS, from RIS to users From AP To users The channels are respectively denoted as , , The RIS is equipped with an intelligent controller, which allows it to share channel state and RIS phase shift information with the base station via a separate fronthaul link.
[0012] For users In comparison to AP The beamforming vector is The magnitude of this vector actually represents AP. Assigned to user The power. Correspondingly, for the AP In this regard, the beamforming matrix composed of the beamforming vectors of all users is: The overall beamforming matrix of the system can merge the beamforming matrices of each user into a single matrix. For RIS, its first... The reflection coefficient of each element is ,in This represents the phase shift coefficient of the component. Based on this, the phase shift vector of the RIS can be written as: For simplicity, a phase shift matrix can also be introduced. The phase shift is represented in the following way: .
[0013] For power optimization problems involving AP selection, since not all APs are operational in the system, the set of operational APs can be defined as follows: , and AP The set of users served can be denoted as When the user Not in the set In the middle, the corresponding beamforming vector According to this definition, the selection of the AP can be represented without introducing additional variables, but is implicit in the beamforming vector between the AP and the user.
[0014] Because the system itself is a non-cellular network, each user can be served by a different access point (AP). The received data is denoted as (0, 1), when receiving a signal, the sum of all noise can be regarded as Gaussian white noise, denoted as (0, Based on the symbols above, we can obtain the user's information. The expression for the received signal:
[0015]
[0016] The first term in the expression represents the user. In terms of valid signals, the second term represents the remaining users' feedback to users. The interference signal, with the last term representing Gaussian white noise, can be used to derive the expression for the signal-to-interference-plus-noise ratio (SINR):
[0017]
[0018] Because when users Not in the set In the middle, the corresponding beamforming vector Therefore, the signal-to-interference-plus-noise ratio (SIR) expression can actually be transformed into the following form:
[0019]
[0020] The system's optimization objectives can be divided into AP transmit power and operating power. The AP's transmit power is actually the sum of the powers of the beamforming vector, which can be denoted as... The operating power of the AP is its static power. Static power can be denoted as The system's operating power is By jointly optimizing the AP working set Beamforming matrix Phase shift vector Ultimately, the goal is to achieve the aforementioned transmit power and operating power... + To reach the minimum.
[0021] During optimization, we simultaneously consider the following constraints: user SINR constraints are needed to ensure the quality of the received signal at the user's location; transmit power constraints and fronthaul link capacity constraints at the AP are needed to ensure that the AP can transmit signals at the appropriate power and rate; and phase shift constraints at the RIS are needed to ensure that the RIS's phase adjustment conforms to the actual situation. Based on these constraints, we present the form of the power optimization problem for smart surface-assisted non-cellular networks:
[0022] +
[0023]
[0024]
[0025]
[0026]
[0027] Step 2 includes: fixing the AP working set and phase shift vector The following algorithm is used for beamforming matrices Optimize
[0028] When variables With phase shift vector When the variables are fixed, removing the constants arising from the fixed variables, the problem transforms into the following form, which is a fractional programming problem:
[0029]
[0030]
[0031]
[0032] Analyzing the simplified optimization problem reveals that, apart from the signal-to-interference-plus-noise ratio (SIR) constraint, the objective function and power constraint have essentially become a convex problem with convex constraints. Furthermore, the non-convex SIR constraint can be transformed into a convex constraint using a quadratic transform in fractional programming. Under these conditions, the original problem can be directly solved using standard convex optimization methods. Through this strategy, we can achieve power optimization.
[0033] Step 3 includes: fixing the beamforming matrix The following algorithm is used for the AP working set and phase shift vector Optimize.
[0034] When beamforming matrix After being fixed, the problem becomes the following form:
[0035] +
[0036]
[0037]
[0038]
[0039]
[0040] Since the problem now transforms into a discrete optimization problem, and because the channel state information changes periodically over time, we need an online algorithm that can dynamically adjust the optimal solution based on parameter changes. For this type of scenario, we can use the PPO algorithm from the DRL domain. Below, we briefly introduce the principle of the PPO algorithm and the approach to solving discrete optimization problems with fixed variables using this type of algorithm:
[0041] The Proximal Policy Optimization (PPO) algorithm is a reinforcement learning algorithm whose core idea is to enable an agent to steadily improve its decision-making strategy by applying appropriate constraints during its interaction with the environment, thus avoiding instability or even divergence caused by a single parameter update being too large.
[0042] Essentially, the core idea of the PPO algorithm is to use pruning to replace the objective function. When an agent updates its behavioral policy, it compares the probability difference between the old and new policies in choosing a particular action. If the new policy differs significantly from the old one, the probability ratio will deviate from the normal range. The core of the PPO algorithm is to improve this problem of overly rapid updates. It prunes drastic changes back to a relatively reasonable range by setting a confidence interval, making each policy adjustment relatively stable. Simultaneously, the PPO algorithm also uses multiple epochs of mini-batch updates to perform multiple randomized small-batch learnings on the empirical data collected from the agent's interactions with the environment. This data utilization strategy ultimately improves sample utilization efficiency, allowing the agent to extract more information from limited empirical data.
[0043] In summary, the PPO algorithm's working steps can be summarized as follows: First, the agent interacts with the environment using the current policy and collects experiential data based on this. Then, a dedicated value network evaluates the merits of each decision made in these experiences. Next, PPO optimizes the policy network multiple times based on the value network's evaluation, using a defined pruning objective function. Finally, PPO can input the updated policy into a new round of interaction, initiating the next learning cycle. In summary, the overall process of the PPO algorithm can be summarized as follows: Figure 2 .
[0044] For this subproblem, the variable to be optimized has now been transformed into the AP working set. and phase shift vector The problem has now become a discrete optimization problem. For discrete optimization problems, the complexity of the algorithm often increases exponentially with the problem size, making traditional search algorithms unsuitable. The PPO algorithm, however, offers a different perspective. Based on the dynamic nature of the problem channel, it transforms a simple combinatorial optimization problem into a sequential decision-making process, allowing the agent to learn from past experiences and explore possible optimal solution strategies.
[0045] Specifically, we can use the system's channel For the state, the system's optimization variable AP working set and phase shift vector For each action, a policy and value network is constructed using the total power consumed by the system as the reward. To learn the spatial correlation features of the channel, the network structure can be built using a CNN + fully connected layers, with the sigmoid function as the activation function. The agent is then repeatedly trained and its policy saved on a large number of problem instances. Ultimately, for each instance, the agent can start from scratch and gradually construct a complete solution using the current policy. The PPO algorithm can then update the policy network based on the quality of the final solution, increasing the probability of selecting actions leading to higher rewards in similar situations in the future.
[0046] Step 4 includes: repeating the above iterative optimization process, using the PPO algorithm and the system's online output for iteration and training, so that the DRL network deployed in step 3 continuously approaches the optimal policy.
[0047] Step 5 includes: obtaining a relatively stable multi-functional optimization algorithm for AP selection, beamforming, and phase shift adjustment based on the network continuously trained in Step 4.
[0048] Beneficial effects:
[0049] Currently, large-scale deployments of non-cellular networks result in significant hardware, deployment, and energy consumption due to the large number of access points (APs) and radio frequency links. To reduce overall network power consumption and improve received signal quality, this invention proposes a RIS-assisted power optimization method for non-cellular networks. This method, by jointly optimizing multiple decision variables such as AP selection, transmit beamforming matrix, and RIS phase shift, can significantly reduce overall system power consumption (including transmit power and AP static operating power) while maintaining user service quality (SINR ratio), ultimately achieving a balance between high capacity and high energy efficiency.
[0050] In terms of specific solutions, this method addresses the MINLP problem, which consists of a continuous variable beamforming matrix and discrete variables AP selection and RIS phase shift. It employs an alternating optimization approach to solve the discrete and continuous variables separately, reducing computational complexity while maintaining accuracy. For the power allocation subproblem in solving the continuous variables, this method uses fractional programming to accurately solve for the optimal beamforming matrix. For the AP selection and RIS phase shift optimization problem in solving the discrete variables, it addresses the challenges of discrete decision-making and time-varying channel conditions. Using a PPO-based DRL algorithm, it learns online and adapts to the dynamically changing channel environment, seeking near-optimal decision strategies from historical experience, thus avoiding the excessive computational complexity of traditional search methods in large-scale problems.
[0051] Meanwhile, the algorithm explicitly incorporates the static operating power of the access points (APs) into the optimization objective and proactively shuts down APs with low contribution to current user service and poor channel gain through an AP selection mechanism. This reduces unnecessary energy loss and fronthaul link power consumption, overcoming the power surge problem caused by node density in non-cellular networks, and laying a technical foundation for building a green 6G network. Furthermore, the optimization methods for both types of variables are adaptive, capable of online operation and adapting to changes in channel conditions. The stable update mechanism of the fractional programming algorithm and the PPO algorithm ensures the reliability of policy learning, enabling the system to continuously provide near-optimal power control and resource allocation schemes in the face of dynamic wireless environments, demonstrating good robustness.
[0052] In summary, this algorithm, employing alternating optimization and combining fractional programming with the PPO algorithm, successfully solves the relatively complex power optimization problem in smart surface-assisted non-cellular networks. It not only significantly reduces total system power consumption but also ensures user communication quality, while exhibiting good online adaptability and robustness. It represents a promising key technology for realizing the high capacity and high energy efficiency vision of future 6G networks. Attached Figure Description
[0053] Figure 1 This is a system model.
[0054] Figure 2 This is the main update process of the PPO algorithm.
[0055] Figure 3 The process for updating the overall power algorithm of the system.
[0056] Figure 4 This is a comparison chart of the cumulative power distribution of the system under different numbers of RIS according to the method of the present invention.
[0057] Figure 5 This is a convergence rate graph showing how the method of the present invention gradually converges as the number of update steps changes.
[0058] Figure 6 This is a performance comparison chart of the method of the present invention with other algorithms under the condition of changing user service requirements. Detailed Implementation
[0059] In traditional non-cellular network systems, power optimization is a crucial and challenging problem. Its main objective is to allocate appropriate transmit power to a large number of distributed access points (APs) in the network to suppress inter-user interference, reduce overall system energy consumption, and improve data rates. However, the introduction of a novel device called a Reflection Array (RIS) alters the form of the power optimization problem, increasing its complexity. For simplicity, in this system model, we assume the RIS is a passive element that adjusts the phase of the transmitted signal only through reflector units, without changing the signal amplitude. Therefore, for a RIS-assisted non-cellular network system, we only need to simultaneously adjust the transmit power of the APs and the phase of each reflector unit on the RIS, thereby improving service quality, reducing overall energy consumption, and ultimately achieving green communication.
[0060] Meanwhile, for non-cellular networks, due to the large number of access points (APs), turning on all APs simultaneously would consume a considerable amount of unnecessary energy. Furthermore, since APs in non-cellular networks need to transmit information via the fronthaul link, the overall power consumption of the fronthaul link would be substantial if all APs were turned on. Therefore, to reduce total system power consumption, power optimization should consider both AP on / off states. This involves appropriately shutting down APs that contribute little to the current system's users or have poor channel conditions, ultimately achieving a balance between system performance and power consumption, avoiding a situation of high performance but low energy efficiency.
[0061] From the above perspectives, for smart surface-assisted noncellular network systems, the power optimization problem can be viewed as a comprehensive network optimization problem integrating AP selection, power allocation, and RIS phase modulation. This problem involves both discrete and continuous decision-making, making it a Mixed Integer Nonlinear Programming (MINLP) problem. Effectively solving this problem can help overcome the power consumption surge caused by node densification in noncellular networks. Through deep and global power optimization, its potential for high capacity and high energy efficiency can be released and realized, laying a solid foundation for 6G green networks.
[0062] The specific simulation parameters are as follows: Consider a smart surface-assisted non-cellular network where four access points (APs) are evenly distributed on a circle with a center of (100m, 100m) and a radius of 300m, each AP equipped with three antennas. Eight single-antenna users are randomly distributed within a circle with a center of (100m, 100m) and a radius of 50m. In this scenario, the direct link between the APs and users will be blocked by buildings or other obstacles, resulting in relatively high loss. Meanwhile, the Reflection Path Receiver (RIS) is placed at a higher position between the APs and users to provide a better reflection path. Based on the actual situation, we mainly consider the case where the phase value of the RIS is discrete, and its values belong to the following set, where... Indicates the number of RIS phase-resolved bits:
[0063]
[0064] Unless otherwise specified, network training and simulation testing will be performed using the parameter settings in Table 1 below by default:
[0065] Table 1 Simulation Parameter Settings
[0066] parameter value Number of APs 4 Number of users 8 Number of antennas per AP 3 Channel and Path loss -30dBm Channel Path loss -90dBm Transmission bandwidth 180kHz Noise power spectral density -170dBm / Hz Rice factor 10 AP maximum transmit power 33dBm AP operating power 5.5w
[0067] For the policy network used in PPO, to preserve the spatial correlation of multi-dimensional inputs such as the channel state matrix, we can choose a CNN (Convolutional Neural Network) for feature extraction and action quantization to avoid spatial information loss. Simultaneously, we use the Sigmoid function as the activation function so that the network can approximate any continuous mapping. During training, the number of epochs is set to 500, the batch size to 256, the empirical replay pool size to 10000, and the learning rate to [value missing]. The Adam optimizer was used. All experiments were implemented using PyTorch on an Intel Core i7-11800H (2.30GHz) and an NVIDIA GeForce RTX 3060 GPU (12GB).
[0068] Furthermore, to evaluate the effectiveness of the proposed power optimization algorithm, this invention sets up two baseline methods for comparison: the AP selection and RIS phase-shift greedy search method and the polynomial approximation optimization method. These methods differ from the algorithm of this invention in that their main optimization idea is based on convex function approximation, aiming to obtain the optimal solution to the problem through multiple iterative iterations. Compared with the algorithm of this invention, these algorithms have relatively higher complexity and a relatively longer convergence time to achieve ideal performance. The final results are illustrated below through a comparison of simulation results:
[0069] Figure 4 The graph shows the cumulative distribution curves of the algorithm across 2000 implementations as the system size and number of RIS (Resource Identifiers) increase. The horizontal axis represents the total power consumption of the system, and the vertical axis represents the cumulative distribution probability. The graph shows that when the system size is small, the algorithm effectively reduces system energy consumption. Furthermore, as the system size increases, the energy consumption increases in an approximately linear manner. Meanwhile, if the number of APs in the system is much larger than the number of users, the energy consumption generated by this algorithm does not increase significantly. Instead, it can dynamically adjust to approximate the cumulative distribution curve of a small system, demonstrating the algorithm's good decision-making performance. Regarding the RIS phase variable, as the number of RIS components increases, the system power consumption decreases accordingly, reflecting the practical role of RIS in actively regulating the channel environment and saving total system energy.
[0070] Figure 5 The graph illustrates the convergence speed of the algorithm, with the horizontal axis representing the number of iterations and the vertical axis representing the total system loss. As can be seen, under slight changes in the channel environment, the algorithm gradually learns the mapping from the channel environment to AP selection and RIS phase shift variables over time, eventually obtaining a relatively stable strategy. Although the system's performance loss fluctuates from the perspective of instantaneous power, on average, its total power continuously decreases with online training, eventually reaching a relatively stable optimal value.
[0071] Figure 6 The performance of the proposed algorithm is compared with other types of power optimization algorithms, where the horizontal axis represents the User Service Quality Requirement (SINR) and the vertical axis represents the total power consumed by the system. The specific power optimization algorithms selected are as described above: the RIS phase-shift greedy search method and the polynomial approximation optimization method are chosen for AP. From the results, we can see that the algorithm proposed in this invention achieves better performance than the other two methods, and due to its online training and low complexity, the time and resources required to obtain the optimal solution are also lower than the other two types of algorithms. Furthermore, combined with… Figure 4 The simulation results also show that the introduction of RIS significantly improves the system power optimization results. This demonstrates that such devices can play a role in practice, helping us to better allocate and optimize resources in communication networks.
[0072] The results above show that when the number of access points (APs) in the system is similar to the number of users, this method can achieve efficient and energy-saving communication in a non-cellular network while ensuring user communication quality. However, if the number of APs in the system is much greater than the number of users, this method will not use all APs simultaneously to serve users in a high-energy-consuming manner. Instead, it will effectively and dynamically control the switching on and off of APs while meeting the basic communication quality requirements of users, thereby reducing the overall energy consumption of the system and achieving the ultimate goal of green communication.
[0073] Furthermore, considering the above results, this method focuses on solving the power optimization problem of a smart surface-assisted non-cellular network system by employing the basic idea of alternating optimization and combining fractional programming with near-end policy optimization. Compared to other methods, this method has the following advantages: 1. For dynamically changing channels in real-world scenarios, a dynamic adjustment algorithm is designed, which can actively adjust decision variables such as beamforming matrix, phase shift vector, and AP selection according to changes in channel state, ensuring that the final output power remains within a relatively low range. 2. The algorithm can be trained and solved online. Compared to many data-driven intelligent algorithms, this method can update the policy network and value network online using currently stored data, and then apply the updated results to the problem solution, without having to separate training and solving into two modules. 3. The fractional programming and near-end policy optimization methods used in the algorithm itself have relatively low complexity. Compared to many algorithms based on traditional optimization methods with relatively high complexity, its efficiency in solving the power optimization problem is significantly improved.
[0074] This invention provides a smart surface-assisted optimization method for a cellular network-free system. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.
Claims
1. A power optimization method for a smart surface-assisted cellular network system, characterized in that, The main steps include: Step 1: The sum of AP transmit power and AP operating power in a smart surface-assisted non-cellular network system is used as the optimization objective. The AP selection, AP transmit power, and phase shift of each RIS unit are used as optimization variables. The power optimization problem is established by combining the AP transmit power constraint, the RIS phase shift constraint, and the data capacity constraint of each fronthaul link. Step 2: Optimize the beamforming matrix of the AP with fixed AP selection and phase shift of each RIS unit, and then use a convex optimization method based on fractional programming to solve for the AP's transmit power. Step 3: Fix the beamforming matrix of AP and iteratively solve for the AP selection and the phase shift variables of each RIS unit; Step 4: After the results of Step 3 stabilize, combine the above steps to finally realize the power optimization method for a smart surface-assisted non-cellular network system.
2. The method for power optimization of a smart surface-assisted non-cellular network system according to claim 1, characterized in that, The optimization objectives of the system described in step 1 are AP transmit power and AP operating power.
3. The method for power optimization of a smart surface-assisted non-cellular network system according to claim 2, characterized in that, The transmit power of the AP is actually the sum of the power of the beamforming vector.
4. The power optimization method for a smart surface-assisted non-cellular network system according to claim 3, characterized in that, The operating power of the AP is its static power.
5. The intelligent surface-assisted power optimization method for a cellular-free network system according to claim 4, characterized in that, The form of the smart surface-assisted non-cellular network power optimization problem is as follows: + in, The sum of the power of the beamforming vector, AP The j-th AP, Indicates AP The set of users served The system's operating power and For beamforming vectors, The phase shift matrix represents the user's constraints from top to bottom. SINR constraint, AP Transmit power constraints, AP The forehaul capacity constraint and the phase shift constraint of RIS, where, For users SINR minimum service quality requirements For AP Maximum transmit power, For AP Maximum fronthaul capacity, This is the candidate set of phase shift vectors.
6. The method for power optimization of a smart surface-assisted cellular network system according to claim 2, characterized in that, Step 2 specifically includes: Step 2-1: Fix the AP working set and phase shift vector For beamforming matrix Optimize: When variables With phase shift vector When the variables are fixed, removing the constants arising from the fixed variables, the problem transforms into the following form, which is a fractional programming problem: Step 2-2: Analyze the fractional programming problem. Its objective function and constraints have become a convex problem and convex constraints, so it can be solved directly using the standard convex optimization method.
7. The method for power optimization of a smart surface-assisted cellular network system according to claim 2, characterized in that, Step 3 involves fixing the beamforming matrix, transforming the optimization problem into a discrete optimization problem, which is then solved using the PPO algorithm.
8. The power optimization method for a smart surface-assisted cellular network system according to claim 7, characterized in that, Specifically, it includes: Step 3-1, using the channel For the state, the system's optimization variable AP working set and phase shift vector For each action, a policy network and a value network are constructed with the total power consumed by the system as the reward. Then, an agent composed of the policy network, the value network, and training parameters is trained and the policy is saved on the problem instance. Finally, for each instance, the agent can start from scratch and gradually build a complete solution through the current policy. Step 3-2: Based on the PPO algorithm, update the policy network according to the quality of the final solution, so that in similar states in the future, the probability of selecting actions that lead to higher rewards will be increased. Step 3-3: Repeat the iterative optimization process of Step 3-2. Through the PPO algorithm and the online output of the system, iterate and train to make the deployed DRL network continuously approach the optimal policy.
9. A power optimization method for a smart surface-assisted non-cellular network system according to claim 6, characterized in that, The policy network is constructed using a CNN with fully connected layers, while the value network is constructed using fully connected layers. Both types of networks use the Sigmoid function as the activation function.
10. A power optimization method for a smart surface-assisted cellular network system according to claim 2, characterized in that, Step 4 specifically involves obtaining a relatively stable multi-functional optimization algorithm for AP selection, beamforming, and phase shift adjustment based on the network continuously trained in Step 3.