DRL-based active RIS-assisted MISO communication system joint optimization method

Through intelligent DRL-based algorithms to optimize the power distribution and location deployment of active RIS, the optimization problem of active RIS in 5G/6G wireless communication systems is solved, significantly improving users and speed.

CN120151897APending Publication Date: 2025-06-13GUILIN UNIV OF AEROSPACE TECH +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510349638.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively optimize the power distribution, location deployment and joint optimization of active RIS in 5G/6G wireless communication systems, resulting in limited system performance.

Method used

Using a deep reinforcement learning (DRL)-based method, intelligent algorithms are designed through DDPG algorithms, the base station transmit beamforming matrix and active RIS phase shift matrix are optimized, and joint optimization of power distribution and position deployment is achieved.

Benefits of technology

It significantly improves users and speed, breaks through the limitations of traditional optimization methods, and improves system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151897A_ABST
    Figure CN120151897A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of wireless communication, in particular to a DRL-based joint optimization method for an active RIS-assisted MISO communication system, which comprises the following steps of: establishing a joint optimization task by taking a base station transmitted beam forming matrix and an active RIS phase shift matrix as optimization variables, integrating the joint optimization task into a DRL framework, and designing a DRL algorithm based on a depth deterministic strategy gradient, so as to realize the joint optimization of the active RIS-assisted MISO communication system. And joint optimization of a base station transmitting beam forming matrix, an RIS phase shift matrix, power distribution and position deployment is realized through an intelligent algorithm. Specifically, communication between a base station and a user is assisted and enhanced by an active RIS, and the active RIS optimizes a wireless channel by amplifying an incident signal and adjusting the phase; according to the system, a base station transmitting beam forming matrix and a phase shift matrix of an active RIS are jointly optimized, and the user sum rate is maximized by optimizing power distribution and position selection design. Through simulation verification, the limitation of a traditional optimization method is broken through, and the user speed is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technologies, and particularly to a joint optimization method for an active RIS-assisted MISO communication system based on DRL. Background Art

[0002] In 5G / 6G wireless communication systems, reconfigurable intelligent surfaces (RISs) have become one of the key technologies for improving communication performance by dynamically regulating the wireless channel environment. Traditional passive RISs consist of low-cost passive reflection elements and can only adjust the signal phase. However, due to the double-path propagation of signals through "base station - RIS - user", multiplicative fading effects are caused, which limits the system capacity gain. Active RISs can break through the performance bottleneck of passive RISs by introducing signal amplification functions for reflection elements, but at the same time, they also bring problems such as high power allocation, location deployment, and high complexity of joint optimization.

[0003] In the prior art, the optimization of active RISs mainly relies on numerical methods such as alternating optimization and convex approximation, which have defects such as high computational complexity, poor real-time performance, and difficulty in dealing with high-dimensional continuous variables. In recent years, deep reinforcement learning (DRL) has been introduced into the field of wireless communication due to its adaptive decision-making ability in complex dynamic environments. However, existing DRL schemes mostly focus on the optimization of passive RISs and do not address challenges specific to active RISs such as power amplification constraints and joint optimization of mixed variables. Summary of the Invention

[0004] The purpose of the present invention is to provide a joint optimization method for an active RIS-assisted MISO communication system based on DRL, aiming to achieve the joint optimization of the base station transmit beamforming matrix, the RIS phase shift matrix, power allocation, and location deployment, and break through the limitations of traditional optimization methods.

[0005] To achieve the above purpose, the present invention provides a joint optimization method for an active RIS-assisted MISO communication system based on DRL, including the following steps:

[0006] Step 1: Taking the base station transmit beamforming matrix and the active RIS phase shift matrix as optimization variables, establish a joint optimization task and transform the optimization problem into a DRL task;

[0007] Step 2: Design a DRL algorithm based on DDPG and use the DDPG algorithm to solve the joint optimization problem of continuous beamforming matrices and phase shift matrices;

[0008] Step 3: Construct an active RIS-assisted MISO system model;

[0009] Step 4: Conduct optimization from two dimensions of power allocation and location optimization to obtain a DRL-based dynamic power allocation strategy and a selection strategy for the deployment location of the active RIS.

[0010] Step 5: Simulation verification.

[0011] Optionally, in Step 1, with the goal of maximizing the system sum rate, the joint optimization problem of the base station transmit beamforming matrix and the active RIS phase shift matrix is modeled as a non-convex optimization problem.

[0012] The process of transforming the optimization problem into a DRL task is specifically to transform the joint optimization problem into a dynamic optimization task under the DRL framework by defining the state space, action space, and reward function, where the state includes the channel state information and the current system configuration, the action is the base station transmit beamforming matrix and the active RIS phase shift matrix, and the reward is the user sum rate.

[0013] Optionally, the process of Step 2 for designing the DRL algorithm based on DDPG includes the following steps:

[0014] Step 2.1: Construct the DDPG network structure, including designing the action network and the policy network, where the action network takes the state as input and outputs continuous actions, and the policy network evaluates the value of the actions and updates the network parameters.

[0015] Step 2.2: Define the reward function, specifically using the user sum rate as the immediate reward, and optimize the base station transmit beamforming matrix and the active RIS phase shift matrix by maximizing the cumulative reward.

[0016] Step 2.3: Store the historical interaction data, i.e., state, action, reward, and next state, in the experience replay buffer, and update the network parameters using the mini-batch random sampling method.

[0017] Step 2.4: Update the parameters of the target network using the soft update method to smooth the network training process.

[0018] Step 2.5: Adjust and optimize the hyperparameters through experiments.

[0019] Optionally, the execution process of Step 3 includes the following steps:

[0020] Step 3.1: Hardware architecture design of the active RIS-assisted MISO system model.

[0021] Step 3.2: Establish the power consumption model.

[0022] Step 3.3: Channel modeling.

[0023] Step 3.4: Establish the signal transmission model.

[0024] Optionally, the active RIS-assisted MISO system model in step 3.1 includes a base station, an active RIS, and multiple single-antenna users. The active RIS consists of multiple reflecting elements RE, and each reflecting element RE amplifies the signal and adjusts the phase through an active load.

[0025] Optionally, the execution process of step 3.2 is specifically to analyze the power consumption characteristics of the active RIS, including the signal amplification power and the phase shift adjustment power, and combine it with the base station transmission power to establish a total system power consumption model. The expression is as follows:

[0026]

[0027] where, P C represents the total system power consumption, P DC represents the DC bias power used by the amplifier in each reflecting element RE of the active RIS, and P SW represents the power consumed by the phase shift of each reflecting element RE of the RIS.

[0028] Optionally, in the process of channel modeling in step 3.3, considering the channel characteristics from the base station to the active RIS and from the user to the active RIS, including path loss and small-scale fading, a channel state information model is established:

[0029]

[0030] where, R 1 is the small-scale Rayleigh fading vector, and its elements follow an independent complex Gaussian distribution. PL 1 (x) is the path loss.

[0031] Optionally, the signal transmission model in step 3.4 is the mathematical model of the received signal at the user. The signal received at the user is

[0032]

[0033] where, a is the amplification factor, H 1 is the incident channel, H 2,k is the reflection channel, Φ is the phase shift matrix of the active RIS, n r is the thermal noise generated by the active RIS, n is the noise at the user reception, g k is the k-th column vector of the base station beamforming matrix G, and x is the data stream of the k-th user.

[0034] Optionally, the process of simulation verification in step 5 includes the following steps:

[0035] Step 5.1: Compare the performance of the active RIS and the passive RIS;

[0036] Step 5.2: Analyze the impact of different base station transmission powers on system performance;

[0037] Step 5.3: Study the impact of the number of active RIS reflection units on the sum rate;

[0038] Step 5.4: Optimize the power allocation of the active RIS;

[0039] Step 5.5: Explore the deployment location of the active RIS.

[0040] The present invention provides a joint optimization method for an active RIS-assisted MISO communication system based on DRL. Taking the base station transmission beamforming matrix and the active RIS phase shift matrix as optimization variables, a joint optimization task is established and integrated into a DRL framework. Then, a DRL algorithm is designed based on the deep deterministic policy gradient to achieve the joint optimization of the base station transmission beamforming matrix, the RIS phase shift matrix, power allocation, and location deployment through an intelligent algorithm. Specifically, the communication between the base station and the user is enhanced by the active RIS, where the active RIS optimizes the wireless channel by amplifying the incident signal and adjusting the phase; the system jointly optimizes the base station transmission beamforming matrix and the phase shift matrix of the active RIS, and maximizes the user sum rate through the optimization of power allocation and location selection design. Through simulation verification, the present invention breaks through the limitations of traditional optimization methods and significantly improves the user sum rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 It is a schematic structural diagram of the active RIS-assisted multi-user MISO system model of the present invention.

[0043] Figure 2 It is a schematic diagram of the principle of the DRL algorithm based on the deep deterministic policy gradient (DDPG) of the present invention.

[0044] Figure 3 It is a schematic diagram of the component composition of the DRL algorithm based on the deep deterministic policy gradient (DDPG) of the present invention.

[0045] Figure 4 It is a comparison schematic diagram between the active RIS-assisted system and the passive RIS-assisted system in the simulation experiment of the specific embodiment of the present invention.

[0046] Figure 5It is a schematic diagram showing the influence of different transmission powers on the system performance in the simulation experiment of the specific embodiment of the present invention.

[0047] Figure 6 It is a schematic diagram showing the influence of the number of reflection units on the sum rate in the simulation experiment of the specific embodiment of the present invention.

[0048] Figure 7 It is a schematic diagram showing the influence of the learning rate and decay rate of the network on the sum rate in the simulation experiment of the specific embodiment of the present invention.

[0049] Figure 8 It is a schematic diagram showing the design of the scheme for optimizing the active RIS power allocation in the simulation experiment of the specific embodiment of the present invention.

[0050] Figure 9 It is a schematic diagram of the system model for optimizing the selection of the active RIS position in the simulation experiment of the specific embodiment of the present invention.

[0051] Figure 10 It is a schematic diagram showing the influence of different power allocations on the sum rate in the simulation experiment of the specific embodiment of the present invention.

[0052] Figure 11 It is a comparison schematic diagram for verifying whether the optimized design related to power allocation is effective in the simulation experiment of the specific embodiment of the present invention.

[0053] Figure 12 It is a diagram showing the influence of the deployment position of the active RIS on the sum rate in the simulation experiment of the specific embodiment of the present invention.

[0054] Figure 13 It is a comparison schematic diagram for verifying whether the optimized design related to position deployment is effective in the simulation experiment of the specific embodiment of the present invention. Detailed implementation manners

[0055] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.

[0056] Explanation of the meanings of the commonly used English term abbreviations in the present invention:

[0057] RIS: Reconfigurable Intelligent Surface, reconfigurable intelligent surface;

[0058] DRL: Deep Reinforcement Learning, deep reinforcement learning;

[0059] MISO: Multiple Input Single Output, which means multiple inputs and single output;

[0060] DDPG: Deep Deterministic Policy Gradient, which is a deep deterministic policy gradient;

[0061] The present invention provides a joint optimization method for an active RIS-assisted MISO communication system based on DRL, including the following steps:

[0062] Step 1: Taking the base station transmission beamforming matrix and the active RIS phase shift matrix as optimization variables, establish a joint optimization task and transform the optimization problem into a DRL task;

[0063] Step 2: Design a DRL algorithm based on DDPG to solve the joint optimization problem of continuous beamforming matrix and phase shift matrix using the DDPG algorithm;

[0064] Step 3: Construct an active RIS-assisted MISO system model;

[0065] Step 4: Conduct optimization from two dimensions of power allocation and location optimization to obtain a DRL-based dynamic power allocation strategy and a selection strategy for the deployment location of the active RIS;

[0066] Step 5: Conduct simulation verification.

[0067] The following further elaborates in combination with specific embodiments and implementation steps:

[0068] Please refer to Figure 1 , Figure 1 which is the constructed active RIS-assisted multi-user MISO system model. It consists of a base station with M antennas, an active RIS with N reflecting elements, and K single-antenna users. As Figure 1 shown, the active RIS can dynamically adjust the reflection coefficients at each reflecting element (this is also known as reflection beamforming), and reconfigure the incident signal by changing the phase shift and amplification factor of the reflecting element. The base station simultaneously transmits K data streams through M antennas, with each data stream corresponding to a user. The signal first reaches the active RIS and then is reflected and amplified by the RIS. Generally, the base station can transmit information to the user through two links, namely the direct path and the RIS-assisted cascaded path. Since the RIS is usually used when the direct path is blocked, the direct path is regarded as completely blocked and not considered here.

[0069] The specific functions of each component are as follows:

[0070] (1) Base station: Configure M antennas and be responsible for sending data streams to K single-antenna users.

[0071] (2) Active RIS: Deploy N reflecting elements, each of which has the functions of signal amplification and phase shift adjustment.

[0072] (3) User: Receive the enhanced signal reflected by the RIS.

[0073] Furthermore, the signal received at the user is

[0074]

[0075] where a is the amplification factor, H 1 is the incident channel, H 2,k is the reflection channel, Φ is the phase shift matrix of the active RIS, n r is the thermal noise generated by the active RIS, n is the noise at the user receiver, g k is the k-th column vector of the base station beamforming matrix G, and x is the data stream of the k-th user.

[0076] The base station transmit power constraint is: Ε{tr{Gx(Gx) H}} ≤ P C , where P C is the maximum transmit power allowed at the base station.

[0077] The amplification factor constraint of the active RIS is

[0078]

[0079] After further processing, it is

[0080]

[0081] Furthermore, the signal-to-interference-plus-noise ratio (SINR) at the k-th user is

[0082]

[0083] The achievable rate of the users in the active RIS-assisted MISO system can be written as

[0084] R k = log 2 (1 + r k )

[0085] The total power consumption of the active RIS-assisted system can be expressed as P C represents the total power consumption of the system, P DC represents the DC bias power used by the amplifier in each RE of the active RIS, P SW represents the power consumed by the phase shift of each RE of the RIS. The formula for the power consumption of the active RIS system can be rewritten as

[0086] The objective function and constraints for maximizing the total rate can be written as

[0087]

[0088] s.t. tr{GG H} ≤ P BS

[0089]

[0090] P BS + P R ≤ C.

[0091] It can be seen that this is a non - convex trivial optimization problem, where the objective function is non - convex and the constraints are also non - convex. It is difficult to obtain the optimal solution using classical mathematical tools, especially for such large - scale networks. In the present invention, instead of directly solving this challenging optimization problem mathematically, a rate optimization problem is formulated based on an advanced DRL method to maximize the sum rate.

[0092] Algorithm design based on DRL in step 2. The proposed algorithm has two neural networks, namely the actor network and the critic network. The former takes the state as input and outputs continuous actions, while the latter takes both continuous actions and the state as input to judge the value of the actions. It is not easy to find the maximizing Q - value of the actions given the next non - convex optimization state. The actor network can eliminate this need by approximating the actions. The goal of the critic network is to minimize the gap between the Q - value and the target Q - value, usually using the mean squared error (MSE) as the loss function to update its parameters. During the training process of DDPG, the two networks are updated alternately. The actor network is updated through policy gradients to improve the value of the generated actions. The critic network is updated by comparing the current Q - value and the target Q - value to better evaluate the value of the actions. As Figure 2 shown in the schematic diagram of the algorithm design.

[0093] Furthermore, given the state s, action a, reward r, the Q - function follows the Bellman equation to calculate the Q - value, which can be expressed as

[0094]

[0095] The soft update of the network can be expressed as

[0096] θ Q′ ← λθ Q + (1 - λ)θ Q′

[0097] θ μ′ ← λθ μ+(1 - λ)θ μ′

[0098] where λ is the learning rate of the target network, usually a very small constant (such as 0.001), indicating how to adjust the parameter θ μ .

[0099] Furthermore, the algorithm design includes four main components, such as Figure 3 shown below:

[0100] (1) Central controller: Coordinates the scheme process and regulates the RIS power allocation.

[0101] (2) Environment module: Simulates the actual communication environment and generates new states according to actions.

[0102] (3) Agent module: Executes algorithm processing, randomly samples from the experience replay buffer to update parameters.

[0103] (4) Experience replay buffer: Stores interaction experiences (state, action, next state, reward).

[0104] Assume that the system can collect channel state information in real time. The central controller is responsible for correctly implementing the scheme process, including adjusting the power allocation of the active RIS. The environment module simulates the actual environment and generates different states through different actions. The main responsibility of the agent module is to execute actions and optimize strategies according to environmental feedback. The algorithm network structure includes an input layer, two hidden layers, and an output layer. The number of neurons in the hidden layers is determined by the number of base station antennas, the number of RIS elements, etc.

[0105] Specifically, the specific steps of the algorithm are as follows:

[0106] Input: Incident channel H 1 , reflected channel H 2,k , thermal noise n r , noise n.

[0107] Initialization: Environment (ENV), DDPG agent (AGE), experience replay buffer (EB), active RIS power consumption P R , G, Φ, S.

[0108] for episode number episode = 0, 1,..., MAX_episode

[0109] Step 1: Collect the initial observation state s (t) , t = 0.

[0110] for time step t = 0, 1,..., MAX_t

[0111] Step 2: The agent AGE determines the action a according to the current state s(t) Generate action a (t) .

[0112] Step 3: Input action a (t) into the environment ENV to obtain the next state s (t+1) .

[0113] Step 4: Obtain the immediate reward r (t+1) .

[0114] Step 5: Store the experience (s (t) , a (t) , s (t+1) , r (t+1) ) into the experience replay buffer EB.

[0115] Step 6: Randomly sample multiple batches of data from the buffer.

[0116] Step 7: Update the network parameters.

[0117] Step 8: Update the current state s (t) = s (t+1)

[0118] end for

[0119] end for

[0120] Output: The optimal beamforming matrix G, phase shift matrix Φ, and user sum rate S.

[0121] The simulation parameters are shown in the following table:

[0122]

[0123]

[0124] The optimization design for maximizing the sum rate. In the application and deployment of active RIS, there are actually more considerations. Generally, it is necessary to maximize the system performance under the premise of controlling costs. Based on these considerations, some new problems can be obtained, such as the cost of deployment and the cost performance of power consumption, the problem of where to deploy RIS to achieve the best effect, etc. Maximize the sum rate by optimizing the power allocation and location deployment of active RIS, which is divided into two sub-schemes for separate optimization.

[0125] First is the power allocation design of active RIS. The system architecture mainly consists of four parts:

[0126] Central controller: Integrates the power allocation function and coordinates the power allocation strategies of the base station and active RIS.

[0127] Environment module: Simulates the actual communication environment and feedbacks channel state, reward, and termination conditions.

[0128] Agent module: Based on the Actor-Critic framework of DDPG, generates actions (power allocation ratios).

[0129] Experience buffer: Stores historical interaction data (state, action, reward), supports mini-batch random sampling to improve algorithm stability.

[0130] Through the DRL framework, the present invention transforms the power allocation problem of the active RIS into a dynamic optimization task, and combines techniques such as experience replay and batch normalization to achieve the best power allocation exploration in complex scenarios. The power allocation ratio of the active RIS is divided into 101 cases from 0% to 100%, and the maximum total rate under each ratio is calculated in turn, and each ratio is trained for 10,000 iterations. The size of the user sum rate reflects the system performance. During the design process, due to the fixed total power, as the power allocated to the active RIS increases, the base station transmission power will gradually decrease. The convergence speed of the total rate under different base station transmission powers is different, which may cause the average total rate to not reflect the best performance, thereby affecting the judgment of the final result. Therefore, the maximum sum rate is selected as the final reference index. Since the algorithm based on DRL is adopted, data fluctuations are inevitable. To facilitate the study of data trends, the data is smoothed, specifically using the moving average method, and its formula is as follows where x[] is the original data, y[] is the processed data, w is the size of the window, and n represents the serial number of the last data. And to ensure the consistency and correspondence of the data length, forward filling is performed at the beginning of the sequence.

[0131] Then there is the design of the active RIS position selection. To simplify the research, the applicant regards the base station, RIS, and user area as a point, and temporarily does not consider the influence of their own sizes. In this case, the initial position of the active RIS is set to (0,0), the base station position is (0,10), and the position of the user area is (100, 10).

[0132] Furthermore, path loss modeling is carried out. The path loss decays with the increase of distance, and the free space path loss model is adopted: where is the carrier wavelength and d is the transmission distance.

[0133] Furthermore, the transmission distance d from the base station to the RIS 1 can be expressed as The transmission distance from the RIS to the user area can be expressed as

[0134] Furthermore, the path loss from the base station to the RIS is The path loss from the RIS to the user is

[0135] Furthermore, the channel from the base station to the RIS can be expressed as The channel from the RIS to the user area can be expressed as where the R 1 and the R 2,k are both small-scale Rayleigh fading vectors, and their elements follow independent complex Gaussian distributions.

[0136] Furthermore, the received signal of user k can be expressed as

[0137]

[0138] Furthermore, the path loss is further expressed as

[0139] PL(d) = 20log 10 (d) + 20log 10 (f) - 147.55.

[0140] where the unit of d is meters and the unit of f is Hertz. d is the distance between the deployment location of the active RIS and the base station, and f is the frequency.

[0141] For further illustration, please refer to Figures 4 to 13 The present invention conducts simulation experiments for verification and comparison:

[0142] Combined with the latest achievements in wireless communication in recent years, millimeter waves are selected for the experiment. The frequency of millimeter waves is generally between 30 GHz and 300 GHz, and millimeter waves of 30 GHz are selected and substituted into f.

[0143] After selecting millimeter waves, it is also necessary to consider its effective propagation range. Even under the visible condition without obstacles, the maximum effective propagation distance of 30 GHz millimeter waves may only be 100 to 300 meters in different scenarios. This is also the reason for setting the deployment range of the active RIS between 0 and 100 meters, so that the total propagation distance is approximately several hundred meters.

[0144] Similarly, some modifications are made to the previously proposed algorithm. It is selected that x increases by 1 each time from 0 until it reaches 100, and each position selection of x is trained for 10,000 time steps. In data processing, 10,000 immediate rewards are processed to obtain the average reward.

[0145] Figure 4 This is the comparison between the active RIS-assisted system and the passive RIS-assisted system in the simulation experiment. This figure compares the sum rates of the active RIS-assisted system and the passive RIS-assisted system. The system parameters mentioned above are used, and the base station transmit power is P t= 20 dB. On the premise that the reference power is 1 W, this power is 100 W. With other conditions being the same, the active RIS consumes an additional 20 W of power to amplify the reflected signal.

[0146] Figure 5 This is the impact of different transmission powers on the system performance. The impact of different base station transmission powers on the system was studied. The same color system is used to represent the immediate reward and average reward under the same transmission power. It can be seen that the performance of the active RIS-assisted communication system is similar to that of the passive RIS: as the time step t increases, the reward value gradually converges; and the convergence speed is faster at low signal-to-noise ratios (= 10 dB), while slower at high signal-to-noise ratios (= 20 dB and = 40 dB). This is because the higher the signal-to-noise ratio, the larger the dynamic range of the immediate reward, resulting in greater fluctuations, so the effect is worse.

[0147] Figure 6 This is the impact of the number of reflection units on the sum rate. The system configuration considering the active RIS is the number of reflection elements N = {4, 10, 40}, and other configurations remain unchanged. To verify the advantages of the active RIS, a passive RIS is added for comparison. The configuration of the passive RIS-assisted system is (M = 4, N = 4, K = 4). The abscissa in the figure is the time step, and the ordinate is the average reward. Compared with the change in transmission power, the proposed algorithm is more robust to the change in the number of reflection units, that is, the change in the system configuration will not significantly affect the convergence time of the reward due to the increase in the number of reflection elements.

[0148] Figure 7 This is to study the performance of the proposed algorithm and the impact of the learning rate and decay rate of the network on the sum rate. The impact of different learning rates on the average reward was compared. In the experimental scenario, compared with 0.0001 and 0.00001, the network with a learning rate of 0.001 has better algorithm performance but a longer convergence time; while when the learning rates are 0.1 and 0.01, although the learning rates are larger, the oscillation of the reward value is more obvious, resulting in sub-optimal final performance. Its performance is similar to that of the learning rate: the optimal decay rate is neither the largest nor the smallest. In this experiment, better performance was achieved when the decay rate was 0.00001.

[0149] Furthermore, Figure 8 This is the design of a scheme to optimize the power allocation of the active RIS. The central controller dynamically allocates the power ratio between the base station and the active RIS to achieve an appropriate balance between the signal amplification gain and power consumption to maximize the system sum rate.

[0150] Figure 9It is a system model for optimizing the position selection of an active RIS. It consists of a base station with M antennas, an active RIS with N reflecting elements, and K single-antenna users in a region. On this basis, a horizontally deployable range that can be arbitrarily arranged is added to the active RIS. The system has different user sum-rates at different horizontal positions. The research objective is to find the optimal deployment position to maximize the sum-rate.

[0151] Figure 10 It is the influence of different power allocations on the sum-rate. Keeping the total system power constant at 100W, on this premise, the influence of different power allocation ratios of the active RIS on the system sum-rate is explored.

[0152] Figure 11 It is the comparison between the selected optimal result and the traditional equal-division scheme. This simulation result proves that evenly dividing the power between the base station and the active RIS (i.e., each accounting for 50%) to compare the differences between active and passive RIS systems is not the optimal allocation scheme. Through experiments, it is verified that the method of the present invention can find a more reasonable power allocation ratio.

[0153] Figure 12 It is the influence of the deployment position of the active RIS on the sum-rate. It can be seen from the figure that the active RIS has better performance when deployed in the middle position between the base station and the user group.

[0154] Figure 13 It is the comparison of the sum-rates at three key points. This simulation result shows that both the immediate reward and the average reward when x = 50 are much greater than the cases when x = 0 and x = 100, which proves that the design of the present invention for finding the best deployment can achieve the goal.

[0155] In summary, the present invention realizes the joint design of the base station transmission beamforming matrix, the RIS phase shift matrix, power allocation, and position deployment through an intelligent algorithm, breaks through the limitations of traditional optimization methods, and significantly improves the user sum-rate.

[0156] What is disclosed above is only one or more preferred embodiments of the present invention. Of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. A DRL-based active RIS-assisted MISO communication system joint optimization method, characterized in that: The following steps are involved: Step 1: Taking the base station transmit beamforming matrix and the active RIS phase shift matrix as optimization variables, a joint optimization task is established and the optimization problem is transformed into a DRL task; Step 2: Design the DRL algorithm based on DDPG, and use the DDPG algorithm to solve the joint optimization problem of the continuous beamforming matrix and the phase shift matrix; Step 3: Build an active RIS-assisted MISO system model; Step 4: Optimize from two dimensions: power allocation and location optimization to obtain a dynamic power allocation strategy based on DRL and a selection strategy for the deployment location of the active RIS. Step 5: Simulation verification.

2. The DRL-based active RIS-assisted MISO communication system joint optimization method according to claim 1, characterized in that: In step 1, the joint optimization problem of the base station transmit beamforming matrix and the active RIS phase shift matrix is ​​modeled as a non-convex optimization problem with the goal of maximizing the system sum rate; The process of converting the optimization problem into a DRL task is to convert the joint optimization problem into a dynamic optimization task under the DRL framework by defining the state space, action space and reward function, where the state includes the channel state information and the current system configuration, the action is the base station transmit beamforming matrix and the active RIS phase shift matrix, and the reward is the user and rate.

3. The DRL-based active RIS-assisted MISO communication system joint optimization method according to claim 2, characterized in that: Step 2 The process of designing a DRL algorithm based on DDPG includes the following steps: Step 2.1: Build the DDPG network structure, including designing the action network and the policy network, where the action network takes the state as input and outputs continuous actions, and the policy network evaluates the value of the action and updates the network parameters; Step 2.2: Define the reward function, specifically taking the user and rate as the immediate reward, and optimize the base station transmit beamforming matrix and active RIS phase shift matrix by maximizing the cumulative reward; Step 2.3: Store historical interaction data, i.e., state, action, reward, and next state, through the experience replay buffer, and use a small batch random sampling method to update network parameters; Step 2.4: Use the soft update method to update the parameters of the target network and smooth the network training process; Step 2.5: Optimize hyperparameters through experimental adjustment.

4. The DRL-based active RIS-assisted MISO communication system joint optimization method according to claim 3, characterized in that: The execution process of step 3 includes the following steps: Step 3.1: Hardware architecture design of active RIS-assisted MISO system model; Step 3.2: Establish power consumption model; Step 3.3: Channel modeling; Step 3.4: Establish a signal transmission model.

5. The DRL-based active RIS-assisted MISO communication system joint optimization method according to claim 4, characterized in that: The active RIS-assisted MISO system model constructed in step 3.1 includes a base station, an active RIS, and multiple single-antenna users, wherein the active RIS is composed of multiple reflective elements RE, each of which amplifies the signal and adjusts the phase through an active load.

6. The DRL-based active RIS-assisted MISO communication system joint optimization method according to claim 5, characterized in that: The execution process of step 3.2 is to analyze the power consumption characteristics of active RIS, including signal amplification power and phase shift adjustment power, and combine them with the base station transmission power to establish the system total power consumption model, which is expressed as follows: Among them, P C Indicates the total power consumption of the system, P DC P represents the DC bias power used by the amplifier in each reflective element RE of the active RIS. SW It represents the power consumed by phase shifting of each reflective element RE of RIS.

7. The DRL-based active RIS-assisted MISO communication system joint optimization method according to claim 6, characterized in that: In the process of channel modeling in step 3.3, the channel characteristics from the base station to the active RIS and from the user to the active RIS, including path loss and small-scale fading, are considered to establish a channel state information model: Where R1 is the small-scale Rayleigh fading vector, whose elements obey independent complex Gaussian distributions, and PL1(x) is the path loss.

8. The DRL-based active RIS-assisted MISO communication system joint optimization method according to claim 7, characterized in that: Step 3.4 The signal transmission model is a mathematical model of the user receiving the signal. The signal received at the user is Where a is the amplification factor, H1 is the incident channel, and H 2,k is the reflection channel, Φ is the phase shift matrix of the active RIS, n r is the thermal noise generated by the active RIS, n is the noise at the user’s receiving location, g k is the k-th column vector of the base station beamforming matrix G, and x is the data stream of the k-th user.

9. The DRL-based active RIS-assisted MISO communication system joint optimization method according to claim 8, characterized in that: The simulation verification process in step 5 includes the following steps: Step 5.1: Compare the performance of active RIS and passive RIS; Step 5.2: Analyze the impact of different base station transmit powers on system performance; Step 5.3: Study the effect of the number of active RIS reflection units on the sum rate; Step 5.4: Optimize power distribution of active RIS; Step 5.5: Explore deployment locations for active RIS.

Citation Information

Cited By

  • Internet of Things security resource optimization method and system based on multi-agent reinforcement learning

    CN120916144A

  • A Method and System for Optimizing IoT Security Resources Based on Multi-Agent Reinforcement Learning

    CN120916144B