STAR-RIS assisted cellular-free mMIMO resource allocation method

By using the DDPG algorithm to optimize the location and parameters of STAR-RIS in the STAR-RIS auxiliary cellular-free large-scale MIMO system, the problems of signal occlusion and multipath reflection in traditional systems are solved, communication quality and network coverage are improved, energy consumption is reduced, and efficient resource allocation is achieved.

CN120583533APending Publication Date: 2025-09-02BEIJING INFORMATION SCI & TECH UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510845515.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Traditional cellular-free large-scale MIMO communication systems face signal occlusion and multipath reflection problems in complex environments, resulting in an increase in communication dead zones, a decrease in service quality, and an increase in network energy consumption. The existing optimization methods are inefficient.

Method used

The STAR-RIS-assisted cellular-free large-scale MIMO system is adopted, combined with the deep deterministic strategy gradient (DDPG) algorithm, and the position, reflection and transmission beamforming matrix, phase shift matrix and amplitude coefficient of STAR-RIS are optimized to improve resource allocation efficiency through deep reinforcement learning.

Benefits of technology

It significantly improves communication quality and network coverage in complex environments, reduces network energy consumption, and improves information transmission rate and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120583533A_ABST
    Figure CN120583533A_ABST
Patent Text Reader

Abstract

The invention discloses a resource allocation method of STAR-RIS assisted cellular-free mMIMO (multiple input multiple output). The method comprises the following steps: acquiring position information of UE, AP and STAR-RIS, acquiring channel state information (CSI) of each communication link, and constructing a state information set of a current system based on the position information and the CSI; based on the state information set, utilizing an Actor network in a depth deterministic policy gradient (DDPG) algorithm to generate an optimization action; adjusting resources among the UE, the AP and the STAR-RIS on the basis of the optimization action, performing downlink signal transmission on the basis of the adjusted resources, and calculating the outage probability, the SINR and the information transmission rate per unit time of the UE; and outputting an instant reward value under the optimization action based on the outage probability, the SINR and the information transmission rate, and performing updating optimization based on the instant reward value. According to the invention, the technical problem of poor performance of a traditional cellular-free large-scale MIMO communication system is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technology, and in particular to a STAR-RIS assisted cellular-free MMIMO resource allocation method. Background Art

[0002] With the rapid development of emerging applications such as virtual reality, augmented reality, autonomous driving, and the Internet of Things, communication device types are becoming increasingly diverse, and users are placing higher demands on the data transmission rate and latency performance of wireless communication systems. However, traditional centralized cellular networks often face problems such as low spectrum utilization, severe signal interference, and prominent cell edge effects when dealing with high-density and complex communication environments, making it difficult to meet current and future network performance requirements.

[0003] Cell-free massive multiple-input, multiple-output (CF-mMIMO) technology effectively improves spectrum utilization, reduces inter-cell interference, and eliminates the cell boundary effect in traditional cellular networks by deploying a large number of distributed access points over a wide area, providing users with more flexible and continuous communication services. However, in densely populated urban environments, signal transmission is susceptible to obstruction and multipath reflections, leading to an increase in communication dead zones, reduced service quality, and increased network energy consumption. These issues urgently need to be addressed.

[0004] To address this issue, the "STAR-RIS" (Reconfigurable Smart Surface with Both Transmitting and Reflecting Capabilities) technology has emerged. Unlike traditional passive smart surfaces that only reflect, STAR-RIS can simultaneously beamform both reflected and transmitted waves, thereby creating an intelligent wireless propagation environment with full spatial coverage. This technology provides more degrees of freedom for signal propagation, enhancing communication link quality while expanding network coverage, making it particularly suitable for complex scenarios such as urban areas.

[0005] Combining STAR-RIS with CF-mMIMO technology to build a STAR-RIS-assisted cell-free massive multiple-input multiple-output system (STAR-RIS-CF-mMIMO) can further improve the overall performance and resource utilization efficiency of the network, demonstrate stronger environmental adaptability and communication robustness in complex environments, and provide users with a more stable and high-quality service experience.

[0006] Currently, performance optimization in such systems mostly relies on traditional iterative optimization methods and alternating optimization algorithms, resulting in poor performance of non-cellular massive MIMO communication systems.

[0007] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0008] The embodiment of the present invention provides a STAR-RIS assisted non-cellular massive MIMO resource allocation method to at least solve the technical problem of poor performance of traditional non-cellular massive MIMO communication systems.

[0009] According to one aspect of an embodiment of the present invention, a STAR-RIS-assisted resource allocation method for non-cellular massive MIMO is provided, comprising: collecting location information of user equipment UE, access point AP, and simultaneously transmitting and reflecting reconfigurable smart surface STAR-RIS in the current system, and obtaining channel state information CSI of each communication link in the current system, constructing a state information set of the current system based on the location information and the CSI; based on the state information set, generating an optimization action using an Actor network in a deep deterministic policy gradient DDPG algorithm, wherein the optimization action comprises at least one of the following: transmitting beamforming matrix, receiving ... The method comprises the following steps: calculating a shaping matrix, a position coordinate of a STAR-RIS, a reflection amplitude coefficient of a STAR-RIS, and a phase shift matrix of a STAR-RIS; adjusting resources between the UE, the AP, and the STAR-RIS based on the optimization action, performing downlink signal transmission based on the adjusted resources, and calculating an interruption probability, a signal-to-noise ratio (SINR), and an information transmission rate per unit time of the UE; outputting an instantaneous reward value under the optimization action based on the interruption probability, the signal-to-noise ratio (SINR), and the information transmission rate; and optimizing weight parameters of the Actor network and the Critic network in the DDPG algorithm through gradient updating based on the instantaneous reward value.

[0010] According to another aspect of an embodiment of the present invention, a STAR-RIS-assisted non-cellular massive MIMO resource allocation device is also provided, comprising: an acquisition module configured to acquire location information of user equipment UE, access point AP and simultaneously transmitting and reflecting reconfigurable smart surface STAR-RIS in the current system, and obtain channel state information CSI of each communication link in the current system, and construct a state information set of the current system based on the location information and the CSI; an optimization module configured to generate an optimization action based on the state information set using the Actor network in the deep deterministic policy gradient DDPG algorithm, wherein the optimization action includes at least one of the following: sending a beamforming matrix, A receiving beamforming matrix, a position coordinate of STAR-RIS, a reflection amplitude coefficient of STAR-RIS, and a phase shift matrix of STAR-RIS are received; an adjustment module is configured to adjust the resources between the UE, the AP, and the STAR-RIS based on the optimization action, and perform downlink signal transmission based on the adjusted resources, and calculate the interruption probability, signal-to-noise ratio (SINR), and information transmission rate per unit time of the UE; an update module is configured to output an instantaneous reward value under the optimization action based on the signal-to-noise ratio (SINR) and the information transmission rate, and optimize the weight parameters of the Actor network and the Critic network in the DDPG algorithm through gradient update based on the instantaneous reward value.

[0011] In an embodiment of the present invention, location information of a user equipment (UE), an access point (AP), and a simultaneously transmitting and reflecting reconfigurable smart surface (STAR-RIS) in a current system is collected, and channel state information (CSI) of each communication link in the current system is obtained. A state information set of the current system is constructed based on the location information and the CSI. Based on the state information set, an optimization action is generated using an actor network in a deep deterministic policy gradient (DDPG) algorithm, wherein the optimization action includes at least one of the following: a transmit beamforming matrix, a receive beamforming matrix, the location coordinates of the STAR-RIS, a reflection amplitude coefficient of the STAR-RIS, and a phase shift matrix of the STAR-RIS. Based on the optimization action, resources between the UE, the AP, and the STAR-RIS are adjusted, and downlink signal transmission is performed based on the adjusted resources. The outage probability, signal-to-noise ratio (SINR), and information transmission rate per unit time of the UE are calculated. Based on the outage probability, the signal-to-noise ratio (SINR), and the information transmission rate, an instantaneous reward value for the optimization action is output, and based on the instantaneous reward value, weight parameters of the actor network and the critic network in the DDPG algorithm are optimized through gradient updating. The above solution solves the technical problem of poor performance of traditional non-cellular massive MIMO communication systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0013] Figure 1 is a flowchart of an optional STAR-RIS assisted non-cellular massive MIMO resource allocation method according to an embodiment of the present invention;

[0014] Figure 2 1 is a schematic structural diagram of an optional user-heterogeneous STAR-RIS-CF-mMIMO system according to an embodiment of the present invention;

[0015] Figure 3 1 is a distribution diagram of UEs, APs, and STAR-RIS after algorithm optimization according to an embodiment of the present invention;

[0016] Figure 4 is a curve diagram showing the impact of different numbers of UEs on information transmission rate according to an embodiment of the present invention;

[0017] Figure 5 is a graph showing the effect of different numbers of APs on information transmission rate according to an embodiment of the present invention;

[0018] Figure 6 is a graph showing the effect of different numbers of RIS on information transmission rate according to an embodiment of the present invention;

[0019] Figure 7 1 is a schematic structural diagram of an optional STAR-RIS-assisted non-cellular massive MIMO resource allocation system according to an embodiment of the present invention;

[0020] Figure 8 A schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0022] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0023] According to an embodiment of the present invention, a method embodiment of a STAR-RIS-assisted cellular-free massive MIMO resource allocation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0024] Existing resource allocation technologies often face signal obstruction and blind spots in complex environments (such as densely populated urban areas), resulting in degraded communication quality. Furthermore, the mathematical modeling and solution requirements of traditional resource allocation methods in such scenarios are more complex, placing higher demands on algorithm efficiency and adaptability. To address the communication degradation caused by signal obstruction, traditional passive reconfigurable smart surfaces (RIS) are limited by their one-sided reflective properties, making it difficult to achieve full coverage in any environment.

[0025] To this end, the present invention proposes a STAR-RIS-assisted resource allocation method for cell-free massive MIMO. By combining deep reinforcement learning with resource optimization strategies, the information transmission rate of the cell-free massive multiple-input multiple-output (STAR-RIS-CF-mMIMO) communication system enabled by a simultaneous transmissive and reflective reconfigurable smart surface is improved. The method uses the deep deterministic policy gradient (DDPG) algorithm to jointly optimize the position coordinates, phase shift matrix, reflection / transmission amplitude coefficients, and transmit and receive beamforming matrices of multiple STAR-RIS. Based on the current location distribution of user equipment (UE) and access point (AP), the present invention can solve the optimal multi-STAR-RIS deployment location and its configuration parameters.

[0026] Figure 1 The STAR-RIS assisted non-cellular massive MIMO resource allocation method according to an embodiment of the present invention is as follows: Figure 1 As shown, the method includes the following steps:

[0027] Step S102: Collect the location information of user equipment UE, access point AP and simultaneously transmitting and reflecting reconfigurable smart surface STAR-RIS in the current system, obtain the channel state information CSI of each communication link in the current system, and construct a state information set of the current system based on the location information and the CSI.

[0028] For example, the location information of the UE, the AP, and the STAR-RIS is collected and normalized to a predefined coordinate system to obtain normalized location information; the channel state information CSI of the communication links between the UE and the AP, between the AP and the STAR-RIS, and between the STAR-RIS and the UE is obtained, wherein the CSI includes: path loss, Rayleigh fading parameters, and LOS / NLOS flags; the location information and the CSI are mapped into a state vector to construct the state information set for input to the DDPG algorithm.

[0029] Step S104: Based on the state information set, an optimization action is generated using the Actor network in the deep deterministic policy gradient (DDPG) algorithm, wherein the optimization action includes at least one of the following: a transmit beamforming matrix, a receive beamforming matrix, a STAR-RIS position coordinate, a STAR-RIS reflection amplitude coefficient, and a STAR-RIS phase shift matrix.

[0030] For example, the state vector in the state information set is input into an Actor network constructed by a deep neural network; the phase shift matrix, reflection amplitude coefficient and position coordinates for controlling the STAR-RIS, as well as the optimization parameters for the transmit and receive beamforming matrices of the AP and UE are generated at the output layer of the Actor network; based on the optimization parameters, the optimization action is generated, and regularization and restriction functions are applied to the optimization action so that the optimization action satisfies power constraints, position constraints, amplitude physical constraints, and phase physical constraints.

[0031] Step S106: Based on the optimization action, adjust the resources between the UE, the AP and the STAR-RIS, perform downlink signal transmission based on the adjusted resources, and calculate the UE's interruption probability, signal-to-noise ratio (SINR), and information transmission rate per unit time.

[0032] The transmit beamforming matrix output by the Actor network is assigned to the corresponding AP to determine its downlink transmission signal structure for each UE; the receive beamforming matrix is ​​assigned to the corresponding UE to determine its weighted combining strategy for the received signal; the reflection and transmission paths are updated according to the position and parameter configuration of the STAR-RIS, and a UE aggregation channel matrix containing direct links and indirect links is constructed.

[0033] Step S108: Based on the interruption probability, the signal-to-noise ratio (SINR), and the information transmission rate, output an immediate reward value under the optimization action, and based on the immediate reward value, optimize the weight parameters of the Actor network and the Critic network in the DDPG algorithm through gradient update.

[0034] The transmission rates of all UEs are summed to form the initial reward value for this time slot. Penalties are introduced to adjust the initial reward value for violations of the lower limit of the signal-to-noise ratio (SINR), the outage probability threshold, or the STAR-RIS deployment restriction, to form the final instantaneous reward value. The penalty terms include at least one of the following: a proportional penalty term is applied to reduce the total reward when the SINR of any UE is below a set threshold; a constraint penalty term is applied when the outage probability of any UE is above a set outage probability threshold; and an area constraint penalty term is applied when the location of any STAR-RIS exceeds a predefined deployment boundary.

[0035] Because the phase matrices involved in STAR-RIS are continuous variables, this embodiment employs deep reinforcement learning algorithms, specifically the Deep Deterministic Policy Gradient (DDPG) algorithm, which demonstrates excellent performance in high-dimensional continuous control scenarios. This algorithm dynamically optimizes STAR-RIS's configuration by adjusting system control policies in real time to adapt to changing network conditions, significantly improving resource allocation efficiency, network energy consumption control, and overall communication quality.

[0036] The present application also provides another STAR-RIS assisted non-cellular massive MIMO resource allocation method, which is applied in the following cases: Figure 2 In the downlink STAR-RIS-CF-mMIMO communication system shown in FIG, multiple APs and STAR-RIS collaborate to provide services for heterogeneous users. Figure 2 As shown in the figure, the STAR-RIS-CF-mMIMO system includes M APs, each of which is equipped with N antennas. L STAR-RIS (Simultaneously Transmitting and Reflecting Reconfigurable Intelligence Surface) are deployed around the AP. Each STAR-RIS consists of P 2elements, where the order of the phase shift matrix is ​​P. This paper assumes that each element on STAR-RIS can realize both signal reflection and refraction transmission. Depending on whether the UE and the APs serving it are on the same side of STAR-RIS or on the opposite side, the signal transmitted from each AP is reflected or refracted by STAR-RIS to reach the UE. STAR-RIS optimizes the transmission path of wireless signals by adjusting the amplitude coefficient of its surface. The K UEs in the system are equipped with different numbers of antennas, where the kth UE is equipped with N k antennas, k∈{1,2,...,K}. After the distance rule is used, the m APs closest to UEk are selected. These APs serve UEk as multi-agents, 1≤m≤M. In addition, the total spatial resources of these multi-agent APs are greater than the total spatial resources of the UE, that is, MN>∑ k N k , thereby supporting more efficient resource management and optimization.

[0037] UE k receives direct signals from m APs and signals processed by L STAR-RISs. Assume that all channels are statistically independent and follow a Rice distribution. l∈{1,2,...,L} represents the lth STAR-RIS serving UE k. During signal transmission, the channels passing through STAR-RIS l are affected by the location coordinates of STAR-RIS l. The impact of This paper defines the deployable range Ω R , which is a two-dimensional spatial region used to limit the possible deployment locations of STAR-RIS. R By jointly optimizing the coordinate set of these L STAR-RIS in the environment Achieve the optimal joint deployment strategy for STAR-RIS.

[0038] i∈{1,2,...,m} represents the i-th AP among the m APs serving UE k. The channel from AP i to UE k is expressed as The channel from AP i to STAR-RIS l is expressed as The channel from STAR-RIS l to UE k is expressed as The aggregated channel of UE k is expressed as It consists of a direct link and a reflection / refraction link via STAR-RIS:

[0039]

[0040] In formula (1), represents the set of all direct links between UE k and the m APs divided according to the rule. is the direct channel from AP i to UE k.

[0041] Defined as the set of reflection / refraction links of m APs and L STAR-RIS participating in serving UE k, define is the channel set from m APs participating in serving UE k to STAR-RIS l, which is composed of is composed of m times of splicing, so It can be simplified as:

[0042]

[0043] In formula (2), is the phase shift matrix of STAR-RIS l, which can be expressed as a diagonal matrix:

[0044]

[0045] In formula (3), d ikl ∈{0, 1} is used to represent the relative positions of API, UE k and STAR-RIS l. When API and UE k are on the same side of STAR-RIS l, d ikl =1; when AP i and UE k are on both sides of STAR-RIS l, d ikl =0. β R ,β T ∈[0,1] represents the amplitude coefficients of reflection and refraction of STAR-RIS l, which are used to represent the attenuation or gain of the signal during reflection and refraction. In this paper, we assume that β R +β T =1, which means that all signal energy is fully utilized and there is no energy loss. lp ∈[0,2π) is the phase offset of STAR-RIS l, represents the phase shift of the pth element on the main diagonal of STAR-RIS l, p∈{1,2,…,P}.

[0046] The application of beamforming matrices introduces additional degrees of freedom, enabling better adaptation to different device configurations and numbers of antennas. Represents the AP's transmit beamforming matrix, each sub-matrix V represents the transmit beamforming matrix of the downlink transmission link from API to UE k. k The definition is as follows:

[0047]

[0048] Where Qk is the sequence length of the unit power signal received by UEk. k,q represents the transmit power of the qth stream allocated to UEk, where q = 1, 2, ..., Q k .p k,q It is given by:

[0049]

[0050] ||·|| of all transmit beamforming matrices F (Frobenius norm squared) is equal to the total transmit power P of all APs T Here, represents the measure of the sum of the squares of all elements in the matrix, which physically reflects the total energy of the signal in the transmit beamforming matrix, that is, the total transmit power. This can be expressed as:

[0051]

[0052] At the same time, define U represents the receive beamforming matrix of the heterogeneous UE at the receiving end, which is used to optimize signal reception at UE k. k The definition is as follows:

[0053]

[0054] Where u k,n is the nth diagonal element in the receive beamforming matrix, n=1,2,...,N k . It is expressed as α k,n is the amplitude gain of UEk on the nth receiving antenna, which represents the strength of the received signal on this antenna, α k,n >0, and θ k,n is the phase offset, which represents the phase of the received signal at the antenna, θ k,n ∈[0,2π). The transmit and receive beamforming matrices work together on the downlink to achieve efficient signal transmission and reception.

[0055] For UE k, it receives the signal It can be expressed as:

[0056]

[0057] In formula (8), It is expressed as the length of UE k is Q k The reflection and refraction of STAR-RIS will generate noise. is the noise introduced by STAR-RIS l, where This paper assumes that n A Follows independent circularly symmetric complex Gaussian distributions, each component has variance σ 2 ,Right now Indicates that the shape is P×Q k The identity matrix of . is the additive white Gaussian noise (AWGN) vector at UE k, that is

[0058] The signal-to-noise ratio of UE k is defined as:

[0059]

[0060] In formula (9), It is expressed as the sum of the interference signal powers from other UEs. is the power of the noise introduced by STAR-RIS, which can also be expressed as N k Q k ·σ 2 is the power of the ambient noise.

[0061] Information transmission rate R of UE k k Defined as:

[0062]

[0063] In formula (10), B is the bandwidth allocated to each UE.

[0064] In order to further optimize the resource allocation of these UEs and ensure that they can meet higher transmission rate requirements, this paper gives the interruption probability constraint under the target rate, as shown in (11h). Where μ is the channel capacity, P out (R k ) represents UE k at its information transmission rate R k The interruption probability under UE k has different maximum interruption probability thresholds P out,max,k .P out,max,k It can be defined as P out,max,k =N k ×P out,max , P out,max is the benchmark maximum outage probability threshold. In (11h), we set the SINR requirement of UEk to Γ k Defined as Γ k =N k ×Γ k,min , where Γ k,minis the baseline SINR requirement. To achieve high-speed data transmission, high-antenna UEs require a higher SINR to ensure data accuracy and stability. Low-antenna UEs can tolerate a certain degree of interruption probability, so we set a higher maximum interruption probability threshold for them. This helps prioritize basic communication services when resources are limited.

[0065] In summary, in the STAR-RIS-CF-mMIMO system, the present invention jointly optimizes the STAR-RIS position coordinates, phase shift matrix and its reflection amplitude coefficient, transmit and receive beamforming matrix, and maximizes the sum of the information transmission rates of all UEs in the system while meeting the differentiated performance requirements of heterogeneous users. The optimization model is:

[0066]

[0067] SINR k ≥Γ k (11c)

[0068]

[0069] β R +β T =1(11f)

[0070]

[0071] Constraint (11b) is V ik The Frobenius square norm of the i-th AP cannot exceed the total power budget P i,max , where K i represents the set of UEs served by the i-th AP. Constraint (11c) represents the SINR k The SINR requirement Γ of UE k must be met k Constraint (11d) states that the Frobenius square norm is equal to the length of the signal transmitted at unit power. Constraint (11e) is a constraint on the deployment location of STAR-RIS. Constraint (11f) is a constraint on the amplitude coefficient of STAR-RIS. Constraint (11g) ensures that each element in STAR-RIS does not change its amplitude when adjusting the phase of the incident signal. Constraint (11h) defines the outage probability constraint at a given target rate.

[0072] The parameters of the optimization problem above involve adjustments to the antenna phase and amplitude, which are continuous variables. The DDPG algorithm dynamically adjusts these continuous variables by combining policy gradients with value function evaluation to optimize signal transmission paths and adapt to different device configurations and antenna numbers.

[0073] This paper proposes a method called DSR-JOS, which is modeled as a Markov decision process (MDP) by defining an agent, state, action, and reward:

[0074] Agents: APs, acting as multi-agents, interact with the environment to collect relevant information and determine strategies for STAR-RIS and UEs. Based on this information, APs collaborate with other APs to optimize resource allocation and improve communication performance.

[0075] Action: The optimization variable in this paper is the position coordinate ω of STAR-RIS R , phase shift matrix Φ and its reflection amplitude coefficient β R , sending and receiving beamforming matrices V and U. The action of the i-th agent at time slot t is defined as:

[0076]

[0077] They represent the corresponding parameters of AP i at time slot t. Therefore, the multi-agent action a at time slot t t Defined as:

[0078] State: At the beginning of time slot t, the multi-agent will collect the location information ω of all APs, UEs and STAR-RIS, the CSIH of all channels, the action a of the previous time slot t-1 and the reward r of the previous time slot t-1 The multi-agents will use the collected information to make decisions for the current time slot, thus accelerating the learning process. The state of the i-th agent at time slot t is is defined as:

[0079]

[0080] So the multi-agent state s at time slot t t Defined as:

[0081] Reward: The optimization goal of this paper is to maximize the sum of the information transmission rates of all UEs. The reward r of time slot t is t Defined as:

[0082]

[0083] In formula (14), It is expressed as the sum of the information transmission rates of all UEs in time slot t. ω Indicates the penalty for violating the STAR-RIS deployment area Ω A Constraints, pω ≤0. When ω A ∈Ω A When p ω = 0. Otherwise p ω <0. It is used to punish UE k for violating the outage probability constraint, and It is used to penalize UE k for falling below the SINR constraint. λ1 and λ2 are the weight parameters of the penalty term, which are used to adjust the intensity of the penalty. is the indicator function, when P out (R k )>P out,max,k hour, The value of P is 1. out (R k )≤P out,max,k hour, The value of is 0. Same thing.

[0084] Based on s t 、a t and r t This paper uses the DDPG algorithm to optimize ω A , Φ, V and U. DDPG consists of four deep neural networks (DNNs): an Actor network, a Critic network, and two corresponding target networks.

[0085] The Actor network is responsible for learning and adjusting the optimization variables. It is based on the state s of the current time slot t. t , select the corresponding action a t The strategy is defined as:

[0086]

[0087] The Actor network usage weight is θ μ The strategy μ(s t |θ μ ) for action selection and adding Ornstein-Uhlenbeck noise To explore more actions. The Actor network optimizes θ by gradient ascent method μ To find the optimal deterministic policy μ(s) based on the action-value function Q(s,a) (Q value) t |θ μ ), the expected long-term reward is defined as:

[0088]

[0089] Where Q(s,a) is the global Q-value function of all agents. The cumulative discounted future reward R t Defined as γ∈[0,1] is the discount factor.

[0090] The weight θ of the Actor network μ Updates as follows:

[0091]

[0092] where α μ is the learning rate.

[0093] The objective function J(θ) of the Actor network represents the t Lower μ(s t The expected return that can be generated by |θ) is defined as:

[0094]

[0095] In formula (18), Q(s t ,μ(s t |θ)) evaluates the state s t The following μ(s t |θ), which is the value estimated by the Critic network. The Critic network uses Q A separate DNN is used to evaluate Q(s t ,μ(s t |θ μ )|θ Q The calculation of Q value is based on the Bellman equation approximation:

[0096]

[0097] r(s t ,a t ) means in state s t Take action a t Instant rewards received.

[0098] Target Q value y t Used to guide the update of the Critic network, defined as:

[0099] y t =r(s t ,a t )+γQ'(s t+1 ,μ'(s t+1 |θ μ' )|θ Q' ) (20)

[0100] The loss function L(θ Q ) is the mean square error between the target Q value and the current estimated Q value:

[0101]

[0102] θ of the Critic network Q Update via gradient descent:

[0103]

[0104] where α Q is the learning rate.

[0105] Since it is difficult to know t The exact probability distribution of , the expectations in Equations (18) and (21) can be approximated by the sample mean. Random samples are taken from the stored tuples Replay buffer t is the current time step, C is the replay buffer capacity.

[0106] Finally, through θ μ' ←τθ μ' +(1-τ)θ μ and θ Q' ←τθ Q' +(1-τ)θ Q Soft updates are performed on the target actor network and the target critic network to further enhance network stability. τ is a hyperparameter that controls the update speed of the target network and is typically set to a small value, τ<<1.

[0107] The settings of the simulation scene are as follows: Figure 3 As shown, the system consists of eight APs equipped with four antennas and 10 UEs, five of which have two antennas and five have four antennas. This setup is defined as a heterogeneous environment. Based on the distribution of APs and UEs in the existing environment, the algorithm derives three optimized STAR-RIS deployment locations and configuration parameters.

[0108] The optimized STAR-RIS is positioned closer to areas with dense user devices, thereby improving signal transmission quality and coverage in these areas. Furthermore, all STAR-RIS are deployed close to the APs. This deployment method helps reduce signal attenuation and interference during signal transmission, improves signal strength, and better controls and optimizes the signal path. Furthermore, different STAR-RIS devices have different reflection / refraction amplitude coefficients, indicating that different areas require different signal enhancement requirements and optimization strategies. This deployment method can more efficiently utilize signal resources and improve overall network performance, especially in situations where users are unevenly distributed but concentrated in certain areas.

[0109] To evaluate the performance of the STAR-RIS-assisted cell-free massive MIMO communication system under different configurations, this application compares multiple algorithms, including:

[0110] 1) Comparison with Algorithm 1: This method replaces the STAR-RIS in the environment with a passive intelligent reflecting surface (PRIS). Therefore, the optimization variables are adjusted to the transmit and receive beamforming matrices and the position coordinates, reflection / refraction amplitude coefficients, and phase shift matrix of the STAR-RIS.

[0111] 2) Comparison with Algorithm 2: The optimization variables of this method are adjusted to the transmit and receive beamforming matrices and the reflection / refraction amplitude coefficients of STAR-RIS and its phase shift matrix.

[0112] 3) Comparison with Algorithm 3: The optimization variables of this method are adjusted to the transmit and receive beamforming matrices and the phase shift matrix of STAR-RIS.

[0113] 4) Comparison with Algorithm 4: This method does not use smart reflective surfaces in its environment.

[0114] Figure 4 The impact of the number of UEs on the system information transmission rate is demonstrated. The evolution of the total system rate due to the number of UEs exhibits a three-stage characteristic. The DSR-JOS algorithm significantly delays the system from entering into performance bottlenecks by jointly optimizing the position coordinates, reflection amplitude coefficient, and phase shift matrix of STAR-RIS. In user-sparse scenarios, the linear growth rate of the DSR-JOS algorithm is ahead of the comparison algorithm because it dynamically adjusts the STAR-RIS position to spatially match the signal coverage with the user distribution. When the user scale expands to medium density, the DSR-JOS algorithm still maintains a good growth rate. It reconstructs the multipath propagation environment through STAR-RIS position optimization and postpones the system saturation threshold. In overload scenarios, the DSR-JOS algorithm achieves better interference suppression by optimizing the amplitude coefficient, resulting in slower rate attenuation.

[0115] Data shows that, under the same STAR-RIS architecture, the DSR-JOS algorithm improves system capacity limits and delays the overload threshold for the number of user devices compared to comparison algorithms 2 and 3. When the number of UEs exceeds 40, the DSR-JOS algorithm exhibits less information transmission rate degradation than the comparison algorithms. Comparison algorithm 1, due to its ability to dynamically optimize PRIS deployment locations, effectively utilizes PRIS to optimize channel and interference conditions as the number of users increases, especially in high-density or overload scenarios, resulting in higher information transmission rates than comparison algorithms 2 and 3. However, due to PRIS's limitations, its performance is inferior to that of the DSR-JOS algorithm. Comparison algorithm 4 exhibits the worst performance. These results demonstrate the synergistic effect of the joint optimization strategy: location optimization expands spatial degrees of freedom by reducing channel matrix correlation, while dynamic amplitude adjustment enables the system to maintain a stable SINR in high-density scenarios. This demonstrates that the coordinated optimization of location and amplitude can fully unleash the proactive control potential of STAR-RIS and further delay system overload.

[0116] Figure 5 The impact of the number of APs on the system information transmission rate was characterized. As the number of APs increased from 4 to 16, the information transmission rate of all schemes showed a trend of initially rapidly increasing and then gradually leveling off. Initially, the increase in APs made it easier for each UE to find a closer AP, significantly reducing path loss. This led to a rapid initial performance improvement. However, as the number of APs continued to increase, system complexity and coordination difficulties gradually increased. The increased number of APs in a complex environment increased inter-AP interference, making interference management and precoding more difficult and degrading the coordinated channel matrix condition number. These factors gradually reduced the improvement in the overall information transmission rate. The DSR-JOS algorithm performed optimally across all AP number variations. This was due to the dynamic optimization of STAR-RIS, which played a significant role in reducing path loss, enhancing spatial diversity and coordination, and compensating for signal loss. Comparison algorithm 1 performed worse than DSR-JOS due to its PRIS performance disadvantage, while comparison algorithms 2 and 3 performed worse than DSR-JOS due to their inability to dynamically adjust STAR-RIS parameters to find the optimal channel connection path. Comparison algorithm 4 performed the worst.

[0117] Figure 6The impact of the number of STAR-RISs on the system information transmission rate was characterized. As the number of RISs increased from 1 to 8, the information transmission rates of all schemes showed an initial upward and then downward trend. When the number of STAR-RISs in the environment initially increased, the DSR-JOS algorithm dynamically optimized the STAR-RIS position coordinates and reflection / refraction amplitude coefficients to find more efficient reflection paths, significantly reducing path loss. However, as the number of STAR-RISs continued to increase, the increased number of STAR-RISs led to an increase in the condition number of the cooperative channel matrix, affecting the performance of the precoding algorithm. The DSR-JOS algorithm performed best in all stages of the RIS number variation, especially in the initial stages, where it exhibited a significant rate improvement advantage. Comparative Algorithm 1 ranked second due to its PRIS performance disadvantage. Comparative Algorithms 2 and 3 offered almost no performance improvement compared to the DSR-JOS algorithm, especially in complex environments. Comparative Algorithm 4 performed the worst.

[0118] In summary, the present invention solves the resource allocation problem in a heterogeneous environment by constructing a STAR-RIS-CF-mMIMO system and proposing a solution. The solution uses the DDPG algorithm to jointly optimize the STAR-RIS position coordinates, phase shift matrix and its reflection amplitude coefficient, and transmit and receive beamforming matrices, thereby improving the information transmission rate of the system under different configurations. Simulation results show that the DSR-JOS algorithm proposed in the embodiment of the present application can maintain good performance in various configurations of RIS, AP and UE. These results fully demonstrate the potential and advantages of the solution proposed in the embodiment of the present application in improving the information transmission rate of the system and adapting to complex user communication environments. Therefore, the solution has broad practical application prospects and important significance, and helps to promote the integration of multi-access environments and various communication technologies.

[0119] The key technology of the present invention lies in the resource optimization method using deep reinforcement learning and its application in the STAR-RIS-CF-mMIMO communication system. The present invention optimizes the following technical parameters by utilizing the deep deterministic policy gradient (DDPG) algorithm: (1) the configuration strategy of the receive and transmit beamforming matrices; (2) the adjustment method of the STAR-RIS position coordinates, phase shift matrix and its reflection amplitude coefficient. The present invention can obtain the optimal multiple STAR-RIS position coordinates and their configuration parameters based on the current position distribution of user equipment and access points. In the scenario where the user equipment is configured with heterogeneous antennas, the above optimization scheme can ensure efficient resource allocation and significantly improve the total information transmission rate of the system.

[0120] This application also provides a STAR-RIS assisted non-cellular massive MIMO resource allocation system, such as Figure 7As shown, it includes: a collection module 72, which is configured to collect the location information of the user equipment UE, access point AP and simultaneously transmitting and reflecting reconfigurable smart surface STAR-RIS in the current system, and obtain the channel state information CSI of each communication link in the current system, and construct the state information set of the current system based on the location information and the CSI; an optimization module 74, which is configured to generate an optimization action based on the state information set using the Actor network in the deep deterministic policy gradient DDPG algorithm, wherein the optimization action includes at least one of the following: transmitting beamforming matrix, receiving beamforming matrix, STAR-RIS The IS position coordinates, the STAR-RIS reflection amplitude coefficient, and the STAR-RIS phase shift matrix; an adjustment module 76 is configured to adjust the resources between the UE and the AP based on the optimization action, perform downlink signal transmission based on the adjusted resources, and calculate the UE's signal-to-noise ratio (SINR) and the information transmission rate per unit time; an update module 78 is configured to output an instantaneous reward value under the optimization action based on the signal-to-noise ratio (SINR) and the information transmission rate, and optimize the weight parameters of the Actor network and the Critic network in the DDPG algorithm through gradient update based on the instantaneous reward value.

[0121] It should be noted that the STAR-RIS-assisted non-cellular massive MIMO resource allocation system provided in the above embodiment is only illustrated by the division of the above-mentioned functional modules. In actual applications, the above-mentioned functional allocation can be completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the STAR-RIS-assisted non-cellular massive MIMO resource allocation system provided in the above embodiment and the STAR-RIS-assisted non-cellular massive MIMO resource allocation method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0122] Figure 8 Schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present disclosure is shown. Figure 8 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0123] like Figure 8As shown, the electronic device includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 into the random access memory (RAM) 1003. Various programs and data required for system operation are also stored in the RAM 1003. The CPU 1001, ROM 1002 and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0124] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a LAN card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1010 as needed, so that computer programs read therefrom can be installed into the storage section 1008 as needed.

[0125] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A STAR-RIS assisted resource allocation method for non-cellular massive MIMO, characterized in that: include: Collecting location information of user equipment (UE), access point (AP), and simultaneously transmitting and reflecting reconfigurable smart surface (STAR-RIS) in the current system, and obtaining channel state information (CSI) of each communication link in the current system, and constructing a state information set of the current system based on the location information and the CSI; Based on the state information set, an optimization action is generated using an Actor network in a deep deterministic policy gradient (DDPG) algorithm, wherein the optimization action includes at least one of the following: a transmit beamforming matrix, a receive beamforming matrix, a position coordinate of STAR-RIS, a reflection amplitude coefficient of STAR-RIS, and a phase shift matrix of STAR-RIS; Adjusting resources between the UE and the AP based on the optimization action, performing downlink signal transmission based on the adjusted resources, and calculating an interruption probability, a signal-to-noise ratio (SINR), and an information transmission rate per unit time of the UE; Based on the interruption probability, the signal-to-noise ratio (SINR), and the information transmission rate, an immediate reward value under the optimization action is output, and based on the immediate reward value, weight parameters of the actor network and the critic network in the DDPG algorithm are optimized through gradient update.

2. The method according to claim 1, characterized in that Collecting location information of a user equipment UE, an access point AP, and a simultaneously transmitting and reflecting reconfigurable smart surface STAR-RIS in a current system, and obtaining channel state information CSI of each communication link in the current system, and constructing a state information set of the current system based on the location information and the CSI, including: Collecting location information of the UE, the AP, and the STAR-RIS, and normalizing the location information to a predefined coordinate system to obtain normalized location information; Acquire channel state information (CSI) of the communication links between the UE and the AP, between the AP and the STAR-RIS, and between the STAR-RIS and the UE, wherein the CSI includes: path loss, Rayleigh fading parameters, and LOS / NLOS flags; The location information and the CSI are mapped into a state vector to construct the state information set for input to the DDPG algorithm.

3. The method according to claim 1, characterized in that Based on the state information set, the Actor network in the Deep Deterministic Policy Gradient (DDPG) algorithm is used to generate optimization actions, including: Inputting the state vector in the state information set into the Actor network constructed by the deep neural network; Generating, at the output layer of the Actor network, a phase shift matrix, a reflection amplitude coefficient, and position coordinates for controlling the STAR-RIS, as well as optimization parameters for transmit and receive beamforming matrices of the AP and UE; The optimization action is generated based on the optimization parameters, and regularization and restriction functions are applied to the optimization action so that the optimization action satisfies power constraints, position constraints, amplitude physical constraints, and phase physical constraints.

4. The method according to claim 1, wherein Adjusting resources between the UE and the AP based on the optimization action includes at least one of the following: Allocate the transmit beamforming matrix output by the Actor network to the corresponding AP to determine its downlink transmit signal structure for each UE; Assigning a receive beamforming matrix to the corresponding UE and determining a weighted combining strategy for its received signal; The reflection and transmission paths are updated according to the position and parameter configuration of the STAR-RIS, and a UE aggregation channel matrix including direct links and indirect links is constructed.

5. The method according to claim 1, wherein Outputting an immediate reward value under the optimization action based on the interruption probability, the signal-to-noise ratio (SINR), and the information transmission rate includes: The sum of the transmission rates of all UEs is used as the initial reward value for this time slot; In case of violating the outage probability threshold, the lower limit of the signal-to-noise ratio SINR, or the STAR-RIS deployment restriction, a penalty term is introduced to adjust the initial reward value to form the final instant reward value.

6. The method according to claim 5, characterized in that The penalty item includes at least one of the following: When the outage probability of any UE is higher than the set outage probability threshold, a constraint penalty term is imposed; When the SINR of any UE is lower than the set threshold, a proportional penalty term is applied to reduce the total reward; When the location of any STAR-RIS exceeds the predefined deployment boundary, a regional constraint penalty is imposed.

7. A STAR-RIS-assisted non-cellular massive MIMO resource allocation system, characterized in that: include: a collection module configured to collect location information of user equipment (UE), access point (AP), and simultaneously transmitting and reflecting reconfigurable smart surface (STAR-RIS) in the current system, obtain channel state information (CSI) of each communication link in the current system, and construct a state information set of the current system based on the location information and the CSI; An optimization module is configured to generate an optimization action based on the state information set and using an Actor network in a deep deterministic policy gradient (DDPG) algorithm, wherein the optimization action includes at least one of the following: a transmit beamforming matrix, a receive beamforming matrix, a position coordinate of STAR-RIS, a reflection amplitude coefficient of STAR-RIS, and a phase shift matrix of STAR-RIS; an adjustment module configured to adjust resources between the UE, the AP, and the STAR-RIS based on the optimization action, perform downlink signal transmission based on the adjusted resources, and calculate an interruption probability, a signal-to-noise ratio (SINR), and an information transmission rate per unit time of the UE; An update module is configured to output an immediate reward value under the optimization action based on the interruption probability, the signal-to-noise ratio (SINR), and the information transmission rate, and optimize the weight parameters of the actor network and the critic network in the DDPG algorithm through gradient update based on the immediate reward value.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 6.

9. A computer device, characterized in that: include: memory and processor, The memory stores a computer program; The processor is configured to execute a computer program stored in the memory, wherein the computer program enables the processor to execute the method according to any one of claims 1 to 6 when the program is executed.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Adaptive adjustment method and system of antenna array and medium

    CN120914507A

  • Adaptive adjustment method, system and medium for antenna array

    CN120914507B