Distributed RIS-assisted ISAC system and resource joint optimization method
By combining distributed RIS and NOMA technologies, and using graph neural networks and multi-agent deep reinforcement learning methods to optimize RIS and user pairing and base station active beamforming, the problems of signal obstruction and small coverage range of the ISAC system in the mmWave frequency band are solved, and the total system throughput is maximized and the computational complexity is reduced.
Patent Information
- Application Number
- CN202510736044.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-19
AI Technical Summary
In existing technologies, the ISAC system, when combined with NOMA technology in the mmWave frequency band, faces problems such as signal obstruction and limited coverage. A single RIS-assisted system cannot effectively cover all communication users and sensing targets. Traditional algorithms have high computational complexity in scenarios with large numbers of users and sensing targets, and lack optimization for RIS and user pairing.
A system framework that combines distributed RIS, NOMA, and ISAC technologies is adopted. Through a two-stage optimization method based on graph neural networks and multi-agent deep reinforcement learning, RIS and user pairing, base station active beamforming, and RIS passive beamforming are optimized to maximize the total system throughput.
It effectively improves the total throughput of the system, achieves a balance between communication and perception performance, reduces computational complexity, and improves spectrum utilization.
Smart Images

Figure CN120676378A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a distributed RIS-assisted ISAC system and a resource joint optimization method. Background Art
[0002] Over the past few years, mmWave communications, with their wide bandwidth capabilities, have become a key technology for meeting the demand for high-speed data transmission and expanding capacity. However, mmWave signals face significant challenges due to their susceptibility to blockage and limited ability to penetrate obstacles. To overcome these inherent limitations, RIS has emerged as a particularly innovative technology. It consists of a planar array with passive reflective elements (REs) that intelligently manipulate electromagnetic signals to overcome channel limitations and provide superior performance compared to traditional relay methods. In complex wireless environments, such as urban areas, multiple RISs can be deployed to facilitate communication between base stations and users, thereby establishing virtual line-of-sight (VLoS) paths between the base stations and users. Furthermore, non-orthogonal multiple access (NOMA) has emerged as a promising technology that further improves spectral efficiency by multiplexing communicating users in the power domain. Unlike traditional orthogonal multiple access (OMA) schemes, NOMA can significantly reduce inter-user interference using successive interference cancellation (SIC), making it particularly well-suited for the massive connectivity requirements of future communication and perception networks. Therefore, integrating NOMA with mmWave communications and RIS technologies can overcome the inherent limitations of mmWave propagation while maximizing spectrum utilization.
[0003] On the other hand, since wireless spectrum resources are limited, next-generation wireless communication systems aim to achieve perception. Ideally, this would reuse existing communication frequency bands, enabling spectrum sharing for both perception and communication, thereby improving spectrum utilization. Consequently, ISAC technology has emerged. It integrates the hardware infrastructure for communication and perception functions into base stations, enabling both communication and perception to operate within the same time-frequency resources. To effectively improve perception and communication performance in a distributed RIS-assisted mmWave ISAC system, an effective method is needed to solve the multivariable joint optimization problems of RIS-user pairing, active beamforming, and passive beamforming. Summary of the Invention
[0004] In order to at least partially solve one of the technical problems existing in the prior art, an object of the present invention is to provide a distributed RIS-assisted ISAC system and a resource joint optimization method.
[0005] The first technical solution adopted by the present invention is:
[0006] A distributed RIS-assisted ISAC system includes a base station, multiple distributed RISs, and multiple user terminals;
[0007] Distributed RIS assists user terminals in mmWave communications. Each user terminal can be served by only one RIS, and each RIS can serve user terminals in one NOMA group. Within each NOMA group served by each RIS, target information for each user is decoded using SIC technology. Targets are directly detected by the base station, and the perceived beam gain is used as an indicator of perception performance.
[0008] The RIS and user pairing, base station active beamforming and RIS passive beamforming are jointly optimized to maximize the total throughput of the system under the premise of minimum beam gain for each sensing target.
[0009] Furthermore, the joint optimization step includes two stages:
[0010] In the first stage, a graph neural network-based method is used to optimize the pairing between RIS and users. The attention mechanism and residual connection technology are used to better extract key graph information to obtain a better pairing strategy and effectively avoid gradient vanishing.
[0011] In the second stage, based on the pairing strategy obtained in the first stage, the active beamforming of the base station and the passive beamforming of the RIS are jointly optimized using the multi-agent deep reinforcement learning (MADDPG) method. The base station and each RIS are constructed as an independent agent, and each agent consists of an actor network, a critic network, a target actor network, and a target critic network.
[0012] The second technical solution adopted by the present invention is:
[0013] A resource joint optimization method, applied to the above-mentioned distributed RIS-assisted ISAC system, comprises:
[0014] Construct a distributed RIS-assisted ISAC system model in the mmWave frequency band;
[0015] Construct mmWave channel modeling, ISAC system signal model, and communication and perception model based on the system model;
[0016] Determine the optimization problem based on the constructed model;
[0017] Based on the optimization problem, RIS and user pairing, base station active beamforming and RIS passive beamforming are jointly optimized to maximize the total throughput of the system under the premise of minimum beam gain for each sensing target.
[0018] Furthermore, the ISAC system model in the mmWave frequency band assisted by the distributed RIS includes:
[0019] K single-antenna communication users complete the communication with the base station with the assistance of J RIS; each RIS with N reflective elements is modeled as a uniform planar array (UPA), and the adjacent reflective elements of the RIS are separated by d x and d y The phase shift matrix of the jth RIS is set to Θ j , and considering the constraints of actual hardware, discrete phase shift is used to represent the phase shift matrix, that is, the phase shift of each reflection unit of RIS can only take values in a limited number of phase shifts;
[0020] The base station is equipped with M antennas forming a uniform linear array (ULA); in order to achieve interawareness integration, T sensing targets are sensed directly through the base station;
[0021] The sets of user terminals, RIS and sensing targets are defined as and
[0022] Furthermore, the mmWave channel modeling is constructed as follows:
[0023] The mmWave channel is characterized by a line-of-sight (LoS) link and multiple non-line-of-sight (NLoS) links; the channel between the base station BS and the j-th RIS is represented by G 0,j Denote by , and the channel between the jth RIS and the kth user terminal is represented by h j,k Represented, and all mmWave channels are modeled using the Saleh-Valenzuela channel model;
[0024] In addition, considering the passivity of RIS, it becomes extremely difficult to obtain channel state information (CSI), which introduces CSI errors and only imperfect CSI can be obtained.
[0025] Furthermore, the ISAC system signal model is constructed as follows:
[0026] In order to simultaneously serve K communication users and T sensing targets, the base station needs to transmit superimposed NOMA and sensing signals, which are expressed as follows:
[0027] x=W c s c +W r s r =Ws
[0028] Where W c represents the active communication beamforming matrix of the base station, W r represents the active sensing beamforming matrix of the base station; s c Indicates that sr represents the transmission signals of K user terminals and T sensing targets; W,s represent the total active beamforming matrix and the total transmission signal respectively;
[0029] Assume that the kth user terminal matches the μth k RIS, then the signal received by the kth user terminal is expressed as:
[0030]
[0031] Where, Indicates μth k The channel from the RIS to the kth user, H represents the conjugate transpose operation of the matrix, w k,c represents the communication beamforming vector of the kth user, s k,c represents the communication signal that the kth user expects to receive, w d,c represents the communication beamforming vector of the dth user, s d,c represents the communication signal that the dth user expects to receive, w t,r represents the sensing beamforming vector of the t-th target, s t,r represents the communication signal that the tth target expects to receive, n k represents Gaussian white noise.
[0032] Furthermore, the communication and perception model is constructed as follows:
[0033] The effective channel vector of the kth user terminal is defined as Assume that the communication users are sorted in descending order of channel gain; the kth user terminal first detects and removes the interference of the weak user of the same RIS service, while considering the signal of the strong user of the same RIS service and the signals of all other RIS service users as interference; then the received signal to interference plus noise ratio (SINR) of the kth user terminal is expressed as:
[0034]
[0035] Where, represents the noise power;
[0036] In order to make SIC decoding proceed smoothly, the signal of the kth user terminal needs to be decoded at the weak user n of the same RIS service, and the corresponding SINR is:
[0037]
[0038] Then, based on the SIC principle, the achievable rate of user k is:
[0039]
[0040] where R n→k =log2(1+SINR n→k ) and R k→k =log2(1+SINR k→k ) represent the data rate obtained by weak user n decoding user k and the data rate obtained by strong user decoding its own target signal, respectively. The total throughput of the system is obtained by summing the rates of all communicating users:
[0041]
[0042] In terms of perception, the active beamforming matrix W of the base station is used to perceive the target; the base station is defined as receiving the angle θ from the perceived target. t The beam gain is:
[0043]
[0044] Where, α L (·) represents the array response of a uniform linear array.
[0045] Furthermore, the optimization problem is expressed as:
[0046]
[0047] In the optimization problem (P0), μ and Θ represent the set of pairing variables between RIS and users and the set of phase shift matrices of RIS, respectively; P max Indicates the maximum transmission power of the base station, R k represents the achievable rate of the kth user, R min represents the minimum achievable rate constraint for each user, Indicates that the base station receives the signal from the sensing target with an angle of θ t The beam gain, Indicates the minimum beam gain received by the base station for each sensing target, φ j,n represents the phase shift of the nth reflector unit of the jth RIS, represents the set of possible values of the discrete phase shift of the RIS reflection unit; constraint C1 limits each user to only one RIS; C2 represents the maximum transmit power limit of the system; C3 guarantees the minimum rate of each communicating user; C4 ensures the minimum beam gain of each sensing target received by the base station; C5 indicates that the considered RIS phase shift is a discrete phase shift.
[0048] Furthermore, the joint optimization of RIS and user pairing, base station active beamforming, and RIS passive beamforming includes:
[0049] In the first stage, a graph neural network-based method is used to optimize the pairing between RIS and users. The attention mechanism and residual connection technology are used to better extract key graph information to obtain a better pairing strategy and effectively avoid gradient vanishing.
[0050] In the second stage, based on the pairing strategy obtained in the first stage, the active beamforming of the base station and the passive beamforming of the RIS are jointly optimized using the multi-agent deep reinforcement learning (MADDPG) method. The base station and each RIS are constructed as an independent agent, and each agent consists of an actor network, a critic network, a target actor network, and a target critic network.
[0051] Furthermore, during the training process, each agent inputs the observed state into the actor network to obtain the action to be implemented. Then the environment evaluates the action of each agent and updates the critic network and actor network. Finally, a soft update of the target actor network and the target critic network is performed.
[0052] The beneficial effects of the present invention are as follows: the present invention proposes a system framework that combines distributed RIS, mmWave, NOMA and ISAC technologies, and performs two-stage optimization based on machine learning on the optimization problems involved in RIS and user pairing, base station active beamforming and RIS passive beamforming, effectively improving the total throughput of the system and achieving a better balance between communication and perception performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following introduction is made to the drawings of the embodiments of the present invention or the related technical solutions in the prior art. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative work.
[0054] Figure 1 Schematic diagram of the structure (system transmission model) of the ISAC system in the mmWave frequency band assisted by a distributed RIS in an embodiment of the present invention;
[0055] Figure 2 A schematic diagram illustrating how the total throughput of the system is affected by the number of base station antennas under different schemes in a distributed RIS-assisted mmWave band ISAC system considered in an embodiment of the present invention;
[0056] Figure 3A schematic diagram illustrating how the total throughput of the system is affected by the maximum transmit power of the base station under different schemes in a distributed RIS-assisted mmWave band ISAC system considered in an embodiment of the present invention;
[0057] Figure 4 A schematic diagram illustrating the impact of the number of sensing targets and the minimum beam gain of sensing targets on the total throughput of the system under different schemes in a distributed RIS-assisted mmWave band ISAC system considered in an embodiment of the present invention;
[0058] Figure 5 Schematic diagram of beam gain at various angles under different schemes in a distributed RIS-assisted mmWave band ISAC system considered in an embodiment of the present invention. DETAILED DESCRIPTION
[0059] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and are not to be construed as limiting the present application. For the step numbers in the following embodiments, they are provided only for the convenience of explanation and are not intended to limit the order of the steps. The order of execution of the steps in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0060] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the embodiments of the present application. The singular forms of "a", "said", and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise clearly defined, words such as setting, installing, and connecting should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.
[0061] In the description of this application, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on this application.
[0062] In the description of this application, "several" means one or more, "many" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, while "above," "below," and "within" are understood to include the number itself. The terms "first" and "second" are used solely to distinguish technical features and are not to be construed as indicating or implying relative importance, or as implicitly specifying the number or order of the technical features indicated.
[0063] In the description of this application, "and / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0064] Explanation of terms:
[0065] ISAC: abbreviation of Integrated Sensing and Communication, integrated communication and perception technology.
[0066] RIS: abbreviation of Reconfigurable Intelligence Surface, intelligent reflective surface.
[0067] Due to the limited availability of wireless spectrum resources and the increasing need for sensing in wireless communication networks, integrated mmWave communication and sensing systems have been extensively researched. To address the issue of mmWave signals being easily obscured by buildings, academics have also conducted extensive research on distributed RIS-assisted mmWave ISAC communication systems. However, existing technical solutions have the following shortcomings:
[0068] 1) Current research on mmWave ISAC systems focuses on the application of NOMA technology in mmWave ISAC systems, or on applying RIS to mmWave-ISAC systems. However, the combination of mmWave ISAC systems and NOMA technology faces challenges: mmWave signals are susceptible to obstruction and have limited coverage. RIS-assisted mmWave ISAC systems are also subject to the limitations of traditional OMA technology, resulting in limited spectrum utilization.
[0069] 2) Current research on RIS-assisted mmWave-ISAC systems focuses on single-RIS-assisted mmWave-ISAC systems. However, in real-world applications, mmWave signals face the challenge of limited coverage due to obstructions such as buildings. This prevents a single RIS from effectively covering all communication users and sensing targets, resulting in severe performance losses.
[0070] 3) Current research on performance optimization of distributed RIS-assisted mmWave-ISAC systems focuses primarily on optimizing active and passive beamforming, with little attention paid to optimizing the pairing between RIS and users. However, in practical applications where multiple RISs assist multiple users in communication and perception, each RIS can only serve a limited number of users, and each user's channel conditions vary. Therefore, it is necessary to match each user with the optimal RIS based on actual conditions to maximize the overall system throughput.
[0071] 4) Current research on performance optimization of distributed RIS-assisted mmWave-ISAC systems primarily focuses on iterative optimization using traditional algorithms, such as the alternating order method (AO) and the BCD method. These traditional methods offer acceptable complexity in systems with a small number of RIS and users. However, in real-world applications, systems often involve a large number of users and sensing targets, necessitating the deployment of large-scale distributed RIS to assist in communication and perception. Traditional algorithms face significant computational complexity when performing joint optimization of multiple parameters in these systems, significantly hindering the realization of real-time communication and perception.
[0072] Taking into account the shortcomings of single RIS in actual application scenarios, in order to solve the problem of limited propagation range faced by the ISAC system in the mmWave frequency band assisted by a single RIS, the present invention proposes a framework that combines distributed RIS, NOMA, mmWave communication and ISAC technology, and studies the application of machine learning methods therein. Among them, communication users achieve high-quality communication with the base station through the assistance of multiple RIS, while the perception targets are directly detected through the base station. Based on this framework, the present invention also proposes an efficient multi-dimensional resource allocation method based on machine learning, which realizes the pairing of RIS and users, and the joint optimization of active beamforming and passive beamforming, maximizing the total throughput of communication users in the system while ensuring the minimum perception performance of the perception targets.
[0073] Specifically, the present invention considers an ISAC system in the mmWave frequency band assisted by a distributed RIS, wherein the distributed RIS assists communication users in performing mmWave communication, each user can only be served by one RIS, and each RIS can serve users in a NOMA group. In the NOMA group served by each RIS, the target information of each user is decoded by the SIC technology. The perceived target is directly detected by the base station, and the perceived beam gain is used as an indicator of the perception performance. In order to maximize the total throughput of the system while ensuring the minimum beam gain of each perceived target, the present invention also proposes a multi-dimensional resource joint optimization method based on machine learning for RIS and user pairing, base station active beamforming, and RIS passive beamforming. Specifically, the joint optimization method can be divided into two stages. In the first stage, the present invention proposes a method based on graph neural networks to optimize the pairing between RIS and users, and adopts advanced attention mechanisms and residual connection technologies, so that the method can better extract graph key information to obtain a better pairing strategy and effectively avoid gradient disappearance. Based on the pairing strategy obtained in the first stage, in the second stage, the present invention proposes a joint optimization of the base station's active beamforming and the RIS's passive beamforming based on the multi-agent deep reinforcement learning MADDPG method. In this method, the base station and each RIS are constructed as independent agents, each consisting of an actor network, a critic network, a target actor network, and a target critic network. During the training process, each agent inputs the observed state into the actor network to obtain the action to be implemented. The environment then evaluates and updates the critic network and actor network based on the actions of each agent. Finally, a soft update of the target actor network and the target critic network is performed. Simulation results show that the two-stage multi-dimensional resource joint optimization method based on machine learning proposed in this invention brings very significant performance improvements in optimizing RIS and user pairing and base station and RIS beamforming in the ISAC system assisted by distributed RIS in the mmWave frequency band.
[0074] The technical solution of the present invention is explained in detail below with reference to the accompanying drawings and specific embodiments.
[0075] (1) System model
[0076] See also Figure 1 In the ISAC system model in the mmWave band assisted by distributed RIS considered in this embodiment, K single-antenna communication users complete the communication with the base station with the assistance of J RIS. x N y The RIS with 1 reflective element can be modeled as a uniform planar array (UPA), and the distance between adjacent reflective elements of the RIS is dx and d y The phase shift matrix of the jth RIS is set to Considering the constraints of actual hardware, discrete phase shift is used to represent the phase shift matrix. That is, the phase shift of each reflector unit of RIS can only take values in a limited number of phase shifts, namely:
[0077]
[0078] The base station is equipped with M antennas to form a uniform linear array (ULA). In order to achieve inter-sensory integration, T sensing targets are directly sensed through the base station. For the convenience of representation, this embodiment defines the set of users, RIS and sensing targets as and In order to be closer to the actual application scenario, the system considers K>J, that is, the number of users is greater than the number of RISs, which further shows that the pairing between RISs and users is very important and meaningful.
[0079] (2) mmWave channel modeling
[0080] The mmWave channel can be characterized by a line-of-sight (LoS) link and multiple non-line-of-sight (NLoS) links. The channel between the base station BS and the j-th RIS is represented by G 0,j and the channel between the j-th RIS and the d-th user is represented by h j,k Represented, and all mmWave channels are modeled using the Saleh-Valenzuela channel model as follows:
[0081]
[0082]
[0083] Among them, L G and L h Represents G 0,j and h j,k The number of paths. l , and Represents G 0,j The complex Gaussian gain of the lth path, arrival azimuth, arrival elevation and exit elevation, ε l , and Respectively represent h j,k The complex Gaussian gain of the lth path, the exit azimuth angle and the exit elevation angle.
[0084]
[0085]
[0086] Where λ represents the wavelength of the millimeter wave signal, d represents the distance between base station antennas, and the distance between adjacent reflection units of RIS is d. x and d y Separated by horizontal and vertical spacing.
[0087] In addition, considering the passivity of RIS, it becomes extremely difficult to obtain channel state information (CSI), so we introduce CSI error. Therefore, this embodiment considers that only imperfect CSI can be obtained, that is:
[0088]
[0089] h j,k =h j,k +Δh j,k , (7)
[0090] Among them, ΔG 0,j and Δh j,k They represent the errors generated when obtaining CSI, and h j,k Indicates the imperfect CSI obtained.
[0091] (3) ISAC system signal model in the mmWave band assisted by distributed RIS
[0092] In order to simultaneously serve K communication users and T sensing targets, the base station needs to transmit superimposed NOMA and sensing signals as follows
[0093] x=W c s c +W r s r =Ws, (8)
[0094] in, represents the active communication beamforming matrix of the base station, It represents the active sensing beamforming matrix of the base station. and Denote the active beamforming vectors of the k-th user and the t-th sensing target respectively. W=[W c ,W r ] and s=[s c ,s r ] T denote the total active beamforming matrix and the total transmitted signal, respectively. and denote the transmission signals of K communication users and T sensing targets respectively. In addition, in order to facilitate the expression of subsequent formulas, we assume that the k-th user matches μ k -th RIS, where Then the signal received by the k-th user can be expressed as:
[0095]
[0096] The first part of the above formula represents the target desired signal received by the k-th user, the second and third parts represent the communication interference signal and the perceived interference signal received by the user, respectively, and the fourth part represents the received additive white Gaussian noise (AWGN), where σ 2 Represents the noise power.
[0097] (4) Communication and perception model
[0098] For the convenience of representation, the effective channel vector of the k-th user is defined as According to NOMA theory, within each NOMA group, the strong user decodes the weak user's signal through SIC before decoding its own target signal. Without loss of generality, assume that the communication users are sorted in descending order of channel gain, that is, In addition, in order to eliminate interference, it is assumed that the perception signal is a priori knowledge for the communication user, that is, the communication user will not be interfered by the perception signal.
[0099] Therefore, the k-th communication user first detects and removes the interference of weak users in the same RIS service, and at the same time considers the signals of strong users in the same RIS service and the signals of all other users in the RIS service as interference. The received signal-to-interference-plus-noise ratio (SINR) of the k-th user can be expressed as:
[0100]
[0101] In addition, in order to make SIC decoding proceed smoothly, the signal of the k-th user needs to be decoded at the weak user n served by the same RIS, that is, The corresponding SINR is:
[0102]
[0103] Then, based on the SIC principle, the achievable rate of user k is:
[0104]
[0105] where R n→k =log2(1+SINR n→k ) and R k→k =log2(1+SINR k→k) represent the data rate obtained by weak user n decoding user k and the data rate obtained by strong user decoding its own target signal. The total throughput of the system can be obtained by summing the rates of all communicating users, that is:
[0106]
[0107] In terms of perception, the base station's active beamforming matrix It is used to sense the target. The base station is defined as receiving the angle from the sensed target as θ t The beam gain is:
[0108]
[0109] (5) Maximizing the total system throughput
[0110] With the above definition completed, the present invention aims to maximize the total throughput while meeting the minimum perceptual performance of each perceptual target by jointly optimizing the pairing relationship between RIS and users, the base station active beamforming matrix, and the RIS passive beamforming. The mathematical expression of the optimization problem is:
[0111]
[0112] In the optimization problem (P0), and The constraints C1 and C2 represent the set of RIS-user pairing variables and the set of RIS phase shift matrices, respectively. Constraint C1 limits each user to only one RIS. C2 represents the maximum transmit power limit for the system. C3 guarantees the minimum rate for each communicating user. C4 ensures the minimum beam gain for each perceived target received by the base station. C5 indicates that the considered RIS phase shift is a discrete phase shift.
[0113] (6) Two-stage resource optimization method
[0114] For the above optimization problem, constraints C2, C3, C4 and C5 are all non-convex constraints. In addition, the phase shift matrices of the J distributed RIS are highly correlated, and the optimization parameters μ, W and Θ are interrelated and highly coupled, which brings challenges to optimizing the above problems. In summary, the optimization problem (P1) is a highly complex non-convex optimization problem. The traditional alternating iterative method requires huge computing resources when solving such problems, and the optimization results obtained are often not optimal. In order to more effectively optimize this multi-parameter and multi-dimensional resource allocation problem, the present invention proposes a two-stage resource optimization method based on machine learning. Specifically, in the first stage, a method based on a graph neural network algorithm is proposed, and an advanced graph attention network framework is adopted to pair RIS and users. In the second stage, based on the pairing results of RIS and users obtained in the first stage, a method based on deep reinforcement learning is adopted to optimize the design of base station active beamforming W and RIS passive beamforming Θ. Specifically:
[0115] 5.1) Phase 1: Optimizing the pairing of RIS and users
[0116] Based on the idea of graph theory, base stations, RIS and communication users are first defined as nodes of the graph, expressed as Where node 0 represents a base station, node j∈{1,2,…,J} represents J distributed RISs, and node k+J∈{1+J,…,K+J} represents K communicating users. It is worth noting that after this definition, node j represents the j-th RIS, and node k+J represents the k-th user. In actual deployment scenarios, a user may be located in the service areas of multiple RISs, so it is very necessary to match an optimal RIS for each user. For the convenience of description, we introduce χ j,k+J ∈{0,1} to describe the possible pairing relationship between RIS and users, where χ j,k+J =1 means there is a possible pairing relationship between the j-th RIS and the k-th user, χ j,k+J = 0 means that the j-th RIS and k-th user cannot be paired. In addition, the node characteristics are defined as is the three-dimensional coordinate of the corresponding node.
[0117] Based on the above definitions, given the base station active beamforming and RIS passive beamforming, the optimization problem (P0) is transformed into:
[0118]
[0119] The objective function It is calculated by substituting the given base station active beamforming matrix W and RIS passive beamforming matrix Θ into formula (13). Constraint C6 ensures that pairing can only be performed between RIS and users that have pairing possibilities.
[0120] Note that, since the objective function is calculated When the NOMA inter-group interference and NOMA intra-group interference are involved, the calculation of the above two parts is closely related to the pairing strategy. It is impossible to obtain prior information before determining the pairing strategy. Therefore, the present invention considers the calculation of the objective function Therefore, the problem of maximizing the total system throughput can be equivalently transformed into the problem of maximizing the total effective channel gain of the system, that is, maximizing:
[0121]
[0122] Given the base station active beamforming matrix W and the RIS passive beamforming matrix Θ, the above expression depends only on the channel between the base station and the RIS and the channel between RIS and users Therefore, in order to maximize the total effective channel gain of the system, we introduce the edge feature ε between the base station and RIS in the graph definition. 0,j and the edge feature ε between RIS and the user j,k+J is the corresponding channel gain, that is:
[0123]
[0124] In summary, the defined graph structure includes nodes, node features, edges and edge features, which can be expressed as in, Represents the set of edges in the graph, expressed as:
[0125]
[0126]
[0127] also, By node feature f n and edge feature ε a,b After completing the above definition, the pairing relationship between RIS and users can be solved by a graph neural network (GNN) method that adopts an attention mechanism. Specifically, the graph neural network model proposed in this embodiment includes an encoder and a decoder. The encoder is responsible for extracting and integrating node features and edge features, while the decoder is responsible for calculating the probability distribution of pairings between RIS and users, thereby selecting an optimal RIS for each user.
[0128] Encoder design: The encoder proposed in this embodiment adopts the residual connection graph attention network (RE-GAT) structure, which is used to integrate edge features into node features. The advantage is that it can more effectively extract and integrate key information in the graph structure to obtain a pairing solution with better performance, while also avoiding the gradient vanishing problem. The encoder converts the graph defined above into As input, the node features and edge features are first processed through the fully connected layer and batch normalization layer as follows:
[0129]
[0130]
[0131] Where Bn(·) and Fc(·) represent the batch normalization layer and the fully connected layer respectively. After the above processing, the node features and edge features is input into the first layer of the RE-GAT network and outputs the updated node features The edge features remain unchanged. and It is input to the second layer of the RE-GAT network, and so on. After being processed by the L-layer RE-GAT network, the node features are updated and the output f (L) .
[0132] To describe the above process more specifically, we will introduce the first layer of the RE-GAT network. First, to describe the importance weight between node a and the connected node b, we introduce the attention coefficient calculation as follows:
[0133]
[0134] where z l and Represent the learnable weight vector and matrix of the lth layer network, δ(·) represents the LeakyReLU activation layer in the RE-GAT network, and (·||·||·) represents the concatenation operation. In addition, between two adjacent layers, a residual connection is used to aggregate the updated node features and the input node features, and the output node features are as follows:
[0135]
[0136] in Represents the updated node features. After L RE-GAT layers complete the above update and aggregation process, the final output node features are obtained From this we can get the final graph embedding features in:
[0137]
[0138] Decoder Design: Based on the graph features extracted and integrated by the encoder, the decoder also adopts an attention mechanism, including an H-head attention layer and a single-head attention layer. The H-head attention layer calculates the context vector based on the graph features, while the single-head attention layer generates a probability distribution for pairing the user with the RIS, thereby helping the user select the optimal RIS. Taking the user node k+J matching the optimal RIS as an example, the corresponding context vector is obtained as follows:
[0139]
[0140] in Represents a learnable weight matrix. After the above context vector is input into the H-head attention layer, the output This paper adopts the (q, k, v) attention mechanism in the Transformer model. Specifically, Represent the query vector, key vector and value vector respectively, where Represents the corresponding learnable weight matrix. During the algorithm operation, the query vector q and the key vector k are used to calculate the attention coefficient between the RIS node i and the user node k+J, as shown below:
[0141]
[0142] where d t The dimension of query vector, key vector and value vector is represented by . Then, the present invention applies softmax activation function to normalize the attention coefficient calculated above.
[0143] The above calculation operation is repeated in each head of the H-head attention layer. For the convenience of representation, we define is the normalized attention coefficient calculated in the h-th head attention layer. Through a series of splicing operations, the updated context vector can be obtained
[0144]
[0145] After obtaining the updated context vector, it is input into the single-head attention layer of the decoder, and the tangent activation function is used to calculate the pairing probability of each user to each RIS as follows:
[0146]
[0147] Based on the above probabilities, this embodiment uses the softmax function to calculate the pairing probability distribution between user node k+J and all RISs, namely:
[0148]
[0149] After obtaining the pairing probability distribution of each user for each RIS, this embodiment also adopts a sampling strategy to determine the optimal RIS for each user node k + J. Based on the encoder and decoder structures defined above, the detailed graph neural network training process will be given below.
[0150] This embodiment uses the REINFORCE algorithm to train the graph neural network and applies a rolling baseline strategy to reduce the estimated variance of the policy gradient. Based on the traditional actor-critic network, the REINFORCE algorithm replaces the critic network with a baseline actor network to form a dual-actor network structure. This structure improves the stability and convergence speed of the algorithm. In each round of training, the trainable parameters of the baseline actor network remain unchanged, while the trainable parameters of the main actor network are continuously updated in real time. After completing each round of training, the training effect of the round is evaluated by the greedy decoding algorithm. Then, a t-test is used to decide whether to update the baseline policy. That is, the baseline policy is updated to the current policy only when the current policy can bring a sufficiently significant improvement. This not only improves the stability of training but also speeds up the convergence speed.
[0151] First, in order to maximize the total effective channel gain of the system, define:
[0152]
[0153] Then, introduce P ξ (μ) represents the probability of adopting the pairing strategy μ, which is calculated as follows:
[0154]
[0155] Where ξ represents the learnable parameters of the graph neural network, p ξ (i, k+J) represents the pairing probability distribution between the RIS output by the decoder and the user. Based on the above definition, the loss function in the training process is further defined as During the training process, the learnable parameter ξ is continuously updated by the policy gradient descent of the rollback baseline Γ, that is:
[0156]
[0157] The rollback baseline Γ is evaluated and updated by the baseline actor network. During training, the system state represents the currently established RIS and user pairings, the action space represents the process of each user selecting the optimal RIS, and the reward function is derived by quantifying the total effective channel gain of the system.
[0158] After effective training, the GNN network proposed in this embodiment can select an optimal RIS for each user to maximize the total channel gain of the system, thereby improving system performance and obtaining an approximately optimal pairing solution in the first stage.
[0159] 5.2) Phase 2: Jointly optimizing base station active beamforming and RIS passive beamforming using deep reinforcement learning algorithms
[0160] Based on the approximate optimal pairing scheme of RIS and users obtained in stage 1, the overall optimization problem can be further transformed into the joint optimization of base station active beamforming and RIS passive beamforming to maximize the total system throughput under various constraints, that is:
[0161]
[0162] To address this optimization problem, this embodiment proposes a joint optimization of base station active beamforming and RIS passive beamforming based on the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm. Leveraging its powerful exploration and learning capabilities, and through extensive training, the MADDPG algorithm can obtain near-optimal designs for base station active beamforming and RIS passive beamforming.
[0163] MADDPG framework: This embodiment models the base station and J RIS as DDPG agents, and each agent learns how to specify a better strategy to maximize the total reward of the system. Specifically, each agent can perform two functions: one is to implement actions through the actor network, and the other is to evaluate and update the strategy through the critic network. This enables multiple agents to perform their own strategy selection and network evaluation in parallel. Such a framework can promote the cooperation between the base station agent and the RIS agent to design effective active beamforming and passive beamforming. During the training process, this embodiment adopts the centralized training, decentralized execution (CTDE) framework, which can improve the training stability of MADDPG and improve the training efficiency of the agent.
[0164] Based on the above analysis, the present invention defines the base station agent as Ξ0 and the RIS agent as For each agent All by an actor network A target actor network A critic network and a target-critic network Composition, of which and Represent the corresponding network learnable parameters, s i and a m , Representing intelligent agents Ξ i The state and action of each agent, S = {s0,s1,…,s J} represents the global state of the system.
[0165] In addition, this embodiment defines an experience replay buffer To store the experience generated by each agent during the training process. Since the goal of the optimization problem is to maximize the total throughput of the system, each agent cooperates to design its own beamforming to maximize a common reward function r, that is, the total throughput of the system.
[0166] Algorithm flow: At the nth time step, each agent Ξ i First, observe your own state The agent's actor network then uses the observed state to determine the action to be taken by the agent, namely:
[0167]
[0168] After each agent obtains the action it wants to take according to the above formula, the entire system takes the global action. Get the current reward r (n) And the global state S after the action is implemented (n+1) . Subsequently, each agent’s experience {S (n) ,S (n+1) ,a (n) ,r (n)} will be stored in the experience playback buffer In the experience replay buffer, information is provided for subsequent training. After there are enough samples in , a mini-batch of samples will be randomly sampled at each training time step for subsequent policy evaluation and necessary updates.
[0169] First, when we start updating the network parameters, the target Q function is updated as follows:
[0170]
[0171] where γ represents the discount factor, Represents the action obtained by the target actor network. Based on the above formula, the critic network is updated by minimizing the loss function, which is defined as follows:
[0172]
[0173] After completing the update of each agent's critic network, each agent's actor network also needs to be updated. The actor network is updated using the sampled policy gradient method, with the goal of maximizing the expected reward, as shown below:
[0174]
[0175] Then, at the end of each time step, this embodiment uses a soft update method to update the target actor network and the target critic network as follows:
[0176]
[0177]
[0178] in and Represent the soft update learning rates of the target actor network and the target critic network respectively.
[0179] Definition of agent, state, action, and reward function:
[0180] Base station active beamforming agent Ξ0: This agent is responsible for optimizing the base station's active beamforming matrix W. At the nth time step, the observed state of Ξ0 is determined by the RIS obtained in the first stage and the user's pairing strategy μ * And the channel state information of the system is as follows:
[0181]
[0182] The action determined by the actor network of agent Ξ0 is composed of the perception beamforming matrix and the communication beamforming matrix, namely:
[0183]
[0184] RIS passive beamforming agent At the nth time step, the state observed by each RIS agent is:
[0185]
[0186] The action obtained by applying the actor network is composed of the phase shift matrix corresponding to each RIS, namely:
[0187]
[0188] Reward function r (n) : Since the goal of each agent in the training process is the same, that is, to maximize the total throughput of the system, the reward function set in the present invention is the total throughput of the system, that is:
[0189] r (n) =R sum . (45)
[0190] (6) Simulation experiment results
[0191] In this example, simulation analysis verified the effectiveness of the proposed two-stage resource optimization method based on machine learning in improving total system throughput. The simulation was implemented using the PyTorch platform in Python. The superiority of the proposed optimization scheme was further demonstrated by comparing it with other comparable solutions.
[0192] The relevant parameters of the simulation experiment are set as follows: the mmWave simulation channel is set with one LoS path and five NLoS paths. The maximum transmission power P max Set to 30dBm. Noise power is set to σ 2 =-90dBm. The number of base station antennas is M=16, and the number of each RIS reflection unit is N=16. In addition, the carrier frequency involved in this embodiment is set to 30GHz. The number of layers L of RE-GAT in the graph neural network is set to 3. The number of heads H of the attention layer of the decoder in RE-GAT is set to 8. The number of training rounds of MADDPG is set to 100, and the training time step of each round is set to 50000. The capacity of the experience replay buffer is set to 10000, the mini-batch sampling size is set to 128, and the discount factor γ is set to 0.99. The learning rate of the actor network for updating the intelligent agent is set to 1e -4 , the learning rate of the updated critic network is set to 1e -3 , the learning rate of the soft-update target actor network and the target critic network is set to 1e -4 .
[0193] Comparison plan settings:
[0194] Solution 1: In the first stage, the graph neural network method proposed in this invention is used to complete the pairing between RIS and users. In the second stage, the traditional single-agent DDPG method is used to optimize the base station active beamforming and RIS passive beamforming.
[0195] Solution 2: In the first stage, the graph neural network method proposed in this invention is used to complete the pairing between RIS and users. In the second stage, the traditional BCD method is used to iteratively optimize the base station active beamforming and RIS passive beamforming.
[0196] Solution 3: In the first stage, the classic pairing algorithm Gale Shapley algorithm is used to pair RIS and users. In the second stage, the MADDPG method proposed in this invention is used to optimize the base station active beamforming and RIS passive beamforming.
[0197] Solution 4: In the first stage, a random pairing method is used to randomly generate pairings between RIS and users. In the second stage, the MADDPG method proposed in this embodiment is used to optimize the base station active beamforming and the RIS passive beamforming.
[0198] Figure 2 This paper demonstrates the impact of changes in the number of base station antennas on the total system throughput of an ISAC system in the mmWave band, using the machine learning-based two-stage resource optimization method proposed in this embodiment. Four different comparison schemes were set up for comparison with the proposed scheme. Simulation results show that the proposed scheme consistently outperforms the other comparison schemes in terms of optimization performance, and this superiority becomes increasingly significant as the number of base station antennas increases. This is because the MADDPG method proposed in this embodiment can flexibly control the states and actions of individual agents when optimizing the complex beamforming problem of base stations and multiple RISs. Compared to traditional DDPG and BCD methods, the proposed MADDPG method, due to its centralized learning and decentralized execution framework, has more stable training performance and faster convergence. Furthermore, the proposed GNN method also demonstrates superior performance in optimizing the pairing between RISs and users, compared to the traditional Gail Shapley algorithm and the baseline random pairing method. This is primarily due to the large-scale training of graph neural networks, which effectively extracts key pairing-related information from the system, better serving the pairing process and thus achieving superior pairing results. The simulation results further illustrate that the optimized pairing of RIS and users has a significant impact on the total system throughput.
[0199] Figure 3 It shows that in the ISAC system under the mmWave band under the consideration of distributed RIS assistance, the total system throughput is affected by the maximum transmit power of the base station. Figure 2The trend of the curve in the figure shows that as the maximum transmit power of the base station increases, the total system throughput of all schemes shows an increasing trend. Among them, the proposed scheme still achieves the highest total system throughput over all other compared schemes. In addition, when the maximum transmit power of the base station increases to a relatively large level, the performance gap between the proposed scheme and Schemes 3 and 4 gradually increases. This further illustrates that the proposed graph neural network method is significantly better than the Gale Shapley method and random pairing method in optimizing the pairing between RIS and users. It also proves that a better pairing scheme is helpful for the subsequent design of base station active beamforming and RIS passive beamforming based on the MADDPG algorithm.
[0200] Figure 4 It shows that in the ISAC system under the distributed RIS-assisted mmWave frequency band, the total system throughput is affected by the number of sensing targets and the minimum beam gain of the sensing targets. As can be seen from the figure, as the number of sensing targets increases, the total system throughput of all schemes decreases. This is because the increase in the number of sensing targets requires the system to invest more resources to meet the minimum sensing performance of each sensing target. Among all schemes, the proposed scheme can achieve the optimal total system throughput under different numbers of sensing targets and different minimum sensing beam gains, which further illustrates the superiority of the proposed scheme. As the minimum sensing beam gain increases, the total system throughput of all schemes shows a downward trend. The main reason is that the sensing performance requirements of the sensing targets are improved, so the system needs to invest more resources in sensing, which reduces the communication performance. The proposed scheme and the comparison scheme always maintain a relatively significant performance gap, which further illustrates that the two-stage resource optimization scheme proposed in the present invention can better balance the sensing and communication in the ISAC system, thereby obtaining better overall system performance.
[0201] Figure 5 The sensing performance of an ISAC system in the considered distributed RIS-assisted mmWave band is demonstrated. During the simulation, three sensing targets were set, located at angles of -45°, 0°, and 45°, centered on a base station. The sensing performance is described by the beam gain at each angle. The figure shows that compared to Solution 1, the proposed solution achieves higher beam gain at the angle of the sensing target and lower beam gain at angles other than the sensing target. This demonstrates that the proposed solution achieves better sensing performance and further illustrates that the proposed solution in this embodiment can better balance sensing and communication, resulting in superior system performance.
[0202] (7) Beneficial effects
[0203] At present, most of the research on RIS-assisted ISAC systems in the mmWave frequency band focuses on the application scenario research of the ISAC system in the mmWave frequency band assisted by a single RIS, and there is relatively little research on distributed RIS. In the research on system performance optimization, existing research mainly focuses on optimizing the active beamforming of base stations and passive beamforming of RIS, and rarely optimizes the pairing between RIS and users. In addition, in the research on ISAC system performance optimization, existing research mainly adopts traditional optimization methods, such as the alternating iteration method and the block descent method, and few studies consider using more advanced machine learning-based methods for optimization. Considering that the existing research on the joint optimization of RIS and user pairing and active and passive beamforming of the ISAC system in the mmWave frequency band assisted by distributed RIS is not in-depth enough, the present invention considers a distributed RIS-assisted mmWave frequency band ISAC system, and performs two-stage optimization based on machine learning on the optimization problems involving RIS and user pairing, base station active beamforming and RIS passive beamforming. The present invention mainly has the following advantages:
[0204] 1) This invention considers a system framework that combines distributed RIS, mmWave, NOMA, and ISAC technologies. Introducing RIS into the ISAC system in the mmWave band can address the issues of mmWave signal susceptibility to obstruction and insufficient coverage. Introducing NOMA technology into the RIS-assisted mmWave system can effectively improve spectrum resource utilization. Furthermore, the application of ISAC technology enables the system to simultaneously implement both perception and communication functions, which is highly consistent with the growing communication and perception needs in real life.
[0205] 2) In practical applications, since each RIS can only serve a limited number of users, introducing a distributed RIS can effectively expand the coverage of mmWave signals, providing high-quality service to more users and increasing overall system throughput. However, the introduction of a distributed RIS introduces a new RIS-user pairing problem, which this paper investigates.
[0206] 3) Based on the proposed system, to maximize the system's total throughput while ensuring the minimum perceptual beam gain for the sensing target, this paper proposes a two-stage, multi-dimensional resource joint optimization method based on machine learning to address RIS-user pairing, base station active beamforming, and RIS passive beamforming. Simulation results demonstrate that the proposed scheme significantly improves the system's total throughput compared to other comparable schemes, while also achieving a better balance between communication and perceptual performance.
[0207] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0208] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made based on the essence of the present invention are intended to be covered by the scope of protection of the present invention.
Claims
1. A distributed RIS-assisted ISAC system, characterized in that: It includes a base station, multiple distributed RIS and multiple user terminals; Distributed RIS assists user terminals in mmWave communications. Each user terminal can be served by only one RIS, and each RIS can serve user terminals in one NOMA group. Within each NOMA group served by each RIS, target information for each user is decoded using SIC technology. Targets are directly detected by the base station, and the perceived beam gain is used as an indicator of perception performance. The RIS and user pairing, base station active beamforming and RIS passive beamforming are jointly optimized to maximize the total throughput of the system under the premise of minimum beam gain for each sensing target.
2. A distributed RIS-assisted ISAC system according to claim 1, characterized in that: The joint optimization process consists of two stages: In the first stage, a graph neural network-based method is used to optimize the pairing between RIS and users. Attention mechanism and residual connection technology are used to extract key graph information to obtain a better pairing strategy. In the second stage, based on the pairing strategy obtained in the first stage, the active beamforming of the base station and the passive beamforming of the RIS are jointly optimized using the multi-agent deep reinforcement learning (MADDPG) method. The base station and each RIS are constructed as an independent agent, and each agent consists of an actor network, a critic network, a target actor network, and a target critic network.
3. A resource joint optimization method, applied to the distributed RIS-assisted ISAC system according to claim 1 or 2, characterized in that: include: Construct a distributed RIS-assisted ISAC system model in the mmWave frequency band; Construct mmWave channel modeling, ISAC system signal model, and communication and perception model based on the system model; Determine the optimization problem based on the constructed model; Based on the optimization problem, RIS and user pairing, base station active beamforming and RIS passive beamforming are jointly optimized to maximize the total throughput of the system under the premise of minimum beam gain for each sensing target.
4. A resource joint optimization method according to claim 3, characterized in that: The ISAC system model in the mmWave frequency band assisted by the distributed RIS includes: K single-antenna communication users complete the communication with the base station with the assistance of J RIS; each RIS with N reflective elements is modeled as a uniform planar array, and the adjacent reflective elements of the RIS are separated by d x and d y The phase shift matrix of the jth RIS is set to Θ j , and considering the constraints of actual hardware, discrete phase shift is used to represent the phase shift matrix, that is, the phase shift of each reflection unit of RIS can only take values in a limited number of phase shifts; The base station is equipped with M antennas forming a uniform linear array. To achieve interawareness integration, T sensing targets are sensed directly through the base station. The sets of user terminals, RIS and sensing targets are defined as and 5. A resource joint optimization method according to claim 3, characterized in that: The mmWave channel modeling is constructed as follows: The mmWave channel is characterized by a line-of-sight link and multiple non-line-of-sight links; the channel between the base station and the j-th RIS is represented by G 0,j Denote by , and the channel between the jth RIS and the kth user terminal is represented by h j,k Represented, and all mmWave channels are modeled using the Saleh-Valenzuela channel model; In addition, considering the passivity of RIS, it becomes extremely difficult to obtain channel state information (CSI), which introduces CSI errors and only imperfect CSI can be obtained.
6. A resource joint optimization method according to claim 3, characterized in that: The ISAC system signal model is constructed as follows: In order to simultaneously serve K communication users and T sensing targets, the base station needs to transmit superimposed NOMA and sensing signals, which are expressed as follows: x=W c s c +W r s r =Ws Where W c represents the active communication beamforming matrix of the base station, W r represents the active sensing beamforming matrix of the base station; s c Indicates that s r represents the transmission signals of K user terminals and T sensing targets; W,s represent the total active beamforming matrix and the total transmission signal respectively; Assume that the kth user terminal matches the μth k RIS, then the signal received by the kth user terminal is expressed as: Where, Indicates μth k The channel from the RIS to the kth user, H represents the conjugate transpose operation of the matrix, w k,c represents the communication beamforming vector of the kth user, s k,c represents the communication signal that the kth user expects to receive, w d,c represents the communication beamforming vector of the dth user, s d,c represents the communication signal that the dth user expects to receive, w t,r represents the sensing beamforming vector of the t-th target, s t,r represents the communication signal that the tth target expects to receive, n k represents Gaussian white noise.
7. A resource joint optimization method according to claim 3, characterized in that: The communication and perception model is constructed as follows: The effective channel vector of the kth user terminal is defined as Assume that the communication users are sorted in descending order of channel gain; the kth user terminal first detects and removes the interference of the weak user of the same RIS service, while considering the signal of the strong user of the same RIS service and the signals of all other RIS service users as interference; then the received signal interference and noise ratio of the kth user terminal is expressed as: Where, represents the noise power; In order to make SIC decoding proceed smoothly, the signal of the kth user terminal needs to be decoded at the weak user n of the same RIS service, and the corresponding SINR is: Then, based on the SIC principle, the achievable rate of user k is: Among them, R n→k and R k→k denote the data rate obtained by weak user n decoding user k and the data rate obtained by strong user decoding its own target signal, respectively. The total throughput of the system is obtained by summing the rates of all user terminals: In terms of perception, the active beamforming matrix W of the base station is used to perceive the target; the base station is defined as receiving the angle θ from the perceived target. t The beam gain is: Where, α L (·) represents the array response of a uniform linear array.
8. A resource joint optimization method according to claim 3, characterized in that: The expression of the optimization problem is: In the optimization problem (P0), μ and Θ represent the set of pairing variables between RIS and users and the set of phase shift matrices of RIS, respectively; P max Indicates the maximum transmission power of the base station, R k represents the achievable rate of the kth user, R min represents the minimum achievable rate constraint for each user, Indicates that the base station receives the signal from the sensing target with an angle of θ t The beam gain, Indicates the minimum beam gain received by the base station for each sensing target, φ j,n represents the phase shift of the nth reflector unit of the jth RIS, represents the set of possible values of the discrete phase shift of the RIS reflection unit; constraint C1 limits each user to only one RIS; C2 represents the maximum transmit power limit of the system; C3 guarantees the minimum rate of each communicating user; C4 ensures the minimum beam gain of each sensing target received by the base station; C5 indicates that the considered RIS phase shift is a discrete phase shift.
9. A resource joint optimization method according to claim 3, characterized in that: The joint optimization of RIS and user pairing, base station active beamforming, and RIS passive beamforming includes: In the first stage, a graph neural network-based method is used to optimize the pairing between RIS and users. Attention mechanism and residual connection technology are used to extract key graph information to obtain a better pairing strategy. In the second stage, based on the pairing strategy obtained in the first stage, the active beamforming of the base station and the passive beamforming of the RIS are jointly optimized using the multi-agent deep reinforcement learning (MADDPG) method. The base station and each RIS are constructed as an independent agent, and each agent consists of an actor network, a critic network, a target actor network, and a target critic network.
10. A resource joint optimization method according to claim 9, characterized in that: During the training process, each agent inputs the observed state into the actor network to obtain the action to be implemented. Then the environment evaluates the action of each agent and updates the critic network and actor network. Finally, the target actor network and target critic network are soft-updated.
Citation Information
Patent Citations
Multi-RIS-assisted MIMO-NOMA system optimization method based on power minimization
CN116471611A
Communication method and system based on difunctional intelligent metasurface
CN116867066A
Beam forming design method for STAR-RIS-assisted multi-user MISO URLLC system
CN117375684A
Internet of vehicles power distribution and user scheduling method, system and device assisted by intelligent reflecting surface relay, and medium
CN117460034A
Resource optimization distribution system and method of wireless energy transfer communication perception integrated network
CN119183130A