Hybrid RIS-assisted ISAC resource optimization method based on DRL

Through the hybrid RIS assisted ISAC resource optimization method based on DRL, the multivariate coupling non-convex problem and sensitive information leakage in the RIS assisted ISAC system are solved, and the confidentiality rate and resource optimization are maximized in complex environments, improving the security and adaptability of the system.

CN120357930AActive Publication Date: 2025-07-22NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510840209.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

The prior art has high complexity of multivariable coupling, limitations of relying on ideal assumptions, and severe problems of sensitive information leakage in RIS-assisted ISAC systems, especially in the downstream fading of 6G high-frequency communication.

Method used

Using a hybrid RIS-assisted ISAC resource optimization method based on DRL, we will improve SAC reinforcement learning by building an ISAC system model, combining deep-directed intrinsic motivation exploration algorithms, and jointly optimize transmit beamforming, AN interference and RIS phase shift, taking into account RIS hardware damage and channel uncertainty to maximize confidentiality.

Benefits of technology

It breaks through the dependence on perfect channel state information and ideal reflection model, improves the confidentiality rate of the system and adaptability in complex environments, effectively suppresses eavesdropping behavior, and optimizes resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120357930A_ABST
    Figure CN120357930A_ABST
Patent Text Reader

Abstract

The invention discloses a DRL-based hybrid RIS-assisted ISAC resource optimization method, and belongs to the technical field of wireless communication. The method comprises the following steps: constructing an ISAC system model; establishing optimization problems in a real environment and a mismatched environment; sAC reinforcement learning is improved by fusing a depth orientation internal motivation exploration algorithm; and solving an optimization problem by using the improved SAC reinforcement learning to realize resource optimization. According to the method, the hybrid RIS and AN interference is jointly designed, and the secrecy rate is maximized under the constraint of RIS hardware damage and channel uncertainty; meanwhile, a phase correlation reflection amplitude model is introduced to adapt to the RIS hardware damage condition, transmitted beam forming, AN interference and RIS phase shifting are jointly optimized through DRL, the dependence of a traditional algorithm on perfect channel state information (CSI) and an ideal reflection model is broken through, and ISAC system resource optimization is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to a DRL-based hybrid RIS-assisted ISAC resource optimization method. Background Art

[0002] With the explosive growth of Internet of Everything applications, spectrum resources are becoming increasingly scarce. Due to its advantages in sharing hardware platforms, improving spectrum efficiency, and reducing interference, ISAC technology has become a key direction for solving the spectrum congestion problem. ISAC realizes the coordination of communication and radar sensing by designing dual-functional waveforms. However, in the scenario of channel degradation, it is difficult for a single active beamforming to balance performance and security requirements. As a green programmable reflection technology, RIS provides additional degrees of freedom for the ISAC system by reconstructing the wireless propagation environment, significantly improving multi-user interference suppression and radar sensing performance. However, the security of the ISAC system faces severe challenges: when the sensing target is a potential eavesdropper, the power concentration strategy of traditional radars may lead to the leakage of sensitive information. Therefore, physical layer security and artificial noise technologies have been introduced to enhance security, but existing research has deficiencies in the security optimization of RIS-assisted ISAC systems.

[0003] Traditional convex optimization-based methods (such as Taylor expansion, semidefinite relaxation, and alternating optimization) face limitations of high complexity and dependence on ideal assumptions when solving non-convex problems with multivariable coupling, especially in the case of RIS hardware impairments (such as phase-dependent reflection amplitude) and dynamic channel environments. Deep reinforcement learning provides a new paradigm for resource optimization of time-varying wireless systems through the interactive learning between the agent and the environment. In the RIS-assisted ISAC scenario, DRL can jointly optimize transmit beamforming, AN interference, and RIS phase shift, breaking through the dependence of traditional algorithms on perfect channel state information CSI and ideal reflection models.

[0004] For example, the DRL framework based on soft actor-critic avoids local optima by maximizing the policy entropy, and significantly improves the system secrecy rate under the guarantee of the minimum radar signal-to-interference-plus-noise ratio and power constraints. In addition, for the performance loss caused by the phase-dependent amplitude of RIS, DRL can provide a data-driven solution for complex system modeling by adaptively learning the optimal strategy under non-ideal hardware characteristics. Recent research has further extended the application of DRL in non-ideal conditions (such as low-resolution RIS, multi-user MIMO systems), and achieved performance close to ideal assumptions by improving the network architecture and training mechanism, opening up a new path for resource optimization of hybrid RIS-assisted ISAC systems.

[0005] A physical layer security optimization method for a RIS-assisted ISAC system disclosed in China (application number: 202410624177.2). This method takes the detection probability of eavesdroppers as the system sensing performance index and the total secrecy rate of users as the system communication performance index. By jointly optimizing the ISAC-BS beamforming and RIS phase shift, the secrecy rate of all users is maximized under the condition of meeting a certain detection probability. During the solution process, the mathematical model based on the user secrecy rate optimization problem is first decoupled into three sub-problems, and the three decision variables of the base station receiving beam, RIS phase, and base station transmitting beam are alternately optimized. Then, methods such as the Rayleigh quotient problem, first-order Taylor expansion, semidefinite relaxation, and successive convex approximation are used to transform the sub-problems into deterministic convex optimization models for solution. Through simulation, it can be seen that increasing the number of RIS elements and the number of base station antennas can improve the system secrecy rate, and this scheme performs better than the schemes without RIS and with random RIS phase in terms of improving the system secrecy rate and reducing power loss.

[0006] A joint beamforming design method for a hybrid RIS-assisted millimeter-wave ISAC system disclosed in China (application number: 202410333259.1). This method constructs an optimization function by considering the sensing performance of the ISAC system, and jointly optimizes the analog beamformer, digital beamformer, and HRIS reflection coefficient matrix to maximize the illumination power in the worst case while ensuring communication quality. When solving, the non-convex problem is first decoupled into three sub-problems using the surrogate optimization technique, and the DFBs and HRIS beamforming optimization sub-problems are transformed into convex problems for solution using quadratic constraint quadratic programming (QCQP) and semidefinite programming relaxation (SDR) techniques. For the DFBs hybrid precoding design sub-problem, the MO-AltMin algorithm based on minimizing the Euclidean distance is used for solution. The simulation results show that the proposed algorithm is effective, and the fully connected hybrid precoding structure can obtain near-optimal performance, and the hybrid RIS can achieve a good balance between system performance and hardware cost.

[0007] The above patent applications have the following deficiencies: 1. The constructed system model involves numerous parameters and complex channel models. During the process of decoupling and solving the mathematical model, the model complexity increases significantly, and the demand for computing resources also increases substantially.

[0008] 2. It depends on the ideal assumption of channel state information and RIS phase amplitude design, which is not applicable to the actual situation where channel state information cannot be obtained or the RIS hardware is damaged. Only passive intelligent reflectors are considered, and there is "multiplicative fading", especially in 6G high-frequency communication, the "multiplicative fading" effect is more significant.

[0009] Therefore, how to solve the non-convex problem of multi-variable coupling, which faces the limitations of high complexity, dependence on ideal assumptions, and the serious problem of sensitive information leakage, is the technical problem that the present invention aims to solve. Summary of the Invention

[0010] The purpose of the present invention is to provide a DRL-based hybrid RIS-assisted ISAC resource optimization method to solve the problems raised in the above background technology.

[0011] The purpose of the present invention is achieved as follows: A DRL-based hybrid RIS-assisted ISAC resource optimization method, characterized in that: the method includes the following steps:

[0012] Step S1: Construct an ISAC system model;

[0013] The ISAC system model includes a hybrid RIS, a base station BS, users, and a sensing target. The base station BS serves the users with the help of the hybrid RIS and simultaneously detects the position of the sensing target by using the sensing target.

[0014] Step S2: Establish optimization problems in the real environment and the mismatched environment;

[0015] Step S3: Improve the SAC reinforcement learning by integrating the deep directional intrinsic motivation exploration algorithm;

[0016] Step S4: Solve the optimization problem by using the improved SAC reinforcement learning to achieve resource optimization.

[0017] Preferably, in step S1, constructing the ISAC system model is specifically as follows:

[0018] Step S1-1: Set the hybrid RIS;

[0019] The hybrid RIS consists of components, including active components and passive components. Among them, the active components need to additionally tune the amplitude of the incident signal compared with the passive components;

[0020] Let represent the set of active components, ; Let represent the relay / reflection coefficient used on the th component:

[0021] ;

[0022] Among them, represents the phase shift, ; When , ; is the th component, where is the Euler's formula;

[0023] Based on the following three diagonal matrices: , and , where:

[0024] ;

[0025] ;

[0026] where is the coefficient matrix containing only passive components, is the coefficient matrix containing only active components, is the coefficient matrix of the entire hybrid RIS; ; and are the passive reflection coefficient containing only passive components and the active relay coefficient of active components, respectively;

[0027] Step S1-2: Construct the transmission signal of the base station BS and the received signal of the user;

[0028] Using precoding design, the transmission signal of the base station BS is expressed as:

[0029] ;

[0030] where represents the desired signal of user , ; represents the AN generated by the base station BS, ; and are the precoding vectors of the base station BS for and , respectively; is the number of users;

[0031] Adopting the AN strategy, by utilizing more DoF to enhance the sensing ability and at the same time counteracting the eavesdropping from the detected sensing target, the received signal at user is expressed as:

[0032] ;

[0033] where and represent the hybrid RIS and user and the channel vector between the hybrid RIS and the base station BS; Transmit signals for the base station BS; Is the conjugate transpose of the matrix;

[0034] , , Is the channel gain per unit distance, Indicates the hybrid RIS and the user The distance between; Indicates the distance between the hybrid RIS and the base station BS; Indicates the hybrid RIS and the user The Rayleigh fading vector between; Indicates the Rayleigh fading vector between the hybrid RIS and the base station BS;

[0035] , Is the additive noise caused by the active elements between the base station BS and the hybrid RIS ; Is the additive white Gaussian noise at the base station BS to the user ; , Indicates the average noise variance of the additive noise caused by the active elements between the base station BS and the hybrid RIS ; Indicates the average noise variance of the additive white Gaussian noise at the base station BS to the user , the noise power spectral density and noise figure at the hybrid RIS and the user Are the same, that is .

[0036] Preferably, in step S2, an optimization problem in the real environment and the mismatch environment is established, specifically:

[0037] Step S2-1: Obtain the phase amplitude model in the reflection loss of the hybrid RIS in the real environment and the mismatch environment;

[0038] In the real environment, the hybrid RIS follows the phase-dependent amplitude model. When , , ; Get:

[0039] ;

[0040] Among them, , Indicates the inherent phase shift of the component, caused by the non-ideal characteristics of the hardware circuit, ; Indicates the dependence intensity of the reflection amplitude on the phase, , the larger its value, the more significant the impact of phase tuning on reflection loss;

[0041] BS knows the independent cascaded channel of each user, denoted as:

[0042] ;

[0043] Therefore, in the real environment, the received signal at user is expressed as:

[0044] ;

[0045] where represents the column vector formed by extracting and arranging the diagonal elements of in sequence; is the transmission signal of base station BS, is the noise at user , is the inverted matrix;

[0046] The hybrid RIS reflection in the mismatch environment is assumed to be lossless and is expressed as follows:

[0047] ;

[0048] The agent can only access an imperfect estimate of the cascaded channel, that is:

[0049] ;

[0050] where represents the channel estimation error matrix of the cascaded channel of each user, with independent and identically distributed terms ;

[0051] Step S2-2: Calculate the achievable rate of the user in the real environment and the mismatch environment;

[0052] The SINR of the user in the real environment is calculated as:

[0053] ;

[0054] where is the precoding vector of the desired signal of base station BS for the current user i, is the precoding vector of the desired signals of base station BS for other users, is the precoding vector of base station BS for AN, is the cascaded channel of the user, represents the one formed by The transpose matrix of the column vector formed by sequentially extracting and arranging the diagonal elements of is the average noise variance, is the coefficient matrix containing only active components; is the number of users;

[0055] User The achievable rate of

[0056] ;

[0057] In a mismatched environment, for user the SINR is calculated as:

[0058] ;

[0059] where, is the hybrid RIS reflection coefficient matrix in the mismatched environment, is the cascaded channel with imperfect estimation in the mismatched environment, is from the first active RIS element coefficients stripped from

[0060] User the achievable rate of

[0061] ;

[0062] Step S2-3: Calculate the radar signal SINR of the sensing target in the real environment and the mismatched environment;

[0063] Step S2-4: Calculate the secure transmission rates of the users in the real environment and the mismatched environment;

[0064] Step S2-5: Obtain the optimization problem.

[0065] Preferably, in step S2-3, calculating the radar signal SINR of the sensing target specifically includes:

[0066] The radar signal at the sensing target is expressed as:

[0067] ;

[0068] where, represents the channel vector between the BS and the sensing target, and ; represents the Gaussian white noise at the sensing target; is the distance between the base station BS and the sensing target; is the antenna array steering vector, expressed as:

[0069] ;

[0070] Among them, is the azimuth angle of the perceived target, and respectively represent the wavelength and the antenna spacing; the echo SINR related to detecting the perceived target is expressed as:

[0071] ;

[0072] Among them, is the conjugate transpose of the antenna array steering vector, is the number of users, is the channel gain per unit distance.

[0073] Preferably, in step S2-4, calculating the secure transmission rate of users in the real environment and the mismatched environment is specifically:

[0074] The secure transmission rate of user in the real environment is expressed as:

[0075] ;

[0076] The secure transmission rate of user in the mismatched environment is expressed as:

[0077] ;

[0078] Among them, is the achievable rate of user , is the eavesdropping rate of the perceived target on user , is the secure transmission rate of user in the real environment, is the secure transmission rate of user in the mismatched environment, .

[0079] Preferably, in step S2-5, obtaining the optimization problem is specifically:

[0080] Collaboratively formulating the beamforming matrix and the reflection coefficient to achieve the secrecy sum rate, the optimization problem in the real environment:

[0081] ;

[0082] ;

[0083] ;

[0084] ;

[0085] ;

[0086] Among them, is the secrecy sum rate in the real environment; is the SINR of the radar echo signal; represents the SINR threshold of the echo signal, represents the transmit power of the BS, is the power budget of the active element, is the transmit power of the hybrid RIS active element, expressed as:

[0087] ;

[0088] Among them, represents the transmit power of the user-expected signal; is the identity matrix of K active elements; is the channel vector between the hybrid RIS and the base station BS; is the conjugate transpose of the channel vector between the hybrid RIS and the base station BS;

[0089] The optimization problem in the mismatch environment is:

[0090] ;

[0091] ;

[0092] ;

[0093] ;

[0094] ;

[0095] Among them, is the secrecy sum rate in the mismatch environment.

[0096] Preferably, in step S3, the fusion depth-directed intrinsic motivation exploration algorithm improves the SAC reinforcement learning, specifically:

[0097] Step S3-1: Modify the reward function for training the agent:

[0098] ;

[0099] Among them, is the instantaneous reward calculated by the environment in the current state, is the average reward up to the current state;

[0100] Step S3-2: Add an explorer network to predict the coefficients of the hybrid RIS , where :

[0101] ;

[0102] where, when , , in the formula, the number of 1s refers to the number of elements of the transmit beamforming and the AN beamforming included in the action vector; each has 2 terms and can measure the real part and the imaginary part of simultaneously;

[0103] Step S3-3: The prediction of the explorer network is used to perturb the action:

[0104] ;

[0105] where is the Hadamard product, and the hyperparameter constrains the explorer network not to perform harmful perturbations on the actions selected by the actor;

[0106] Step S3-4: Correct the losses of the Q-network and the actor network as follows:

[0107] ;

[0108] ;

[0109] ;

[0110] ;

[0111] where ; the explorer network is optimized to maximize the sum of the absolute values of the TD errors:

[0112] ;

[0113] ;

[0114] Step S3-5: Update the deterministic exploration network through the deterministic policy gradient algorithm:

[0115] ;

[0116] ;

[0117] where is the calculation of the mathematical expectation, is the temporal difference error after action perturbation, is the gradient operator for the perturbation vector ; is the environmental state observed by the agent at time step t, is the original action selected by the agent at time step t, is the perturbation vector generated by the explorer network, is the explorer network with parameters ; is the gradient operator for the parameter ; is the parameter of the explorer network ; is the learning rate of the explorer network, The perception objective function of the explorer network.

[0118] Compared with the prior art, the present invention has the following improvements and advantages:

[0119] 1. By jointly designing the hybrid RIS and AN interference, the secrecy rate is maximized under the constraints of considering RIS hardware impairments and channel uncertainties; meanwhile, a phase-related reflection amplitude model is introduced to adapt to the RIS hardware impairment situation, and DRL is used to jointly optimize the transmit beamforming, AN interference, and RIS phase shift, breaking through the dependence of traditional algorithms on perfect channel state information CSI and ideal reflection models, and realizing the resource optimization of the ISAC system.

[0120] 2. Through the SAC algorithm and the deep directed intrinsic motivation exploration algorithm, the balance between exploration and exploitation efficiency is achieved, and better adaptability is demonstrated in the complex design of transmit beamforming, AN, and hybrid RIS coefficient matrices. BRIEF DESCRIPTION OF THE DRAWINGS

[0121] Figure 1 is a schematic flow diagram of the method of the present invention.

[0122] Figure 2 is a schematic diagram of the ISAC system model structure.

[0123] Figure 3 is a schematic diagram of the instantaneous secrecy sum rate in the case of

[0124] Figure 4 is a schematic diagram of the instantaneous secrecy sum rate in the case of

[0125] Figure 5 is a schematic diagram of the present transmit beam direction. DETAILED DESCRIPTION OF THE INVENTION

[0126] The following further outlines the present invention with reference to the accompanying drawings.

[0127] As shown Figure 1 in the figure, a hybrid RIS-assisted ISAC resource optimization method based on DRL, the method includes the following steps:

[0128] Step S1: Construct an ISAC system model;

[0129] As shown Figure 2 in the figure, the ISAC system model includes a hybrid RIS, a base station BS, users, and a sensing target. The base station BS with antennas simultaneously serves single-antenna users with the help of hybrid RISs, and at the same time detects the location of the sensing target; The hybrid RISs are strategically placed near the users to enhance downlink multi-user communication, and it is assumed that the hybrid RISs are located far from the sensing target, and the radar echo signals reflected by the hybrid RISs are weak and can be ignored; The sensing target is considered as a potential eavesdropper Eve who may eavesdrop on the downlink transmission of the users.

[0130] Constructing the ISAC system model specifically includes:

[0131] Step S1-1: Set the hybrid RIS;

[0132] The hybrid RIS consists of components, including active components and passive components. Among them, the active components need to additionally tune the amplitude of the incident signal compared with the passive components;

[0133] Let represent the set of active components, ; Let represent the relay / reflection coefficient used on the th component:

[0134] ;

[0135] Among them, represents the phase shift, ; When , ; is the th component, is a complex number, is Euler's formula;

[0136] Based on three diagonal matrices: , and , where:

[0137] ;

[0138] ;

[0139] where, is the coefficient matrix containing only passive components, is the coefficient matrix containing only active components, is the coefficient matrix of the entire hybrid RIS; ; and are the passive reflection coefficient containing only passive components and the active relay coefficient of the active components, respectively;

[0140] Step S1-2: Construct the transmission signal of the base station BS and the received signal of the user;

[0141] Using precoding design, the transmission signal of the base station BS is expressed as:

[0142] ;

[0143] where, represents the desired signal of user , ; represents the AN generated by the base station BS, ; and are the precoding vectors of the base station BS for and , respectively; is the number of users;

[0144] Adopting the strategy of AN, by exploiting more DoF to enhance the sensing ability and at the same time counteracting the eavesdropping from the detected sensing target, the received signal at user is expressed as:

[0145] ;

[0146] where, and represent the channel vectors between the hybrid RIS and user and between the hybrid RIS and the base station BS, respectively; is the transmission signal of the base station BS; is the conjugate transpose of the matrix;

[0147] , , is the channel gain at unit distance, represents the hybrid RIS and user The distance between represents the distance between the hybrid RIS and the base station BS; represents the Rayleigh fading vector between the hybrid RIS and the user and represents the Rayleigh fading vector between the hybrid RIS and the base station BS;

[0148] , is the additive noise caused by active elements between the base station BS and the hybrid RIS, is the additive white Gaussian noise at the user ; , represents the average noise variance of the additive noise caused by active elements between the base station BS and the hybrid RIS, represents the average noise variance of the additive white Gaussian noise at the user , and the noise power spectral density and noise figure at the hybrid RIS and the user are the same, that is .

[0149] Step S2: Establish the optimization problems in the real environment and the mismatched environment, specifically:

[0150] Step S2-1: Obtain the phase-amplitude model in the reflection loss of the hybrid RIS in the real environment and the mismatched environment;

[0151] In the real environment, the hybrid RIS follows the phase-dependent amplitude model. When , , ; we get:

[0152] ;

[0153] where , represents the inherent phase offset of the element, which is caused by the non-ideal characteristics of the hardware circuit, ; represents the dependence strength of the reflection amplitude on the phase, and the larger its value, the more significant the impact of phase tuning on the reflection loss;

[0154] The BS knows the independent cascaded channel of each user, denoted as:

[0155] ;

[0156] Therefore, in the real environment, the received signal at the user is expressed as:

[0157] ;

[0158] Among them, represents the column vector formed by sequentially extracting and arranging the diagonal elements of ; is the signal transmitted by the base station BS, is the user at the noise, is the inverted matrix;

[0159] In the mismatched environment, the hybrid RIS reflection is assumed to be lossless, and the manifestation is as follows:

[0160] ;

[0161] The agent can only access an imperfect estimate of the cascaded channel, that is:

[0162] ;

[0163] Among them, represents the channel estimation error matrix of the cascaded channel of each user, with independent and identically distributed terms ;

[0164] Step S2-2: Calculate the achievable rate of the user in the real environment and the mismatched environment;

[0165] In the real environment, the SINR of user is calculated as:

[0166] ;

[0167] Among them, is the precoding vector of the desired signal of the base station BS for the current user i, is the precoding vector of the desired signals of the base station BS for other users, is the precoding vector of the base station BS for the AN, is the cascaded channel of the user, is the transpose matrix of the column vector formed by sequentially extracting and arranging the diagonal elements of ; is the average noise variance, is the coefficient matrix containing only active elements; is the number of users;

[0168] User 's achievable rate is calculated as:

[0169] ;

[0170] The SINR of the user in a mismatch environment is calculated as follows:

[0171] ;

[0172] where, is the hybrid RIS reflection coefficient matrix in the mismatch environment, is the cascaded channel with imperfect estimation in the mismatch environment, is from the first active RIS element coefficients stripped out;

[0173] The achievable rate of user is calculated as follows:

[0174] .

[0175] Step S2-3: Calculate the radar signal SINR of the sensing target in the real environment and the mismatch environment;

[0176] The radar signal at the sensing target is expressed as:

[0177] ;

[0178] where, represents the channel vector between the BS and the sensing target, and ; represents the Gaussian white noise at the sensing target; is the distance between the base station BS and the sensing target; is the antenna array steering vector, expressed as:

[0179] ;

[0180] where, is the azimuth angle of the sensing target, and represent the wavelength and antenna spacing respectively; The echo SINR related to detecting the sensing target is expressed as:

[0181] ;

[0182] where, is the conjugate transpose of the antenna array steering vector, is the number of users, is the channel gain per unit distance.

[0183] In step S2-4, calculate the secure transmission rate of the user in the real environment and the mismatch environment, specifically:

[0184] In the real environment, user The secure transmission rate at is expressed as:

[0185] ;

[0186] The secure transmission rate of the user in the mismatch environment at is expressed as:

[0187] ;

[0188] where, is the achievable rate of the user ; is the eavesdropping rate of the sensing target on the user ; is the secure transmission rate of the user in the real environment at ; is the secure transmission rate of the user in the mismatch environment at ; .

[0189] The optimization problem obtained in step S2-5 is specifically:

[0190] Cooperatively formulate the beamforming matrix and reflection coefficient to achieve the secrecy sum rate. The optimization problem in the real environment:

[0191] ;

[0192] ;

[0193] ;

[0194] ;

[0195] ;

[0196] where, is the secrecy sum rate in the real environment; is the SINR of the radar echo signal; represents the SINR threshold of the echo signal, represents the transmit power of the BS, is the power budget of the active element, is the transmit power of the hybrid RIS active element, expressed as:

[0197] ;

[0198] where, represents the transmit power of the user desired signal; is the identity matrix of K active elements; is the channel vector between the hybrid RIS and the base station BS; is the conjugate transpose of the channel vector between the hybrid RIS and the base station BS;

[0199] The optimization problem in the mismatch environment is:

[0200] ;

[0201] ;

[0202] ;

[0203] ;

[0204] ;

[0205] where, is the secrecy sum rate in the mismatch environment.

[0206] In step S3, the fusion depth-directed intrinsic motivation exploration algorithm is used to improve the SAC reinforcement learning, specifically:

[0207] A SAC agent maintains three networks: two Q networks and one actor network, and each network is a multi-layer perceptron; the purpose of using two Q networks is to reduce the overestimation of the Q value. The Q network takes the state provided by the environment and the action generated by the actor network as inputs and outputs Q value estimates, and these estimates are scalars; an explorer network is additionally added through the depth-directed intrinsic motivation exploration algorithm to more deeply explore potential actions not selected by the actor network.

[0208] Step S3-1: Modify the reward function for training the agent:

[0209] ;

[0210] where, is the instantaneous reward calculated by the environment in the current state, is the average reward up to the current state;

[0211] Step S3-2: Add an explorer network to predict the coefficients of the hybrid RIS , where :

[0212] ;

[0213] where, when , , in the formula, the number of 1s refers to the number of elements of the transmit beamforming and the AN beamforming included in the action vector; each There are two terms to measure simultaneously the real and imaginary parts of;

[0214] Step S3-3: The prediction of the explorer network is used to perturb the action:

[0215] ;

[0216] where, is the Hadamard product, and the hyperparameter constrains the explorer network from making harmful perturbations to the actions selected by the actor;

[0217] Step S3-4: The losses of the Q-network and the actor network are corrected as follows:

[0218] ;

[0219] ;

[0220] ;

[0221] ;

[0222] where, ; The explorer network is optimized to maximize the sum of the absolute values of the TD errors:

[0223] ;

[0224] ;

[0225] Step S3-5: Update the deterministic exploration network through the deterministic policy gradient algorithm:

[0226] ;

[0227] ;

[0228] where, is the calculation of the mathematical expectation, is the temporal difference error after action perturbation, is the gradient operator for the perturbation vector , is the environmental state observed by the agent at time step t, is the original action selected by the agent at time step t, is the perturbation vector generated by the explorer network, is the explorer network with parameter , is the gradient operator for the parameter , For the Explorer Network parameters, is the learning rate of the Explorer Network, the perceived objective function of the Explorer Network.

[0229] Step S4: Use the improved SAC reinforcement learning to solve the optimization problem and achieve resource optimization, specifically as follows:

[0230] Step S4-1: Action setting: The policy network flattens and concatenates the transmit beamforming , AN beamforming and the coefficient matrix of the hybrid RIS as the action vector output;

[0231] The actor network generates the real and imaginary parts respectively, and then constructs , and . To satisfy the constraint conditions, the agent normalizes the output of the action. Therefore, an action vector consists of elements;

[0232] Step S4-2: State setting: The state vector includes the SINR of the radar echo signal, the user achievable rate, the perceived target eavesdropping rate, the previous action, and the cascaded channel matrix or its estimate , depending on whether the CSI of the BS is perfect or not; considering the SINR of the perceived target echo signal, the achievable rate of each user, and the eavesdropping rate of the perceived target on each user, therefore, terms related to the secrecy rate and radar detection performance can be obtained;

[0233] The cascaded channel estimation contribution of each user elements, and the previous action vector contributes elements, thus forming a -dimensional state vector;

[0234] Furthermore, the correlation between state dimensions will reduce the learning performance of the reinforcement learning agent. Therefore, after each environmental step, the state vector is whitened. Finally, the initial state of training still requires the previous action. Therefore, , and are initialized as the identity matrix to form the initial environmental state.

[0235] Step S4-3: At each time step, the reward is determined by the secrecy sum rate represented by the optimization problem in the real environment and the optimization problem in the mismatched environment.

[0236] To prove the effectiveness and effects of this invention application, a further detailed description of this invention application is made and simulation experiments are conducted for verification:

[0237] In the constructed ISAC system model, there are a total of 16 hybrid RIS elements, among which 5 are active elements and 11 are passive elements. There are a total of 4 users, the BS transmitting antenna is 4, and the BS transmitting power is 30 dBm;

[0238] The number of BS antennas is set to 4, the BS position is at (0,0), 4 users are set, and the users are randomly distributed on the circles with the BS as the center and radii of 10, 20, 30, and 40 m respectively. The sensing target Eve is placed at (0,50), that is, 50 m due north of the BS.

[0239] As Figure 3 、 4 shown, it can be inferred from the evaluation results that - Spatial exploration can achieve near-optimal results in all test scenarios; specifically, when , the performance of the SAC agent trained in the mismatched environment is significantly worse than that in the real environment; when is increased to 0.6, its performance is close to the real environment, which is expected because as the range shrinks, the possible interval of the hybrid RIS loss factor will also decrease; the method proposed in this invention application is not affected by the value. For each value of , this method shows robust performance, and the achieved secrecy sum rate is slightly lower than that in the real environment; in addition, - Spatial exploration has no problem in terms of convergence speed, that is, the learning curve is basically parallel to the real environment, which means that the exploration strategy can implicitly learn how its action selection affects the loss in RIS reflection, and compared with the total training time, the time spent in this process can be ignored.

[0240] The transmitting beam pattern is as Figure 5 shown, and the detected sensing target is located at the position; the experimental results show that the total radar transmission gain exhibits a highly focused characteristic in the direction; the normalized gain reaches the maximum value, indicating that the radar signal energy is precisely concentrated in the main beam direction, verifying the efficient energy projection ability of the scheme for the sensing direction of the sensing target; as the angle deviates from , the radar gain shows a symmetric attenuation trend and drops to about 0.6 at and at The edge region is close to zero, showing a significant sidelobe suppression effect, which is beneficial to reducing the negative impact of multipath interference on the radar detection accuracy; in addition, the proposed method forces the AN gain to be concentrated to the greatest extent on , that is, the main lobe direction of the sensing beam. Therefore, the solution proposed in this invention application not only meets the requirements of radar sensing, but also effectively suppresses Eve's eavesdropping behavior on users.

[0241] The above are only the embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A hybrid RIS-assisted ISAC resource optimization method based on DRL, characterized in that: The method includes the following steps: Step S1: Construct an ISAC system model; The ISAC system model includes a hybrid RIS, a base station BS, users, and a sensing target. The base station BS serves the users with the help of the hybrid RIS and simultaneously detects the position of the sensing target by using the sensing target; Step S2: Establish optimization problems in the real environment and the mismatched environment; Step S3: Improve SAC reinforcement learning by integrating the deep directed intrinsic motivation exploration algorithm; Step S4: Solve the optimization problem by using the improved SAC reinforcement learning to achieve resource optimization.

2. The hybrid RIS-aided ISAC resource optimization method based on DRL according to claim 1, wherein: In step S1, constructing the ISAC system model specifically includes: Step S1-1: Set the hybrid RIS; Hybrid RIS in total consists of active components and passive components. Among them, the active components need to additionally tune the amplitude of the incident signal compared with the passive components; Let represent a set of active components, ; Let represent the relay / reflection coefficient used on the th component: ; Among them, represents a phase shift, ; when then ; is the th component, is a complex number, is Euler's formula; Based on the following three diagonal matrices: , and where: ; ; Among them, is the coefficient matrix containing only passive components, is the coefficient matrix containing only active components, is the coefficient matrix of the entire hybrid RIS; ; and are the passive reflection coefficient containing only passive components and the active relay coefficient of active components, respectively; Step S1-2: Construct the transmission signal of the base station BS and the received signal of the users; Using precoding design, the transmission signal of the base station BS is expressed as: ; Among them, represents the expected signal of the user ; ; represents the AN generated by the base station BS ; and are the precoding vectors of the base station BS for and respectively; is the number of users; Adopting the AN strategy, by enhancing the sensing ability using the DoF and simultaneously counteracting the eavesdropping from the detected sensing target, the received signal at the user is expressed as: ; Among them, and respectively represent the channel vectors between the hybrid RIS and the user and between the hybrid RIS and the base station BS; transmits signals for the base station BS; is the conjugate transpose of the matrix; , , is the channel gain per unit distance, represents the distance between the hybrid RIS and the user ; represents the distance between the hybrid RIS and the base station BS; represents the Rayleigh fading vector between the hybrid RIS and the user ; represents the Rayleigh fading vector between the hybrid RIS and the base station BS; , is the additive noise caused by active components between the base station BS and the hybrid RIS, and is the additive white Gaussian noise at the user ; , denotes the average noise variance of the additive noise caused by active components between the base station BS and the hybrid RIS, and denotes the average noise variance of the additive white Gaussian noise at the user ; the noise power spectral density and the noise figure at the hybrid RIS and the user are the same, i.e., .

3. A hybrid RIS-assisted ISAC resource optimization method based on DRL according to claim 1, characterized in that: In step S2, establishing the optimization problems in the real environment and the mismatched environment specifically includes: Step S2-1: Obtain the phase amplitude model in the reflection loss of the hybrid RIS in the real environment and the mismatched environment; In the real environment, the hybrid RIS follows the phase-dependent amplitude model. When , , ; we get: ; Among them, , represents the inherent phase offset of the component, which is caused by the non-ideal characteristics of the hardware circuit, ; represents the dependence intensity of the reflection amplitude on the phase, and the larger its value, the more significant the influence of phase tuning on the reflection loss; The BS knows the independent cascaded channel of each user, denoted as: ; Therefore, in a real environment, the user The received signal at is expressed as: ; Among them, denotes the column vector formed by sequentially extracting and arranging the diagonal elements of ; is the transmission signal of base station BS, is the user at the noise, is the inverse matrix; The reflection of the hybrid RIS in the mismatched environment is assumed to be lossless, and the expression form is as follows: ; The agent can only access an imperfect estimate of the cascaded channel, that is: ; Among them, represents the channel estimation error matrix of the cascaded channel for each user, with independent and identically distributed terms ; Step S2-2: Calculate the achievable rate of the users in the real environment and the mismatched environment; User in the real environment The SINR calculation is as follows: ; Among them, is the precoding vector of the desired signal of the base station BS for the current user i, is the precoding vector of the desired signals of the base station BS for other users, is the precoding vector of the base station BS for the AN, is the cascaded channel of the user, is used to represent the transpose matrix of the column vector formed by extracting and arranging the diagonal elements of is the average noise variance, is the coefficient matrix containing only active elements; is the number of users; User The achievable rate is calculated as follows: ; The SINR calculation for the user in a mismatched environment is as follows: ; Among them, is the hybrid RIS reflection coefficient matrix in the mismatch environment, is the cascaded channel with imperfect estimation in the mismatch environment, is from stripped out the first active RIS element coefficients; User The achievable rate is calculated as follows: ; Step S2-3: Calculate the radar signal SINR of the sensing target in the real environment and the mismatched environment; Step S2-4: Calculate the secure transmission rate of the users in the real environment and the mismatched environment; Step S2-5: Obtain the optimization problem.

4. The hybrid RIS-aided ISAC resource optimization method based on DRL according to claim 3, characterized in that: In step S2-3, calculating the radar signal SINR of the sensing target specifically includes: The radar signal at the sensing target is expressed as: ; Among them, represents the channel vector between the BS and the sensed target, and ; represents the Gaussian white noise at the sensed target; is the distance between the base station BS and the sensed target; is the antenna array steering vector, expressed as: ; Among them, is the azimuth angle of the sensing target, and respectively represent the wavelength and the antenna spacing; the echo SINR related to detecting the sensing target is expressed as: ; wherein, is the conjugate transpose of the antenna array steering vector, is the number of users, is the channel gain per unit distance.

5. A hybrid RIS-assisted ISAC resource optimization method based on DRL according to claim 3, characterized in that: In step S2-4, calculating the secure transmission rate of the users in the real environment and the mismatched environment specifically includes: The secure transmission rate at the user is expressed as: ; The secure transmission rate at the user is expressed as follows: ; wherein, is the achievable rate of the user , is the eavesdropping rate of the sensing target to the user , is the secure transmission rate at the user in the real environment is the secure transmission rate at the user in the mismatched environment .

6. A hybrid RIS-aided ISAC resource optimization method based on DRL according to claim 3, characterized in that: In step S2-5, obtaining the optimization problem specifically includes: Collaboratively formulate the beamforming matrix and the reflection coefficient to achieve the secrecy sum rate. The optimization problem in the real environment: ; ; ; ; ; Among them, is the secrecy sum rate in the real environment; is the SINR of the radar echo signal; represents the SINR threshold of the echo signal, represents the transmit power of the BS, is the power budget of the active element, is the transmit power of the hybrid RIS active element, expressed as: ; wherein, represents the transmit power of the user desired signal; is the identity matrix of K active elements; is the channel vector between the hybrid RIS and the base station BS; is the conjugate transpose of the channel vector between the hybrid RIS and the base station BS; The optimization problem in the mismatched environment is: ; ; ; ; ; Among them, is the secrecy sum rate in the mismatched environment.

7. A hybrid RIS-aided ISAC resource optimization method based on DRL according to claim 1, characterized in that: In step S3, improving SAC reinforcement learning by integrating the deep directed intrinsic motivation exploration algorithm specifically includes: Step S3-1: Modify the reward function for training the agent: ; where, is the instantaneous reward calculated by the environment in the current state, is the average reward up to the current state; Step S3-2: Add an explorer network to predict the coefficients of the hybrid RIS , where : ; Among them, when , , where the number of 1s refers to the number of elements of the transmit beamforming and the AN beamforming included in the action vector; each has two terms, simultaneously measuring the real and imaginary parts of Step S3-3: The prediction of the explorer network is used to perturb the action: ; Among them, is the Hadamard product, and the hyperparameter constrains the explorer network not to make harmful perturbations to the actions selected by the actor; Step S3-4: Correct the losses of the Q network and the actor network as follows: ; ; ; ; Among them, ; The explorer network is optimized to maximize the sum of the absolute values of TD errors: ; ; Step S3-5: Update the deterministic exploration network by using the deterministic policy gradient algorithm: ; ; Among them, is for mathematical expectation calculation, is the temporal difference error after action perturbation, is the gradient operator for the perturbation vector ; is the environmental state observed by the agent at time step t, is the original action selected by the agent at time step t, is the perturbation vector generated by the explorer network, is the explorer network with parameter ; is the gradient operator for the parameter ; is the parameter of the explorer network ; is the learning rate of the explorer network, is the perception objective function of the explorer network.

Citation Information

Patent Citations

  • RIS-assisted backscatter communication perception integration method

    CN116248173A

  • Communication-centered RIS-assisted cellular-free ISAC network joint beam forming method

    CN117375683A

  • Combined beam forming design method for hybrid RIS-assisted millimeter wave ISAC system

    CN118118072A

  • Transmission and reflection dual-function reconfigurable intelligent surface auxiliary sensing integrated safety communication system and method

    CN118214469A

  • Physical layer security optimization method of RIS-assisted ISAC system

    CN118540694A

Cited By

  • Internet of Things security resource optimization method and system based on multi-agent reinforcement learning

    CN120916144A