A hybrid RIS-assisted ISAC resource optimization method based on DRL

Through the hybrid RIS assisted ISAC resource optimization method based on DRL, the multivariate coupling non-convex problem and sensitive information leakage in RIS assisted ISAC system are solved, and efficient resource allocation and confidentiality improvement in complex environments are achieved.

CN120357930BActive Publication Date: 2025-08-22NANJING UNIV OF POSTS & TELECOMM
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510840209.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-08-22
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

The prior art has high complexity of multivariable coupling, limitations of relying on ideal assumptions, and severe problems of sensitive information leakage in RIS-assisted ISAC systems, especially in the downstream fading of 6G high-frequency communication.

Method used

Using a hybrid RIS-assisted ISAC resource optimization method based on DRL, we will improve SAC reinforcement learning by building an ISAC system model, combining deep-directed intrinsic motivation exploration algorithms, and jointly optimize transmit beamforming, AN interference and RIS phase shift, taking into account RIS hardware damage and channel uncertainty to maximize confidentiality.

Benefits of technology

It breaks through the dependence of traditional algorithms on perfect channel state information and ideal reflection model, improves the confidentiality of the system and adaptability in complex environments, optimizes resource configuration, and effectively suppresses eavesdropping behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120357930B_ABST
    Figure CN120357930B_ABST
Patent Text Reader

Abstract

The present invention discloses a hybrid RIS-assisted ISAC resource optimization method based on DRL, belonging to the field of wireless communication technology. The method comprises: constructing an ISAC system model; establishing optimization problems in real environments and mismatched environments; integrating a deep directed intrinsic motivation exploration algorithm to improve SAC reinforcement learning; and using the improved SAC reinforcement learning to solve the optimization problem and achieve resource optimization. The present invention achieves maximum confidentiality rate under the constraints of RIS hardware damage and channel uncertainty by jointly designing hybrid RIS and AN interference. Simultaneously, a phase-dependent reflection amplitude model is introduced to adapt to RIS hardware damage. Using DRL to jointly optimize transmit beamforming, AN interference, and RIS phase shift, the method overcomes the traditional algorithm's reliance on perfect channel state information (CSI) and ideal reflection models, achieving ISAC system resource optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communications, and in particular relates to a DRL-based hybrid RIS-assisted ISAC resource optimization method. Background Art

[0002] With the explosive growth of Internet of Things (IoT) applications, spectrum resources are becoming increasingly scarce. ISAC technology, with its advantages in sharing hardware platforms, improving spectrum efficiency, and reducing interference, has become a key approach to addressing spectrum congestion. ISAC achieves synergy between communication and radar sensing by designing dual-function waveforms. However, in channel degradation scenarios, active beamforming alone struggles to balance performance and security requirements. RIS, a green, programmable reflection technology, provides ISAC systems with additional degrees of freedom by reconfiguring the wireless propagation environment, significantly improving multi-user interference suppression and radar sensing performance. However, the security of ISAC systems faces significant challenges: when the perceived target is a potential eavesdropper, the power concentration strategy of traditional radars can lead to the leakage of sensitive information. To this end, physical layer security and artificial noise technologies have been introduced to enhance security, but existing research on the security optimization of RIS-assisted ISAC systems is insufficient.

[0003] Traditional convex optimization-based methods (such as Taylor expansion, semidefinite relaxation, and alternating optimization) face limitations due to high complexity and reliance on ideal assumptions when solving non-convex problems with multivariable coupling. This is particularly true in the presence of RIS hardware impairments (such as phase-dependent reflection amplitude) and dynamic channel environments. Deep reinforcement learning, through interactive learning between an agent and its environment, offers a new paradigm for resource optimization in time-varying wireless systems. In RIS-assisted ISAC scenarios, DRL can jointly optimize transmit beamforming, AN interference, and RIS phase shift, transcending the traditional algorithms' reliance on perfect channel state information (CSI) and ideal reflection models.

[0004] For example, the soft actor-critic DRL framework avoids local optima by maximizing policy entropy, significantly improving system confidentiality while ensuring radar minimum signal-to-interference-noise ratio and power constraints. Furthermore, to address the performance loss caused by RIS phase-dependent amplitude, DRL can adaptively learn optimal policies under non-ideal hardware characteristics, providing a data-driven solution for complex system modeling. Recent research has further expanded the application of DRL to non-ideal conditions (such as low-resolution RIS and multi-user MIMO systems). By improving network architecture and training mechanisms, it achieves performance close to ideal assumptions, opening up new avenues for resource optimization in hybrid RIS-assisted ISAC systems.

[0005] A physical layer security optimization method for RIS-assisted ISAC systems, published in China (Application No. 202410624177.2), uses the eavesdropper's detection probability as the system perception performance metric and the user's total confidentiality rate as the system communication performance metric. By jointly optimizing the ISAC-BS beamforming and RIS phase shift, the method maximizes the confidentiality rate for all users while maintaining a certain detection probability. The solution is first decoupled into three subproblems, each of which alternately optimizes the three decision variables: the base station receive beam, the RIS phase, and the base station transmit beam. These subproblems are then transformed into deterministic convex optimization models using Rayleigh quotient problems, first-order Taylor expansions, semidefinite relaxations, and successive convex approximations. Simulations show that increasing the number of RIS elements and base station antennas improves the system confidentiality rate. This solution also outperforms solutions without RIS or with random RIS phases in terms of improved confidentiality rate and reduced power consumption.

[0006] A hybrid RIS-assisted millimeter-wave ISAC system joint beamforming design method (application number: 202410333259.1) was published in China. This method constructs an optimization function based on the ISAC system's sensor performance to jointly optimize the analog beamformer, digital beamformer, and HRIS reflection coefficient matrix, maximizing the worst-case illumination power while ensuring communication quality. The solution first uses surrogate optimization techniques to decouple the non-convex problem into three subproblems. Quadratically constrained quadratic programming (QCQP) and semidefinite programming relaxation (SDR) techniques are then used to transform the DFBS and HRIS beamforming optimization subproblems into convex problems. The DFBS hybrid precoding design subproblem is solved using the MO-AltMin algorithm based on minimizing the Euclidean distance. Simulation results demonstrate the effectiveness of the proposed algorithm, demonstrating that a fully connected hybrid precoding architecture achieves near-optimal performance. Hybrid RIS achieves a good balance between system performance and hardware cost.

[0007] The above patent application has the following deficiencies: 1. The constructed system model involves numerous parameters and complex channel models. In the process of decoupling and solving the mathematical model, the model complexity increases significantly, and the demand for computing resources also increases significantly;

[0008] 2. The channel state information and RIS phase amplitude design, which rely on ideal assumptions, are not applicable in actual situations where channel state information cannot be obtained or the RIS hardware is damaged. It only considers passive intelligent reflectors, which leads to the existence of "multiplicative fading", especially in 6G high-frequency communications, where the "multiplicative fading" effect is more significant.

[0009] Therefore, how to solve the non-convex problem of multivariable coupling, which faces the high complexity, the limitation of relying on ideal assumptions, and the serious problem of leaking sensitive information, is the technical problem that the present invention aims to solve. Summary of the Invention

[0010] The object of the present invention is to provide a DRL-based hybrid RIS-assisted ISAC resource optimization method to solve the problems raised in the above background technology.

[0011] The object of the present invention is achieved by: a hybrid RIS-assisted ISAC resource optimization method based on DRL, characterized in that the method comprises the following steps:

[0012] Step S1: Construct ISAC system model;

[0013] The ISAC system model includes a hybrid RIS, a base station BS, a user, and a sensing target. The base station BS serves the user with the help of the hybrid RIS and uses the sensing target to detect the location of the sensing target.

[0014] Step S2: Establish the optimization problem in the real environment and the mismatch environment;

[0015] Step S3: Integrate the deep directed intrinsic motivation exploration algorithm to improve SAC reinforcement learning;

[0016] Step S4: Use the improved SAC reinforcement learning to solve the optimization problem and achieve resource optimization.

[0017] Preferably, the ISAC system model is constructed in step S1, specifically:

[0018] Step S1-1: Setting up hybrid RIS;

[0019] Hybrid RIS Shared components, including Active components and passive components, wherein active components require additional tuning of the amplitude of the incident signal compared to passive components;

[0020] set up express A collection of active components, ;set up Indicates the Relay / reflection coefficient used on each element:

[0021] ;

[0022] in, represents the phase shift, ;when hour, ; For the components, is plural, is Euler's formula;

[0023] based on The three diagonal matrices of : , and ,in:

[0024] ;

[0025] ;

[0026] in, is the coefficient matrix containing only passive components, is the coefficient matrix containing only active elements, is the coefficient matrix of the entire hybrid RIS; ; and They are the passive reflection coefficient including only passive components and the active relay coefficient including active components;

[0027] Step S1-2: constructing a base station BS transmission signal and a user reception signal;

[0028] Using the precoding design, the base station BS transmission signal is expressed as:

[0029] ;

[0030] in, Represents a user The expected signal, ; represents the AN generated by the base station BS, ; and Base station BS pair and Precoding vector of is the number of users;

[0031] Adopting the AN strategy, by utilizing more DoF to enhance the perception capability, while resisting the eavesdropping from the detection perception target, the user The received signal at is expressed as:

[0032] ;

[0033] in, and Represents hybrid RIS and user and the channel vector between the hybrid RIS and the base station BS; Transmitting signals to the base station BS; is the conjugate transpose of the matrix;

[0034] , , is the channel gain per unit distance, Indicates a mix of RIS and user The distance between represents the distance between the hybrid RIS and the base station BS; Indicates a mix of RIS and user The Rayleigh fading vector between represents the Rayleigh fading vector between the hybrid RIS and the base station BS;

[0035] , Between base station BS and hybrid RIS The additive noise caused by the active components, From base station BS to user Additive white Gaussian noise at ; , Indicates the connection between base station BS and hybrid RIS The average noise variance of the additive noise caused by the active components, Indicates the base station BS to the user The average noise variance of the additive Gaussian white noise at , mixed RIS and user The noise power spectral density and noise figure at are the same, that is, .

[0036] Preferably, the optimization problem in the real environment and the mismatch environment is established in step S2, specifically:

[0037] Step S2-1: Obtaining the phase amplitude model of the hybrid RIS reflection loss in the real environment and the mismatch environment;

[0038] In real environments, the hybrid RIS follows a phase-dependent amplitude model. , , ;get:

[0039] ;

[0040] in, , Represents the inherent phase shift of the component, caused by the non-ideal characteristics of the hardware circuit, ; represents the dependence of the reflection amplitude on the phase, , the larger its value is, the more significant the effect of phase tuning on return loss is;

[0041] The BS knows the independent concatenated channel of each user, which is recorded as:

[0042] ;

[0043] Therefore, in real-world environments, users The received signal at is expressed as:

[0044] ;

[0045] in, Indicated by Extract and arrange the diagonal elements of in sequence to form a column vector; Transmitting signals to the base station BS, For users The noise at is the inverted matrix;

[0046] The hybrid RIS reflection in the mismatched environment is assumed to be lossless and is expressed as follows:

[0047] ;

[0048] The agent only has access to an imperfect estimate of the cascade channel, namely:

[0049] ;

[0050] in, The channel estimation error matrix of the concatenated channel for each user is independent and identically distributed. item ;

[0051] Step S2-2: Calculate the achievable rates of users in the real environment and the mismatched environment;

[0052] Users in real environment The SINR is calculated as:

[0053] ;

[0054] in, is the precoding vector of the desired signal from the base station BS to the current user i, is the precoding vector of the base station BS for other users’ desired signals, is the precoding vector of the base station BS to AN, is the user's cascade channel, To indicate The diagonal elements of are extracted and arranged in sequence to form the transposed matrix of the column vector. is the average noise variance, is the coefficient matrix containing only active elements; is the number of users;

[0055] user The achievable rate is calculated as:

[0056] ;

[0057] In a mismatched environment, users The SINR is calculated as:

[0058] ;

[0059] in, is the hybrid RIS reflection coefficient matrix under mismatched environment, is the imperfectly estimated cascade channel in a mismatched environment, For The front part separated from the active RIS element coefficients;

[0060] user The achievable rate is calculated as:

[0061] ;

[0062] Step S2-3: Calculate the radar signal SINR of the target under the real environment and the mismatch environment;

[0063] Step S2-4: Calculate the secure transmission rate of the user in the real environment and the mismatched environment;

[0064] Step S2-5: Obtain the optimization problem.

[0065] Preferably, the calculation of the radar signal SINR of the perceived target in step S2-3 is specifically as follows:

[0066] The radar signal at the perceived target is expressed as:

[0067] ;

[0068] in, represents the channel vector between the BS and the sensing target, and ; Represents Gaussian white noise at the perceived target; is the distance between the base station BS and the sensing target; is the antenna array steering vector, expressed as:

[0069] ;

[0070] in, is the azimuth of the perceived target, and denote wavelength and antenna spacing respectively; the echo SINR related to the detection and perception target is expressed as:

[0071] ;

[0072] in, is the conjugate transpose of the antenna array steering vector, is the number of users, is the channel gain per unit distance.

[0073] Preferably, the calculation of the secure transmission rate of the user in the real environment and the mismatch environment in step S2-4 is specifically as follows:

[0074] In real-world environments, users The safe transmission rate at is expressed as:

[0075] ;

[0076] Users in mismatched environments The safe transmission rate at is expressed as:

[0077] ;

[0078] in, For users The achievable rate, To perceive the target user The eavesdropping rate, For users in real environment The safe transmission rate at For users in mismatched environments The safe transmission rate at .

[0079] Preferably, the optimization problem obtained in step S2-5 is specifically:

[0080] Collaboratively formulate beamforming matrices and reflection coefficients to achieve confidentiality and speed, and optimize the problem in real-world environments:

[0081] ;

[0082] ;

[0083] ;

[0084] ;

[0085] ;

[0086] in, To ensure confidentiality and speed in real-world environments; is the SINR of the radar echo signal; represents the SINR threshold of the echo signal, represents the transmit power of the BS, is the power budget for active components, is the transmission power of the hybrid RIS active element, expressed as:

[0087] ;

[0088] in, Indicates the transmission power of the user's desired signal; is the identity matrix of K active elements; is the channel vector between the hybrid RIS and the base station BS; is the conjugate transpose of the channel vector between the hybrid RIS and the base station BS;

[0089] The optimization problem under the mismatch environment is:

[0090] ;

[0091] ;

[0092] ;

[0093] ;

[0094] ;

[0095] in, For confidentiality and speed in mismatched environments.

[0096] Preferably, in step S3, the deep directed intrinsic motivation exploration algorithm is integrated to improve the SAC reinforcement learning, specifically:

[0097] Step S3-1: Modify the reward function used to train the agent:

[0098] ;

[0099] in, is the instantaneous reward calculated by the environment in the current state, is the average reward up to the current state;

[0100] Step S3-2: Add an explorer network to predict the coefficients of the hybrid RIS ,in :

[0101] ;

[0102] Among them, when , , where the number of 1s refers to the transmit beamforming and AN beamforming The number of elements included in the action vector; each There are two items that can be measured simultaneously The real and imaginary parts of

[0103] Step S3-3: Explorer Network The prediction of is used to perturb the action:

[0104] ;

[0105] in, is the Hadamard product, the hyperparameter Constrain the explorer network to not make harmful perturbations to the actions chosen by the actors;

[0106] Step S3-4: Modify the losses of the Q network and the actor network as follows:

[0107] ;

[0108] ;

[0109] ;

[0110] ;

[0111] in, The explorer network is optimized to maximize the sum of the absolute values ​​of the TD errors:

[0112] ;

[0113] ;

[0114] Step S3-5: Update the deterministic exploration network through the deterministic policy gradient algorithm:

[0115] ;

[0116] ;

[0117] in, is the mathematical expectation calculation, is the timing difference error after motion disturbance, is the perturbation vector The gradient operator, is the state of the environment observed by the agent at time step t, is the original action chosen by the agent at time step t, The perturbation vector generated for the explorer network, The parameters are The Explorer Network, For the parameters The gradient operator, For Explorer Network Parameters, is the learning rate of the explorer network, Perception objective function of the explorer network.

[0118] Compared with the prior art, the present invention has the following improvements and advantages:

[0119] 1. By jointly designing hybrid RIS and AN jammers, the confidentiality rate is maximized while taking into account RIS hardware impairments and channel uncertainty. Furthermore, a phase-dependent reflection amplitude model is introduced to adapt to RIS hardware impairments. Using DRL, transmit beamforming, AN jammers, and RIS phase shifts are jointly optimized, breaking through the traditional algorithm's reliance on perfect channel state information (CSI) and ideal reflection models, thereby optimizing ISAC system resources.

[0120] 2. Through the SAC algorithm and the deep directional intrinsic motivation exploration algorithm, a balance between exploration and utilization efficiency is achieved, showing better adaptability in complex transmit beamforming, AN and hybrid RIS coefficient matrix design. BRIEF DESCRIPTION OF THE DRAWINGS

[0121] Figure 1 Schematic diagram of the process of the present invention.

[0122] Figure 2 This is a schematic diagram of the ISAC system model structure.

[0123] Figure 3 for Schematic diagram of instantaneous confidentiality and rate in the case of .

[0124] Figure 4 for Schematic diagram of instantaneous confidentiality and rate in the case.

[0125] Figure 5 Schematic diagram of the transmission beam direction. DETAILED DESCRIPTION

[0126] The present invention is further summarized below with reference to the accompanying drawings.

[0127] like Figure 1 As shown, a hybrid RIS-assisted ISAC resource optimization method based on DRL includes the following steps:

[0128] Step S1: Construct ISAC system model;

[0129] like Figure 2 As shown, the ISAC system model includes hybrid RIS, base station BS, users and sensing targets, and the number of antennas is The base stations BS With the help of a hybrid RIS The hybrid RIS is strategically placed near the user to enhance downlink multi-user communication, and assuming that the hybrid RIS is located far away from the sensing target, the radar echo signal reflected by the hybrid RIS is weak and can be ignored. The sensing target is considered to be a potential eavesdropper Eve, who may eavesdrop on the user's downlink transmission.

[0130] Construct the ISAC system model, specifically:

[0131] Step S1-1: Setting up hybrid RIS;

[0132] Hybrid RIS Shared components, including Active components and passive components, wherein active components require additional tuning of the amplitude of the incident signal compared to passive components;

[0133] set up express A collection of active components, ;set up Indicates the Relay / reflection coefficient used on each element:

[0134] ;

[0135] in, represents the phase shift, ;when hour, ; For the components, is plural, is Euler's formula;

[0136] based on The three diagonal matrices of : , and ,in:

[0137] ;

[0138] ;

[0139] in, is the coefficient matrix containing only passive components, is the coefficient matrix containing only active elements, is the coefficient matrix of the entire hybrid RIS; ; and They are the passive reflection coefficient including only passive components and the active relay coefficient including active components;

[0140] Step S1-2: constructing a base station BS transmission signal and a user reception signal;

[0141] Using the precoding design, the base station BS transmission signal is expressed as:

[0142] ;

[0143] in, Represents a user The expected signal, ; represents the AN generated by the base station BS, ; and Base station BS pair and Precoding vector of is the number of users;

[0144] Adopting the AN strategy, by utilizing more DoF to enhance the perception capability, while resisting the eavesdropping from the detection perception target, the user The received signal at is expressed as:

[0145] ;

[0146] in, and Represents hybrid RIS and user and the channel vector between the hybrid RIS and the base station BS; Transmitting signals to the base station BS; is the conjugate transpose of the matrix;

[0147] , , is the channel gain per unit distance, Indicates a mix of RIS and user The distance between represents the distance between the hybrid RIS and the base station BS; Indicates a mix of RIS and user The Rayleigh fading vector between represents the Rayleigh fading vector between the hybrid RIS and the base station BS;

[0148] , Between base station BS and hybrid RIS The additive noise caused by the active components, From base station BS to user Additive white Gaussian noise at ; , Indicates the connection between base station BS and hybrid RIS The average noise variance of the additive noise caused by the active components, Indicates the base station BS to the user The average noise variance of the additive Gaussian white noise at , mixed RIS and user The noise power spectral density and noise figure at are the same, that is, .

[0149] Step S2: Establish the optimization problem in the real environment and the mismatch environment, specifically:

[0150] Step S2-1: Obtaining the phase amplitude model of the hybrid RIS reflection loss in the real environment and the mismatch environment;

[0151] In real environments, the hybrid RIS follows a phase-dependent amplitude model. , , ;get:

[0152] ;

[0153] in, , Represents the inherent phase shift of the component, caused by the non-ideal characteristics of the hardware circuit, ; represents the dependence of the reflection amplitude on the phase, , the larger its value is, the more significant the effect of phase tuning on return loss is;

[0154] The BS knows the independent concatenated channel of each user, which is recorded as:

[0155] ;

[0156] Therefore, in real-world environments, users The received signal at is expressed as:

[0157] ;

[0158] in, Indicated by Extract and arrange the diagonal elements of in sequence to form a column vector; Transmitting signals to the base station BS, For users The noise at is the inverted matrix;

[0159] The hybrid RIS reflection in the mismatched environment is assumed to be lossless and is expressed as follows:

[0160] ;

[0161] The agent only has access to an imperfect estimate of the cascade channel, namely:

[0162] ;

[0163] in, The channel estimation error matrix of the concatenated channel for each user is independent and identically distributed. item ;

[0164] Step S2-2: Calculate the achievable rates of users in the real environment and the mismatched environment;

[0165] Users in real environment The SINR is calculated as:

[0166] ;

[0167] in, is the precoding vector of the desired signal from the base station BS to the current user i, is the precoding vector of the base station BS for other users’ desired signals, is the precoding vector of the base station BS to AN, is the user's cascade channel, To indicate The diagonal elements of are extracted and arranged in sequence to form the transposed matrix of the column vector. is the average noise variance, is the coefficient matrix containing only active elements; is the number of users;

[0168] user The achievable rate is calculated as:

[0169] ;

[0170] In a mismatched environment, users The SINR is calculated as:

[0171] ;

[0172] in, is the hybrid RIS reflection coefficient matrix under mismatched environment, is the imperfectly estimated cascade channel in a mismatched environment, For The front part separated from the active RIS element coefficients;

[0173] user The achievable rate is calculated as:

[0174] .

[0175] Step S2-3: Calculate the radar signal SINR of the target under the real environment and the mismatch environment;

[0176] The radar signal at the perceived target is expressed as:

[0177] ;

[0178] in, represents the channel vector between the BS and the sensing target, and ; Represents Gaussian white noise at the perceived target; is the distance between the base station BS and the sensing target; is the antenna array steering vector, expressed as:

[0179] ;

[0180] in, is the azimuth of the perceived target, and denote wavelength and antenna spacing respectively; the echo SINR related to the detection and perception target is expressed as:

[0181] ;

[0182] in, is the conjugate transpose of the antenna array steering vector, is the number of users, is the channel gain per unit distance.

[0183] In step S2-4, the secure transmission rates of users in the real environment and the mismatched environment are calculated as follows:

[0184] In real-world environments, users The safe transmission rate at is expressed as:

[0185] ;

[0186] Users in mismatched environments The safe transmission rate at is expressed as:

[0187] ;

[0188] in, For users The achievable rate, To perceive the target user The eavesdropping rate, For users in real environment The safe transmission rate at For users in mismatched environments The safe transmission rate at .

[0189] In step S2-5, the optimization problem is obtained, specifically:

[0190] Collaboratively formulate beamforming matrices and reflection coefficients to achieve confidentiality and speed, and optimize the problem in real-world environments:

[0191] ;

[0192] ;

[0193] ;

[0194] ;

[0195] ;

[0196] in, To ensure confidentiality and speed in real-world environments; is the SINR of the radar echo signal; represents the SINR threshold of the echo signal, represents the transmit power of the BS, is the power budget for active components, is the transmission power of the hybrid RIS active element, expressed as:

[0197] ;

[0198] in, Indicates the transmission power of the user's desired signal; is the identity matrix of K active elements; is the channel vector between the hybrid RIS and the base station BS; is the conjugate transpose of the channel vector between the hybrid RIS and the base station BS;

[0199] The optimization problem under the mismatch environment is:

[0200] ;

[0201] ;

[0202] ;

[0203] ;

[0204] ;

[0205] in, For confidentiality and speed in mismatched environments.

[0206] In step S3, the deep directed intrinsic motivation exploration algorithm is integrated to improve SAC reinforcement learning, specifically:

[0207] A SAC agent maintains three networks: two Q networks and an actor network, each of which is a multi-layer perceptron. The purpose of using two Q networks is to reduce the overestimation of Q values. The Q network takes the state provided by the environment and the action generated by the actor network as input and outputs Q value estimates, which are scalars. An additional explorer network is added through a deep directed intrinsic motivation exploration algorithm to explore potential actions that are not selected by the actor network in more depth.

[0208] Step S3-1: Modify the reward function used to train the agent:

[0209] ;

[0210] in, is the instantaneous reward calculated by the environment in the current state, is the average reward up to the current state;

[0211] Step S3-2: Add an explorer network to predict the coefficients of the hybrid RIS ,in :

[0212] ;

[0213] Among them, when , , where the number of 1s refers to the transmit beamforming and AN beamforming The number of elements included in the action vector; each There are 2 items to measure simultaneously The real and imaginary parts of

[0214] Step S3-3: Explorer Network The prediction of is used to perturb the action:

[0215] ;

[0216] in, is the Hadamard product, the hyperparameter Constrain the explorer network to not make harmful perturbations to the actions chosen by the actors;

[0217] Step S3-4: Modify the losses of the Q network and the actor network as follows:

[0218] ;

[0219] ;

[0220] ;

[0221] ;

[0222] in, The explorer network is optimized to maximize the sum of the absolute values ​​of the TD errors:

[0223] ;

[0224] ;

[0225] Step S3-5: Update the deterministic exploration network through the deterministic policy gradient algorithm:

[0226] ;

[0227] ;

[0228] in, is the mathematical expectation calculation, is the timing difference error after motion disturbance, is the perturbation vector The gradient operator, is the state of the environment observed by the agent at time step t, is the original action chosen by the agent at time step t, The perturbation vector generated for the explorer network, The parameters are The Explorer Network, For the parameters The gradient operator, For Explorer Network Parameters, is the learning rate of the explorer network, Perception objective function of the explorer network.

[0229] Step S4: Use the improved SAC reinforcement learning to solve the optimization problem and achieve resource optimization, specifically:

[0230] Step S4-1: Action setting: The policy network will transmit beamforming , AN beamforming and the coefficient matrix of the hybrid RIS After flattening, concatenate them and output them as action vectors;

[0231] The actor network generates the real part and the imaginary part respectively, and then constructs 、 and , in order to satisfy the constraints, the agent normalizes the output of the action. Therefore, an action vector is composed of Elements

[0232] Step S4-2: State setting: The state vector includes the radar echo signal SINR, user reachable rate, perceived target eavesdropping rate, the action of the previous step, and the cascade channel matrix or its estimated value , depending on whether the BS's CSI is perfect or imperfect; considering the SINR of the sensing target echo signal, the achievable rate of each user, and the eavesdropping rate of the sensing target to each user, we can get Items related to confidentiality rate and radar detection performance;

[0233] Concatenated channel estimation contribution of each user elements, and the action vector of the previous step contributes elements, thus forming a -dimensional state vector;

[0234] Furthermore, the correlation between state dimensions will reduce the learning performance of the reinforcement learning agent. Therefore, after each environment step, the state vector is whitened. Finally, the initial state of the training still requires the action of the previous step, so 、 and Initialized to the identity matrix, it constitutes the initial environment state.

[0235] Step S4-3: At each time step, the confidentiality and rate decisions represented by the optimization problem under the true environment and the optimization problem under the mismatched environment are rewarded.

[0236] In order to prove the validity and effect of the present invention, the present invention is further described in detail and verified by simulation experiments:

[0237] In the constructed ISAC system model, there are 16 hybrid RIS components, including 5 active components and 11 passive components. There are 4 users, 4 BS transmitting antennas, and BS transmitting power of 30dBm.

[0238] The number of BS antennas is set to 4, the BS is located at (0,0), and 4 users are set. The users are randomly distributed on circles with radii of 10, 20, 30, and 40 meters with the BS as the center. The sensing target Eve is placed at (0,50), which is 50 meters north of the BS.

[0239] like Figure 3 、 4 As shown, it can be inferred from the evaluation results that - Spatial exploration achieves near-optimal results in all tested scenarios; specifically, when When , the performance of the SAC agent trained in the mismatched environment is significantly worse than that in the real environment; when When it is increased to 0.6, its performance is close to the real environment, which is expected because As the range is reduced, the possible range of the mixed RIS loss factor will also be reduced; the method proposed in the present invention is not affected by The impact of the value, for For each value of , the method shows robust performance, achieving confidentiality and rate slightly lower than the real environment; in addition, - Spatial exploration has no problem with convergence speed, i.e., the learning curve is roughly parallel to the real environment, which means that the exploration policy can implicitly learn how its action selection affects the loss in RIS reflections, and the time spent on this process is negligible compared to the total training time.

[0240] The transmit beam pattern is as follows Figure 5 As shown, the detection perception target is located at Position; Experimental results show that the total transmission gain of radar is The direction shows a highly focused characteristic; the normalized gain reaches the maximum value, indicating that the radar star energy is accurately concentrated in the main beam direction, verifying the efficient energy projection capability of the scheme to perceive the direction of the target; as the angle deviates , the radar gain shows a symmetrical attenuation trend. It drops to about 0.6 at The edge area is close to the zero point, showing a significant sidelobe suppression effect, which is beneficial to reduce the negative impact of multipath interference on radar detection accuracy; in addition, the proposed method forces the AN gain to be concentrated to the maximum extent , that is, the direction of the main lobe of the sensing beam. Therefore, the solution proposed in the present invention not only meets the needs of radar perception, but also effectively suppresses Eve's eavesdropping on users.

[0241] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A hybrid RIS-assisted ISAC resource optimization method based on DRL, characterized by: The method comprises the following steps: Step S1: Construct ISAC system model; The ISAC system model includes a hybrid RIS, a base station BS, a user, and a sensing target. The base station BS serves the user with the help of the hybrid RIS and uses the sensing target to detect the location of the sensing target. Step S1-1: Setting up hybrid RIS; Hybrid RIS Shared components, including Active components and passive components, wherein active components require additional tuning of the amplitude of the incident signal compared to passive components; set up express A collection of active components, ;set up Indicates the Relay / reflection coefficient used on each element: ; in, represents the phase shift, ;when hour, ; For the components, is plural, is Euler's formula; based on The three diagonal matrices of : , and ,in: ; ; in, is the coefficient matrix containing only passive components, is the coefficient matrix containing only active elements, is the coefficient matrix of the entire hybrid RIS; ; and They are the passive reflection coefficient including only passive components and the active relay coefficient including active components; Step S1-2: constructing a base station BS transmission signal and a user reception signal; Using the precoding design, the base station BS transmission signal is expressed as: ; in, Represents a user The expected signal, ; represents the AN generated by the base station BS, ; and Base station BS pair and Precoding vector of is the number of users; Adopting AN strategy, by utilizing DoF to enhance perception capability, while resisting eavesdropping from detecting perception targets, users The received signal at is expressed as: ; in, and Represents hybrid RIS and user and the channel vector between the hybrid RIS and the base station BS; Transmitting signals to the base station BS; is the conjugate transpose of the matrix; , , is the channel gain per unit distance, Indicates a mix of RIS and user The distance between represents the distance between the hybrid RIS and the base station BS; Indicates a mix of RIS and user The Rayleigh fading vector between represents the Rayleigh fading vector between the hybrid RIS and the base station BS; , Between base station BS and hybrid RIS The additive noise caused by the active components, From base station BS to user Additive white Gaussian noise at ; , Indicates the connection between base station BS and hybrid RIS The average noise variance of the additive noise caused by the active components, Indicates the base station BS to the user The average noise variance of the additive Gaussian white noise at , mixed RIS and user The noise power spectral density and noise figure at are the same, that is, ; Step S2: Establish the optimization problem in the real environment and the mismatch environment; Step S3: Integrate the deep directed intrinsic motivation exploration algorithm to improve SAC reinforcement learning; Step S4: Use the improved SAC reinforcement learning to solve the optimization problem and achieve resource optimization.

2. The DRL-based hybrid RIS-assisted ISAC resource optimization method according to claim 1, characterized in that: In step S2, the optimization problem under the real environment and the mismatch environment is established, specifically: Step S2-1: Obtaining the phase amplitude model of the hybrid RIS reflection loss in the real environment and the mismatch environment; In real environments, the hybrid RIS follows a phase-dependent amplitude model. , , ;get: ; in, , Represents the inherent phase shift of the component, caused by the non-ideal characteristics of the hardware circuit, ; represents the dependence of the reflection amplitude on the phase, , the larger its value is, the more significant the effect of phase tuning on return loss is; The BS knows the independent concatenated channel of each user, which is recorded as: ; Therefore, in real-world environments, users The received signal at is expressed as: ; in, Indicated by Extract and arrange the diagonal elements of in sequence to form a column vector; Transmitting signals to the base station BS, For users The noise at is the inverted matrix; The hybrid RIS reflection in the mismatched environment is assumed to be lossless and is expressed as follows: ; The agent only has access to an imperfect estimate of the cascade channel, namely: ; in, The channel estimation error matrix of the concatenated channel for each user is independent and identically distributed. item ; Step S2-2: Calculate the achievable rates of users in the real environment and the mismatched environment; Users in real environment The SINR is calculated as: ; in, is the precoding vector of the desired signal from the base station BS to the current user i, is the precoding vector of the base station BS for other users’ desired signals, is the precoding vector of the base station BS to AN, is the user's cascade channel, To indicate The diagonal elements of are extracted and arranged in sequence to form the transposed matrix of the column vector. is the average noise variance, is the coefficient matrix containing only active elements; is the number of users; user The achievable rate is calculated as: ; In a mismatched environment, users The SINR is calculated as: ; in, is the hybrid RIS reflection coefficient matrix under mismatched environment, is the imperfectly estimated cascade channel in a mismatched environment, For The front part separated from the active RIS element coefficients; user The achievable rate is calculated as: ; Step S2-3: Calculate the radar signal SINR of the target under the real environment and the mismatch environment; Step S2-4: Calculate the secure transmission rate of the user in the real environment and the mismatched environment; Step S2-5: Obtain the optimization problem.

3. The DRL-based hybrid RIS-assisted ISAC resource optimization method according to claim 2, characterized in that: The calculation of the radar signal SINR of the perceived target in step S2-3 is specifically as follows: The radar signal at the perceived target is expressed as: ; in, represents the channel vector between the BS and the sensing target, and ; Represents Gaussian white noise at the perceived target; is the distance between the base station BS and the sensing target; is the antenna array steering vector, expressed as: ; in, is the azimuth of the perceived target, and denote wavelength and antenna spacing respectively; the echo SINR related to the detection and perception target is expressed as: ; in, is the conjugate transpose of the antenna array steering vector, is the number of users, is the channel gain per unit distance.

4. The DRL-based hybrid RIS-assisted ISAC resource optimization method according to claim 2, characterized in that: The calculation of the secure transmission rates of users in the real environment and the mismatched environment in step S2-4 is specifically as follows: In real-world environments, users The safe transmission rate at is expressed as: ; Users in mismatched environments The safe transmission rate at is expressed as: ; in, For users The achievable rate, To perceive the target user The eavesdropping rate, For users in real environment The safe transmission rate at For users in mismatched environments The safe transmission rate at .

5. The DRL-based hybrid RIS-assisted ISAC resource optimization method according to claim 2, characterized in that: The optimization problem obtained in step S2-5 is specifically: Collaboratively formulate beamforming matrices and reflection coefficients to achieve confidentiality and speed, and optimize the problem in real-world environments: ; ; ; ; ; in, To ensure confidentiality and speed in real-world environments; is the SINR of the radar echo signal; represents the SINR threshold of the echo signal, represents the transmit power of the BS, is the power budget for active components, is the transmission power of the hybrid RIS active element, expressed as: ; in, Indicates the transmit power of the user's desired signal; is the identity matrix of K active elements; is the channel vector between the hybrid RIS and the base station BS; is the conjugate transpose of the channel vector between the hybrid RIS and the base station BS; The optimization problem under the mismatch environment is: ; ; ; ; ; in, For confidentiality and speed in mismatched environments.

6. The DRL-based hybrid RIS-assisted ISAC resource optimization method according to claim 1, characterized in that: In step S3, the deep directed intrinsic motivation exploration algorithm is integrated to improve the SAC reinforcement learning, specifically: Step S3-1: Modify the reward function used to train the agent: ; in, is the instantaneous reward calculated by the environment in the current state, is the average reward up to the current state; Step S3-2: Add an explorer network to predict the coefficients of the hybrid RIS ,in : ; Among them, when , , where the number of 1s refers to the transmit beamforming and AN beamforming The number of elements included in the action vector; each There are 2 items, measuring simultaneously The real and imaginary parts of Step S3-3: Explorer Network The prediction of is used to perturb the action: ; in, is the Hadamard product, the hyperparameter Constrain the explorer network to not make harmful perturbations to the actions chosen by the actors; Step S3-4: Modify the losses of the Q network and the actor network as follows: ; ; ; ; in, The explorer network is optimized to maximize the sum of the absolute values ​​of the TD errors: ; ; Step S3-5: Update the deterministic exploration network through the deterministic policy gradient algorithm: ; ; in, is the mathematical expectation calculation, is the timing difference error after motion disturbance, is the perturbation vector The gradient operator, is the state of the environment observed by the agent at time step t, is the original action chosen by the agent at time step t, The perturbation vector generated for the explorer network, The parameters are The Explorer Network, For the parameters The gradient operator, For Explorer Network Parameters, is the learning rate of the explorer network, Perception objective function of the explorer network.

Citation Information

Patent Citations

  • Combined beam forming design method for hybrid RIS-assisted millimeter wave ISAC system

    CN118118072A

  • Physical layer security optimization method of RIS-assisted ISAC system

    CN118540694A

  • RIS-assisted backscatter communication perception integration method

    CN116248173A

  • Communication-centered RIS-assisted cellular-free ISAC network joint beam forming method

    CN117375683A