A STAR-RIS-assisted joint optimization method for concealed synesthesia based on deep reinforcement learning
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-08
- Publication Date
- 2026-08-14
AI Technical Summary
然而,传统DRL算法(如DDPG、TD3、SAC)存在明显局限:一方面缺乏对通感系统特有结构(如波形耦合与安全约束)的针对性设计,导致学习效率低下;另一方面对动态无线环境适应性不足,在部分可观或非平稳信道中难以保持通信与感知性能的稳定协同
[0081](1)我们首先提出一种新型模型,用于在同时反射与投射可重构智能表面(STAR-RIS)辅助的通信与感知一体化(ISAC)系统中实现隐蔽波束成形优化,在对抗潜在威胁和干扰时,能够在复杂的环境中保持低可检测性,保证通信的安全性,同时使用非正交多址接入(NOMA)技术有效地提高通信的效率;
Smart Images

Figure CN122577932A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication technology, and in particular to a STAR-RIS-assisted integrated optimization method for covert sensing based on deep reinforcement learning. Background Technology
[0002] Integrated Sensing and Communication (ISAC) technology is becoming a key technology beyond fifth-generation (B5G) and sixth-generation (6G) networks, enabling high-precision sensing, seamless connectivity, and ubiquitous intelligence. It supports advanced applications such as autonomous driving, robot navigation, and the Internet of Things (IoT).
[0003] Despite the numerous attractive advantages of sensor-integrated technology, it still faces several challenges. Adverse propagation environments caused by obstacles and scattering can degrade sensor-integrated performance, especially in target perception. Unlike traditional reconfigurable smart surfaces, the STAR-RIS reconfigurable smart surface, which simultaneously reflects and refracts transmitted signals to users on both sides of the surface, extends service coverage from 180 degrees to 360 degrees. This enhances the system's adaptability to different wireless environments and improves the flexibility of signal propagation control. Furthermore, due to the broadcast nature of wireless networks, when confidential signals are embedded with a uniform waveform, they are easily detected or intercepted. This problem is further exacerbated when the target is a suspected eavesdropper—the non-cooperative or passive nature of eavesdroppers makes obtaining their channel state information extremely challenging, resulting in severe incompleteness of the eavesdropper's channel state information. However, the simultaneous transmit / receive reconfigurable smart surface can intelligently alter the electromagnetic properties of the incident signal by independently adjusting the amplitude and phase shift parameters of each unit. This capability can both enhance the effective signal at the target receiver and directionally reflect / refract interference signals to disrupt potential eavesdropping.
[0004] While traditional convex optimization methods have theoretical advantages in sensor-integrated systems, their high computational complexity makes them difficult to handle non-convex problems and large-scale real-time scenarios. This has prompted researchers to turn to deep reinforcement learning (DRL) in an attempt to circumvent complex computations through a data-driven approach. However, traditional DRL algorithms (such as DDPG, TD3, and SAC) have significant limitations: on the one hand, they lack specific designs for the unique structures of sensor-integrated systems (such as waveform coupling and security constraints), resulting in low learning efficiency; on the other hand, they are not adaptable enough to dynamic wireless environments, making it difficult to maintain stable coordination between communication and sensing performance in some observable or non-stationary channels. Therefore, designing DRL algorithms that combine adaptability, efficiency, and stability remains a core challenge. Summary of the Invention
[0005] The purpose of this invention is to provide a STAR-RIS-assisted integrated optimization method for concealed synesthesia based on deep reinforcement learning, so as to improve the problems existing in the prior art.
[0006] This invention is implemented as follows: a STAR-RIS-assisted joint optimization method for hidden synesthesia based on deep reinforcement learning, comprising the following steps:
[0007] S1: Establish a concealed sensing integrated system model based on the STAR-RIS reconfigurable smart surface that simultaneously reflects and projects. The concealed communication and sensing integrated system model includes a communication and sensing integrated base station model, a reconfigurable smart surface model, a channel model, a concealed communication signal model, and a sensing signal model. The communication and sensing integrated base station model is used to transmit dual-function communication and sensing signal waveforms. The reconfigurable smart surface model is used to overcome the obstruction of surrounding obstacles when the communication signal propagates. The channel model is used to establish the real communication channel in the system scenario. The concealed communication signal model is used to establish the communication signals received by concealed users and public users. The sensing signal model is used to establish the echo signal containing target information received by the base station.
[0008] S2: Based on the aforementioned integrated covert communication and sensing system model, an optimization problem is established under the constraints of communication, sensing, and non-orthogonal multiple access (NOMA) of the system. Specifically, the system's energy efficiency is used as the objective function, and the beamforming vector of the base station is used as the optimization objective. Reflection and projection phase of reconfigurable smart surfaces and the covariance matrix of the sensed signal To find the optimal solution for the optimized variables;
[0009] S3: Based on the optimization objective function and optimization variables, a deep reinforcement algorithm is proposed—a dual-delay deep deterministic policy gradient-first empirical replay algorithm based on a diffusion model. The algorithm constructs the optimization problem as a Markov decision process (MDP), takes the hardware-constrained reconfigurable smart surface STAR-RIS-assisted covert wireless network as the environment, and the central controller of the base station as the agent. Through continuous interaction between the agent and the environment, the system gradually learns the optimal policy to maximize the cumulative reward, thereby achieving joint optimization of beamforming, STAR-RIS phase shift, and sensing covariance.
[0010] S4: Based on the Markov Decision Process (MDP) in step S3, the agent obtains the current state from the environment. , will state The input is fed into a pre-trained policy network, which then calculates and outputs an action that it believes will maximize long-term reward in the current state. After the agent performs this action, it acts on the environment, and the agent enters the next decision cycle until convergence is achieved. When convergence occurs, it means that the optimal solution to the problem has been obtained, and the model of the concealed communication and perception integrated system is optimized.
[0011] 2. The STAR-RIS-assisted integrated covert sensing joint optimization method based on deep reinforcement learning according to claim 1, characterized in that, in step S1, the integrated communication and sensing base station model has hardware impairments at both the transceiver and receiver ends, including the equipment equipped with One transmitting antenna and A base station with one receiving antenna, a single-antenna public user, a single-antenna covert user, and a potential hostile guardian Willie, who is also a target of detection; the base station transmits a dual-function signal to communicate with the single-antenna public user Carol and the single-antenna covert user Bob via a reconfigurable smart surface, while simultaneously detecting the potential hostile guardian Willie, who intends to monitor the covert user's communications.
[0012] More preferably, in step S1, the specific expression for the reflection and projection phase matrix of the reconfigurable smart surface model is as follows:
[0013] ;
[0014] in, It is the first The reflection transmission amplitude coefficient of each component, Indicates the first Phase offset corresponding to each component The coefficient is used for transmission ( ) or reflection ( ), This indicates the number of components on the reconfigurable smart surface. This indicates that the vector elements inside the parentheses are extracted to form a new diagonal matrix, and j represents the imaginary unit;
[0015] Considering the practical hardware implementation of reconfigurable smart surfaces, their phase noise matrix is modeled using a von Mises distribution: ,in Represents lumped parameters, Indicates phase noise power;
[0016] The channel model includes the channel from a dual-function base station to a reconfigurable smart surface. Channel from reconfigurable smart surface to public user Carol Channel from reconfigurable smart surface to covert user Bob A channel from the reconfigurable smart surface to the hostile guardian Willie. and the echo channel from the reconfigurable smart surface to the base station Among them, the channels from the reconfigurable smart surface to the public user Carol, the covert user Bob, and the hostile guardian Willie are... as well as The channel model is a line-of-sight channel, the channel from the dual-function base station to the reconfigurable smart surface. and echo channel The channel is modeled as a Rayleigh channel;
[0017] The covert communication signal model includes a dual-function signal consisting of N symbols transmitted by a base station, where the base station transmits the signal in the Nth symbol series. The transmitted signal at each symbol is:
[0018] ;
[0019] in, Indicates that no covert signal is sent. The transmitted signal at each symbol Indicates the launch of public information The corresponding beamforming vector is given to the public user Carol. Indicates the first time when sending a covert signal The transmitted signal at each symbol Indicates the transmission of covert information The corresponding beamforming vector is given to the covert user Bob. This indicates that the dedicated sensing signal follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution, and parameter... It is the covariance of the complex Gaussian distribution of the dedicated sensing signal. This indicates that the hardware impairments at the base station follow a complex Gaussian distribution when no covert signals are transmitted, with parameter 0 being the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This means extracting the vector elements inside the parentheses to form a new diagonal matrix. This indicates that the hardware impairment at the base station when transmitting covert signals follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This represents the hardware impairment coefficient of the base station;
[0020] Hidden users in the first The received signal at each symbol is:
[0021] ;
[0022] in, This indicates that the integrated communication and sensing base station uses a composite channel between a reconfigurable smart surface and the concealed user Bob, where... This represents the channel from the reconfigurable smart surface to the covert user Bob. The conjugate transpose of . This represents the projected phase matrix of the reconfigurable smart surface. This represents the projected phase noise matrix of the reconfigurable smart surface. This represents the channel from a dual-function base station to a reconfigurable smart surface. Indicates the first The Gaussian white noise at each symbol follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This indicates that the hardware impairment at the location of the hidden user Bob follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. The mathematical expectation of the signal. Represents the covariance matrix of the transmitted signal;
[0023] The covariance matrix of the transmitted signal is expressed as:
[0024] ;
[0025] in, Indicates the launch of public information The corresponding beamforming vector is given to the public user Carol. Indicates the transmission of covert information The corresponding beamforming vector is given to the covert user Bob. It is the covariance of the complex Gaussian distribution of the dedicated sensing signal. This represents the hardware impairment coefficient of the base station. This means extracting the vector elements inside the parentheses to form a new diagonal matrix;
[0026] Public users in The received signal at each symbol is:
[0027] ;
[0028] in, This indicates that the integrated communication and sensing base station uses a composite channel between a reconfigurable smart surface and the public user Carol, where... This represents the channel from the reconfigurable smart surface to the public user Carol. The conjugate transpose of . This represents the reflection phase matrix of the reconfigurable smart surface. This represents the reflection phase noise matrix of the reconfigurable smart surface. This represents the channel from a dual-function base station to a reconfigurable smart surface. Indicates the first The Gaussian white noise at each symbol follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This indicates that the hardware impairment at the public user Carol's location follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution, where The mathematical expectation of the signal. Represents the covariance matrix of the transmitted signal;
[0029] The rival guardian Willie in the The received signal at each symbol and the total error detection probability are:
[0030] ;
[0031] ;
[0032] in, This indicates that the integrated communication and sensing base station uses a composite channel between the reconfigurable smart surface and the hostile guardian Willie, where... This represents the channel from the reconfigurable smart surface to the hostile guardian Willie. The conjugate transpose of . This represents the projected phase matrix of the reconfigurable smart surface. This represents the projected phase noise matrix of the reconfigurable smart surface. This represents the channel from a dual-function base station to a reconfigurable smart surface. Indicates the first The Gaussian white noise at each symbol follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This represents the total error detection probability. This represents the probability of a false positive. Indicates the probability of missed detection;
[0033] The sensing signal model includes the echo signal received by the base station, and its specific expression is as follows:
[0034] ;
[0035] in, Represents the echo channel from the reconfigurable smart surface to the base station. transpose, This represents the transpose of the projected phase noise matrix of the reconfigurable smart surface. This represents the channel coefficients related to the target's radar cross-section. This represents the turning vector of the hostile guardian Willie. This indicates Willie's azimuth. This represents the channel from a dual-function base station to a reconfigurable smart surface. Indicates the first time when sending a covert signal The transmitted signal at each symbol This indicates that the Gaussian white noise at the base station follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance matrix of a complex Gaussian distribution. This indicates that the hardware impairments at the base station receiver follow a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution, where The mathematical expectation of the signal. This represents the hardware impairment coefficient of the base station.
[0036] More preferably, in step S2, the specific expression of the optimization objective function is:
[0037]
[0038] in, This represents the total power consumption of the static circuitry at the base station. This indicates the power amplification efficiency at the base station. This indicates that the hidden user Bob is transmitting covert information. The achievable rate at that time This indicates that the public user Carol transmits hidden information. The achievable rate at that time;
[0039] The original problem of the optimization problem is specifically described as follows:
[0040] ;
[0041] Indicates the system's energy efficiency. Indicates the launch of public information The corresponding beamforming vector is given to the public user Carol. Indicates the transmission of covert information The corresponding beamforming vector is given to the covert user Bob. This indicates the maximum power threshold of the base station. This represents the threshold of the CRB (Crame-Rao Boundary). This represents the conjugate transpose of the composite channel between the integrated communication and sensing base station and the concealed user Bob via a reconfigurable smart surface. This indicates that the hidden user Bob is transmitting covert information. The achievable rate at that time This indicates that the user can decode the signal. , No. The transmission amplitude coefficient of each component Indicates the first The reflection transmission amplitude coefficient of each component, Indicates the first The phase shift corresponding to the transmission of each component Indicates the first The phase shift corresponding to the reflection of each component express The solution.
[0042] More preferably, in step S3, the specific process of the depth enhancement algorithm is as follows:
[0043] A Markov decision process model is constructed using the base station central controller as an agent and a covert wireless network assisted by a hardware-constrained reconfigurable smart surface (STAR-RIS) as the environment. The model includes states, actions, and immediate rewards. The agent observes current state information from the environment, including dynamic channel state information. The agent inputs this current state information into a pre-trained optimal policy network to obtain the current action, which includes the beamforming vector, STAR-RIS phase shift, and sensing covariance matrix to be optimized. The agent executes the action and receives immediate rewards from the environment, which are associated with the system's objective function and constraints. Through continuous interaction between the agent and the environment, the policy network is optimized based on the immediate rewards to maximize cumulative rewards, thereby achieving joint optimization of the beamforming vector, STAR-RIS phase shift matrix, and sensing covariance matrix.
[0044] More preferably, the action space in the Markov decision process Defined as the set of all possible actions that an agent can perform in any given state, the actions ∈ The joint control variables to be optimized include the beamforming vector of the base station, the discrete phase shift values of each unit of the simultaneously transmitted and projected reconfigurable smart surface, and the covariance matrix of the sensed waveform. The expression is as follows:
[0045] ;
[0046] in, Indicates the first The agent's actions during step training Indicates the first Public information is launched during training. The corresponding beamforming vector is given to the public user Carol. Indicates the first Launching covert information during training The corresponding beamforming vector is given to the covert user Bob. Indicates the first Dedicated sensing signals during step training Indicates the first The reflection transmission amplitude coefficient of the components during step training. No. The projection phase offset of the components during step training. No. The reflection phase shift of the components during step training;
[0047] The state space It consists of the set of all possible states of the environment, the states ∈ Satisfying the Markov property, including the system's dynamic channel state information and the agent's action in the previous time step, the expression is:
[0048] ;
[0049] in, Indicates the first -1 step of training: the agent's actions. Indicates the first The channel from the dual-function base station to the reconfigurable smart surface during step training. Indicates the first During training, the channel from the smart surface to the public user Carol can be reconstructed. Indicates the first During training, the channel from the smart surface to the hidden user Bob can be reconstructed. This represents the channel from the reconfigurable smart surface to the hostile guardian Willie. Indicates the first The echo channel from the smart surface to the base station can be reconstructed during step training;
[0050] The instant reward function is a mapping mechanism used to determine the state of an agent. Next action Then, the environment generates a scalar feedback signal. The expression is:
[0051] ;
[0052] in, Indicates the system's energy efficiency. This indicates that the hidden user Bob is transmitting covert information. The achievable rate at that time This indicates that the user can decode the signal. This indicates that the integrated communication and sensing base station uses the conjugate transpose of the composite channel between the reconfigurable smart surface and the hostile guardian Willie. express The solution, It is the covariance of the complex Gaussian distribution. It is a step function. The value is 1. The value is 0. This represents the coefficient of the penalty term.
[0053] More preferably, the dual-delay deep deterministic policy gradient-first experience replay algorithm based on the diffusion model replaces the actor network in the dual-delay deep deterministic policy gradient algorithm with a generative diffusion model, while employing priority experience replay to improve algorithm performance; the inverse process of the generative diffusion model involves a series of iterative denoising steps from the noise vector Reconstructing the original data sample Based on the original data sample distribution Select Action The approximate expression for the reverse process is:
[0054] ;
[0055] in, Follows a standard Gaussian distribution. This represents the mean of a standard Gaussian distribution. Indicates the first Noise samples at each diffusion step Indicates a time step. Indicates state, Denotes the predetermined variance coefficients, where , express Factors The cumulative product, Indicates the covariance scheduling factor;
[0056] Through gradual noise reduction, the original data sample Represented as:
[0057] ;
[0058] in, Represents a deep neural network. Indicates the first Noise samples at each diffusion step express Factors The cumulative product, Indicates state;
[0059] Deep Neural Networks In a given state The denoised noise is generated under the given conditions, and its mean is subsequently approximated indirectly:
[0060] ;
[0061] in, Indicates the noise mean. Represents a deep neural network. Indicates the first Noise samples at each diffusion step express Factors The cumulative product, Indicates state;
[0062] By reparameterizing to achieve differentiable sampling and stable backpropagation, the reverse process proceeds as follows: After noise reduction, it can be expressed as:
[0063] ;
[0064] in, Indicates the noise mean. Indicates the first Noise samples at each diffusion step Indicates state, It is the operator for the Hadamard product. This indicates that it follows a standard Gaussian distribution. It is the identity matrix. ,in express Factors The cumulative product, This represents the covariance scheduling factor.
[0065] More preferably, the gradient-first empirical replay algorithm for a dual-delay deep deterministic strategy based on a diffusion model includes the following steps:
[0066] S31: System initialization; initialize two online commentator networks. , An online actor network based on a generative diffusion model and the corresponding parameters Initialize the target network and its parameters corresponding to the online network. Simultaneously, initialize a priority experience replay buffer. ;
[0067] S32: Network Training and Motion Generation; Regarding Training Steps =1 to Perform the following sub-steps:
[0068] S321: Observe the current environmental status and initialize random variables. ,in Follows a standard Gaussian distribution. It is the identity matrix;
[0069] S322: Diffusion denoising process; for the number of diffusion steps = Step 1, then perform the following sub-steps:
[0070] S3221: Generating a denoised distribution using a deep neural network ,in Represented as conditional information;
[0071] S3222: Calculate the mean of a Gaussian distribution. and ;
[0072] S33: Action selection and execution; calculation And based on this, randomly select an action. Perform the selected action in the environment. Observe the next state and instant rewards ;
[0073] S34: Experience storage and extraction; storing and extracting experience tuples Store in priority playback buffer And set its priority to the current maximum priority, from the priority experience PER buffer. Extract a batch of empirical data from D;
[0074] S35: Experience priority update; for experience = Calculate up to D, the first... Sampling probability of a piece of experience ,in Representing experience priority, Indicates the capacity of the experience buffer. This is an introduced positive constant used to ensure that the probability of each experience being sampled is not zero. Prioritize experience replay factors; calculate experience weights and update experience priorities. Each experience tuple is then weighted according to the absolute value of its temporal error (TD). Weights are assigned, and samples are then drawn in order of weight. The weights are defined as follows: Weighted normalization representation , Indicates hyperparameters, Indicates the first Sampling probability of a piece of experience This represents the maximum value among the K weights;
[0075] Compensation for the high variance caused by importance sampling weights; the first The timing error (TD) of the empirical data. Represented as ,in , This refers to an online network of critics. An online actor network representing a generative diffusion model. This indicates the corresponding online network parameters; This represents the target network parameters corresponding to the online network, and the noise. Indicates noise It follows a Gaussian distribution, where the parameter 0 is the mean. Let Variance be the variance, and its value be [value]. Indicates the discount factor;
[0076] S36: Update the online critic network parameters; based on the loss function. Update the parameters of the online critic network, including Indicates the normalized weights. , This refers to an online network of critics. An online actor network representing a generative diffusion model. This indicates the corresponding online network parameters; This indicates the target network parameters corresponding to the online network;
[0077] S37: Update online actor network parameters; based on the loss function. Update online actor network parameters;
[0078] S38: Soft update target network parameters; based on soft update , , Update the target network parameters, where Indicates the soft update weight. This indicates the corresponding online network parameters; This represents the target network parameters corresponding to the online network.
[0079] More preferably, in step S4, the agent is trained through continuous interaction with the environment, and after multiple iterations, the policy network of the deep reinforcement learning algorithm converges; the convergence of the policy network specifically refers to the convergence of the overall energy efficiency index of the concealed sensory integrated system to a stable value.
[0080] Compared with the prior art, the present invention has the following advantages:
[0081] (1) We first propose a novel model for achieving covert beamforming optimization in a communication and sensing integrated (ISAC) system assisted by simultaneous reflection and projection reconfigurable smart surfaces (STAR-RIS). This model can maintain low detectability in complex environments while countering potential threats and interference, ensuring communication security. At the same time, it effectively improves communication efficiency by using non-orthogonal multiple access (NOMA) technology.
[0082] (2) We propose an improved dual-delay deep deterministic policy gradient (TD3) deep reinforcement learning algorithm to address the inherent conflict between perception, communication, and computation metrics within a system exacerbated by hardware impairment. Through joint optimization and intelligent decision-making, the proposed method significantly reduces the probability of communication interruption while achieving an overall improvement in the system's comprehensive performance. Attached Figure Description
[0083] Figure 1 This is a flowchart of a method according to an embodiment of the present invention;
[0084] Figure 2 This is the process of the double-delay deep deterministic policy gradient (TD3) priority experience replay algorithm based on the diffusion model in an embodiment of the present invention.
[0085] Figure 3 This is a comparison chart of the convergence results and algorithms of the embodiments of the present invention;
[0086] Figure 4 The number of STAR-RIS components M and the maximum power threshold of the base station are different in different embodiments of the present invention. Below is a simulation diagram showing the relationship between energy efficiency (EE) and the CRB threshold Γ.
[0087] Figure 5 Different numbers of hardware damage coefficients are given in the embodiments of the present invention. With noise intensity index Below is a simulation diagram showing the relationship between energy efficiency (EE) and the number of STAR-RIS components;
[0088] Figure 6 The objective function energy efficiency (EE) constructed for embodiments of the present invention is used to measure the total power consumption of different base station static circuits. and power amplification efficiency at the base station Simulation diagram of the relationship between energy efficiency (EE) and concealment factor (o) under the given conditions;
[0089] Figure 7 The objective function constructed for the embodiments of the present invention relates to the energy efficiency (EE) and the total power consumption of the base station static circuit under the assistance of orthogonal multiple access (NOMA) and non-NOMA. The relationship simulation diagram. Detailed Implementation
[0090] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Unless otherwise defined, the technical or scientific terms used herein should have the ordinary meaning understood by those skilled in the art. The terms "comprising" and similar expressions used herein mean that the element or object preceding the word covers the element or object listed after the word and its equivalents, but does not exclude other elements or objects.
[0091] Example 1: This example provides a joint optimization method for escaping synesthesia based on simultaneous reflection and projection reconfigurable smart surfaces (STAR-RIS) assisted by deep reinforcement learning (DRL), such as... Figure 1 As shown, it includes the following steps:
[0092] S1: Establish a concealed sensing integrated system model based on the STAR-RIS reconfigurable smart surface that simultaneously reflects and projects. The concealed communication and sensing integrated system model includes a communication and sensing integrated base station model, a reconfigurable smart surface model, a channel model, a concealed communication signal model, and a sensing signal model. The communication and sensing integrated base station model is used to transmit communication and sensing dual-function signal waveforms. The reconfigurable smart surface model is used to overcome the obstruction of surrounding obstacles when the communication signal propagates. The channel model is used to establish the real communication channel in the system scenario. The concealed communication signal model is used to establish the communication signals received by concealed users and public users. The sensing signal model is used to establish the echo signal containing target information received by the base station.
[0093] In some embodiments, the integrated communication and sensing base station model has hardware impairments (HWIs) at both the transceiver and receiver ends, including those equipped with... One transmitting antenna and A base station with one receiving antenna, a single-antenna public user, a single-antenna covert user, and a potential hostile guardian Willie, who is also a target of detection; the base station transmits a dual-function signal to communicate with the single-antenna public user Carol and the single-antenna covert user Bob via a reconfigurable smart surface, while simultaneously detecting the potential hostile guardian Willie, who intends to monitor the covert user's communications.
[0094] In some embodiments, the specific expressions for the reflection and projection phase matrices of the reconfigurable smart surface model are as follows:
[0095] ;
[0096] in, It is the first The reflection transmission amplitude coefficient of each component, Indicates the first Phase offset corresponding to each component The coefficient is used for transmission ( ) or reflection ( ), This indicates the number of components on the reconfigurable smart surface. This indicates that the vector elements inside the parentheses are extracted to form a new diagonal matrix, and j represents the imaginary unit.
[0097] In some embodiments, considering the actual hardware implementation of the reconfigurable smart surface, its phase noise matrix can be modeled using the following model:
[0098] ;
[0099] Specifically, the phase noise of the reconfigurable smart surface is modeled using the von Mises distribution: The phase noise of each component is ,in Represents lumped parameters, This represents the phase noise power.
[0100] In some embodiments, the channel model includes a channel from a dual-function base station to a reconfigurable smart surface. Channel from reconfigurable smart surface to public user Carol Channel from reconfigurable smart surface to covert user Bob A channel from the reconfigurable smart surface to the hostile guardian Willie. and the echo channel from the reconfigurable smart surface to the base station Among them, the channels from the reconfigurable smart surface to the public user Carol, the covert user Bob, and the hostile guardian Willie are... as well as The channel model is a line-of-sight channel, the channel from the dual-function base station to the reconfigurable smart surface. and echo channel The channel is modeled as a Rayleigh channel;
[0101] In some embodiments, the covert communication signal model includes a dual-function signal consisting of N symbols transmitted by a base station, wherein the base station in the... The transmitted signal at each symbol is:
[0102] ;
[0103] in, Indicates that no covert signal is sent. The transmitted signal at each symbol Indicates the launch of public information The corresponding beamforming vector is given to the public user Carol. Indicates the first time when sending a covert signal The transmitted signal at each symbol Indicates the transmission of covert information The corresponding beamforming vector is given to the covert user Bob. This indicates that the dedicated sensing signal follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution, and parameter... It is the covariance of the complex Gaussian distribution of the dedicated sensing signal. This indicates that the hardware impairments at the base station follow a complex Gaussian distribution when no covert signals are transmitted, with parameter 0 being the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This means extracting the vector elements inside the parentheses to form a new diagonal matrix. This indicates that the hardware impairment at the base station when transmitting covert signals follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This represents the hardware impairment coefficient of the base station;
[0104] Hidden users in the first The received signal at each symbol is:
[0105] ;
[0106] in, This indicates that the integrated communication and sensing base station uses a composite channel between a reconfigurable smart surface and the concealed user Bob, where... This represents the channel from the reconfigurable smart surface to the covert user Bob. The conjugate transpose of . This represents the projected phase matrix of the reconfigurable smart surface. This represents the projected phase noise matrix of the reconfigurable smart surface. This represents the channel from a dual-function base station to a reconfigurable smart surface. Indicates the first The Gaussian white noise at each symbol follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This indicates that the hardware impairment at the location of the hidden user Bob follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. The mathematical expectation of the signal. Represents the covariance matrix of the transmitted signal;
[0107] The covariance matrix of the transmitted signal is expressed as:
[0108] ;
[0109] in, Indicates the launch of public information The corresponding beamforming vector is given to the public user Carol. Indicates the transmission of covert information The corresponding beamforming vector is given to the covert user Bob. It is the covariance of the complex Gaussian distribution of the dedicated sensing signal. This represents the hardware impairment coefficient of the base station. This means extracting the vector elements inside the parentheses to form a new diagonal matrix.
[0110] Public users in The received signal at each symbol is:
[0111] ;
[0112] in, This indicates that the integrated communication and sensing base station uses a composite channel between a reconfigurable smart surface and the public user Carol, where... This represents the channel from the reconfigurable smart surface to the public user Carol. The conjugate transpose of . This represents the reflection phase matrix of the reconfigurable smart surface. This represents the reflection phase noise matrix of the reconfigurable smart surface. This represents the channel from a dual-function base station to a reconfigurable smart surface. Indicates the first The Gaussian white noise at each symbol follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This indicates that the hardware impairment at the public user Carol's location follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution, where The mathematical expectation of the signal. Represents the covariance matrix of the transmitted signal;
[0113] Assume the channel gain of the two users is as follows: The order of the users (i.e., Bob, the hidden user, is the near-end user, and Carol, the public user, is the remote user) is as follows: This indicates that the integrated communication and sensing base station uses a composite channel between a reconfigurable smart surface and the concealed user Bob. This indicates that the integrated communication and sensing base station, through a composite channel between the reconfigurable smart surface and the public user Carol, allocates greater transmission power to the signal of the remote user. ,in Indicates the launch of public information The corresponding beamforming vector is given to the public user Carol. Indicates the transmission of covert information The appropriate beamforming vector is given to the covert user Bob. According to Non-Orthogonal Multiple Access (NOMA) technology, covert user Bob will perform Step-by-Step Interference Cancellation (SIC) operations. This process first decodes the signal destined for public user Carol. This signal is then subtracted from the composite received signal in order to decode the signal of the covert user Bob himself. Meanwhile, Carol directly decodes its signal. It treats the signals of all other users as interference;
[0114] The hidden user Bob transmits hidden information. The achievable rate at that time is:
[0115] ;
[0116] The hidden user Bob transmits public information. The achievable rate at that time is:
[0117] ;
[0118] The public user Carol transmits public information. The achievable rate at that time is:
[0119] ;
[0120] To ensure that all users can decode the signal Conditions must be met According to the Non-Orthogonal Multiple Access (NOMA) criterion, the decoded signal... The achievable rate allows the covert user Bob to transmit public information. The achievable rate and public information transmitted by the public user Carol Only the minimum achievable speed at a given time can enable effective signal transmission.
[0121] The rival guardian Willie in the The received signal at each symbol and the total error detection probability are:
[0122] ;
[0123] ;
[0124] in, This indicates that the integrated communication and sensing base station uses a composite channel between the reconfigurable smart surface and the hostile guardian Willie, where... This represents the channel from the reconfigurable smart surface to the hostile guardian Willie. The conjugate transpose of . This represents the projected phase matrix of the reconfigurable smart surface. This represents the projected phase noise matrix of the reconfigurable smart surface. This represents the channel from a dual-function base station to a reconfigurable smart surface. Indicates the first The Gaussian white noise at each symbol follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This represents the total error detection probability. This represents the probability of a false positive. Indicates the probability of missed detection;
[0125] Therefore, the hidden constraints of the system can be expressed as:
[0126] ;
[0127] in, This represents the concealment constant.
[0128] In some embodiments, the sensing signal model includes the echo signal received by the base station, specifically expressed as:
[0129] ;
[0130] in, Represents the echo channel from the reconfigurable smart surface to the base station. transpose, This represents the transpose of the projected phase noise matrix of the reconfigurable smart surface. This represents the channel coefficients related to the target's radar cross-section. This represents the turning vector of the hostile guardian Willie. This indicates Willie's azimuth. This represents the channel from a dual-function base station to a reconfigurable smart surface. Indicates the first time when sending a covert signal The transmitted signal at each symbol This indicates that the Gaussian white noise at the base station follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance matrix of a complex Gaussian distribution. This indicates that the hardware impairments at the base station receiver follow a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution, where The mathematical expectation of the signal. This represents the hardware impairment coefficient of the base station;
[0131] In the field of radar sensing, the Cramer-Rao bound (CRB) provides a lower bound for the variance of unbiased estimators and is widely used as a performance evaluation metric. This invention, for the first time, derives the theoretical expression of the Cramer-Rao bound under hardware constraints, quantifying its impact on sensing performance. The specific expression is as follows:
[0132]
[0133] in, Represents the Fisher Information Matrix (FIM). This represents the element in the first row and first column of the Fisher Information Matrix (FIM) inverse.
[0134] The Fisher Information Matrix (FIM) is specifically expressed as follows:
[0135]
[0136] in This represents the channel coefficients related to the target's radar cross-section. This indicates Willie's azimuth. express The real part, express The imaginary part;
[0137] Each element of the Fisher Information Matrix (FIM) is represented as:
[0138]
[0139] in, Representing the Fisher information matrix OK Column elements, tr This indicates taking the trace of the matrix inside the parentheses. This indicates that the number inside the parentheses is the real part. This represents the channel coefficients related to the target's radar cross-section. The inverse of the noise covariance matrix is given by the noise covariance matrix. Represented as:
[0140] ;
[0141] in , , Represents the echo channel from the reconfigurable smart surface to the base station. transpose, This represents the conjugate of the channel coefficients related to the target's radar cross-section. This means extracting the vector elements inside the parentheses to form a new diagonal matrix. Indicates the direction steering vector of the signal. The covariance matrix of the transmitted signal is represented. It represents the Kronecker product.
[0142] S2: Based on the aforementioned integrated covert communication and sensing system model, an optimization problem is established under the constraints of communication, sensing, and non-orthogonal multiple access (NOMA) of the system. Specifically, the system's energy efficiency is used as the objective function, and the beamforming vector of the base station is used as the optimization objective. Reflection and projection phase of reconfigurable smart surfaces and the covariance matrix of the sensed signal To find the optimal solution for the optimized variables.
[0143] In some embodiments, the specific expression of the optimization objective function is:
[0144]
[0145] in, This represents the total power consumption of the static circuitry at the base station. This indicates the power amplification efficiency at the base station. This indicates that the hidden user Bob is transmitting covert information. The achievable rate at that time This indicates that the public user Carol transmits hidden information. The achievable rate at that time;
[0146] The original problem of the optimization problem is specifically described as follows:
[0147] ;
[0148] in, Indicates the system's energy efficiency. Indicates the launch of public information The corresponding beamforming vector is given to the public user Carol. Indicates the transmission of covert information The corresponding beamforming vector is given to the covert user Bob. This indicates the maximum power threshold of the base station. This represents the threshold of the CRB (Crame-Rao Boundary). This represents the conjugate transpose of the composite channel between the integrated communication and sensing base station and the concealed user Bob via a reconfigurable smart surface. This indicates that the hidden user Bob is transmitting covert information. The achievable rate at that time This indicates that the user can decode the signal. , No. The transmission amplitude coefficient of each component Indicates the first The reflection transmission amplitude coefficient of each component, Indicates the first The phase shift corresponding to the transmission of each component Indicates the first The phase shift corresponding to the reflection of each component express The solution.
[0149] S3: Based on the optimization objective function and optimization variables, a deep reinforcement algorithm is proposed—the Diffusion Model-Based Dual-Delay Deep Deterministic Policy Gradient (TD3) Priority Experience Replay Algorithm. The algorithm constructs the optimization problem as a Markov Decision Process (MDP), takes the hardware-constrained reconfigurable smart surface STAR-RIS-assisted covert wireless network as the environment, and the central controller of the base station as the agent. Through continuous interaction between the agent and the environment, the system gradually learns the optimal policy to maximize the cumulative reward, thereby achieving joint optimization of beamforming, STAR-RIS phase shift, and sensing covariance.
[0150] In some embodiments, the specific process of the deep reinforcement algorithm is as follows: The base station central controller is used as an agent, and a covert wireless network containing hardware-constrained STAR-RIS is used as the environment. A Markov decision process model is constructed, including states, actions, and immediate rewards. The agent observes current state information from the environment, including dynamic channel state information. The agent inputs the current state information into a pre-trained optimal policy network to obtain the current action, which includes the beamforming vector, STAR-RIS phase shift, and perceptual covariance matrix to be optimized. The agent executes the action and receives immediate rewards from the environment, which are associated with the system's objective function and constraints. Through continuous interaction between the agent and the environment, the policy network is optimized based on the immediate rewards to maximize cumulative rewards, thereby achieving joint optimization of the beamforming vector, STAR-RIS phase shift matrix, and perceptual covariance matrix.
[0151] In some embodiments, the Markov decision process specifically includes: an action space, a state space, and an immediate reward function; the action space in the Markov decision process... Defined as the set of all possible actions that an agent can perform in any given state, the actions ∈ The agent is in the state according to the current policy. The actions taken directly affect the environment, triggering state transitions. In this embodiment, the actions are specifically the joint control variables to be optimized, including the beamforming vector of the base station, the discrete phase shift values of each unit of the simultaneously transmitted and projected reconfigurable smart surface, and the covariance matrix of the sensed waveform, expressed as:
[0152] ;
[0153] in, Indicates the first The agent's actions during step training Indicates the first Public information is launched during training. The corresponding beamforming vector is given to the public user Carol. Indicates the first Launching covert information during training The corresponding beamforming vector is given to the covert user Bob. Indicates the first Dedicated sensing signals during step training Indicates the first The reflection transmission amplitude coefficient of the components during step training. No. The projection phase offset of the components during step training. No. The reflection phase shift of the components during step training.
[0154] The state space It consists of the set of all possible states of the environment, the states ∈ At any moment The complete state of the environment is characterized, which satisfies the Markov property, i.e., the next state. The probability distribution depends only on the current state. With the selected action Regardless of the historical state sequence, in this embodiment, the state information preferably includes, but is not limited to: the system's dynamic channel state information and the action performed by the agent in the previous moment, expressed as:
[0155] ;
[0156] in, Indicates the first -1 step of training: the agent's actions. Indicates the first The channel from the dual-function base station to the reconfigurable smart surface during step training. Indicates the first During training, the channel from the smart surface to the public user Carol can be reconstructed. Indicates the first During training, the channel from the smart surface to the hidden user Bob can be reconstructed. This represents the channel from the reconfigurable smart surface to the hostile guardian Willie. Indicates the first The echo channel from the smart surface to the base station can be reconstructed during step training;
[0157] The instant reward function is a mapping mechanism used to determine the state of an agent. Next action Then, the environment generates a scalar feedback signal. This immediate reward is used to quantitatively evaluate the merits of the action in a specific state and is a key signal guiding the agent to learn the optimal strategy. Its expression is:
[0158] ;
[0159] in, Indicates the system's energy efficiency. This indicates that the hidden user Bob is transmitting covert information. The achievable rate at that time This indicates that the user can decode the signal. This indicates that the integrated communication and sensing base station uses the conjugate transpose of the composite channel between the reconfigurable smart surface and the hostile guardian Willie. express The solution, It is the covariance of the complex Gaussian distribution. It is a step function. The value is 1. The value is 0. This represents the coefficient of the penalty term.
[0160] In some embodiments, the dual-delay deep deterministic policy gradient-first empirical replay algorithm based on the diffusion model replaces the actor network in the dual-delay deep deterministic policy gradient algorithm with a generative diffusion model, while employing priority empirical replay to improve algorithm performance. The generative diffusion model adopts a two-stage working mode: first, Gaussian noise is gradually injected through a forward process to transform the original data into pure noise; then, iterative denoising is performed through a reverse process to finally reconstruct the original data. The algorithm in this embodiment only involves the reverse process. The reverse process of the generative diffusion model involves a series of iterative denoising steps from the noise vector... Reconstructing the original data sample Based on the original data sample distribution Select Action Due to the distribution of real conditions Since direct calculation is not feasible due to the unknown data distribution, we use a parametric model for approximation. The approximate expression for the reverse process is:
[0161] ;
[0162] in, Follows a standard Gaussian distribution. This represents the mean of a standard Gaussian distribution. Indicates the first Noise samples at each diffusion step Indicates a time step. Indicates state, Denotes the predetermined variance coefficients, where , express Factors The cumulative product, Indicates the covariance scheduling factor;
[0163] Through gradual noise reduction, the original data sample Represented as:
[0164] ;
[0165] in, Represents a deep neural network. Indicates the first Noise samples at each diffusion step express Factors The cumulative product, Indicates state;
[0166] Deep Neural Networks In a given state The denoised noise is generated under the given conditions, and its mean is subsequently approximated indirectly:
[0167] ;
[0168] in, Indicates the noise mean. Represents a deep neural network. Indicates the first Noise samples at each diffusion step express Factors The cumulative product, Indicates state;
[0169] By reparameterizing to achieve differentiable sampling and stable backpropagation, the reverse process proceeds as follows: After noise reduction, it can be expressed as:
[0170] ;
[0171] in, Indicates the noise mean. Indicates the first Noise samples at each diffusion step Indicates state, It is the operator for the Hadamard product. This indicates that it follows a standard Gaussian distribution. It is the identity matrix. ,in express Factors The cumulative product, This represents the covariance scheduling factor.
[0172] In some embodiments, the gradient-first empirical replay algorithm based on a diffusion model and a dual-delay deep deterministic strategy is used. Figure 2 The process of the deep reinforcement learning algorithm is described, including the following steps:
[0173] S31: System initialization; initialize two online commentator networks. , An online actor network based on a generative diffusion model and the corresponding parameters Initialize the target network and its parameters corresponding to the online network. Simultaneously, initialize a priority experience replay buffer. ;
[0174] S32: Network Training and Motion Generation; Regarding Training Steps =1 to Perform the following sub-steps:
[0175] S321: Observe the current environmental status and initialize random variables. ,in Follows a standard Gaussian distribution. It is the identity matrix;
[0176] S322: Diffusion denoising process; for the number of diffusion steps = Step 1, then perform the following sub-steps:
[0177] S3221: Generating a denoised distribution using a deep neural network ,in Represented as conditional information;
[0178] S3222: Calculate the mean of a Gaussian distribution. and ;
[0179] S33: Action selection and execution; calculation And based on this, randomly select an action. Perform the selected action in the environment. Observe the next state and instant rewards ;
[0180] S34: Experience storage and extraction; storing and extracting experience tuples Store in priority playback buffer And set its priority to the current maximum priority, from the priority experience PER buffer. Extract a batch of empirical data from D;
[0181] S35: Experience priority update; for experience = Calculate up to D, the first... Sampling probability of a piece of experience ,in Representing experience priority, Indicates the capacity of the experience buffer. This is an introduced positive constant used to ensure that the probability of each experience being sampled is not zero. Prioritize experience replay factors; calculate experience weights and update experience priorities. Each experience tuple is then weighted according to the absolute value of its temporal error (TD). Weights are assigned, and samples are then drawn in order of weight. The weights are defined as follows: Weighted normalization representation , Indicates hyperparameters, Indicates the first Sampling probability of a piece of experience This represents the maximum value among the K weights;
[0182] Compensation for the high variance caused by importance sampling weights; the first The timing error (TD) of the empirical data. Represented as ,in , This refers to an online network of critics. An online actor network representing a generative diffusion model. This indicates the corresponding online network parameters; This represents the target network parameters corresponding to the online network, and the noise. Indicates noise It follows a Gaussian distribution, where the parameter 0 is the mean. Let Variance be the variance, and its value be [value]. Indicates the discount factor;
[0183] S36: Update the online critic network parameters; based on the loss function. Update the parameters of the online critic network, including Indicates the normalized weights. , This refers to an online network of critics. An online actor network representing a generative diffusion model. This indicates the corresponding online network parameters; This indicates the target network parameters corresponding to the online network;
[0184] S37: Update online actor network parameters; based on the loss function. Update online actor network parameters;
[0185] S38: Soft update target network parameters; based on soft update , , Update the target network parameters, where Indicates the soft update weight. This indicates the corresponding online network parameters; This represents the target network parameters corresponding to the online network.
[0186] S4: Based on the Markov Decision Process (MDP) in step S3, the agent obtains the current state from the environment. , will state The input is fed into a pre-trained policy network, which then calculates and outputs an action that it believes will maximize long-term reward in the current state. After the agent performs this action, it acts on the environment, and the agent enters the next decision cycle until convergence is achieved. When convergence occurs, it means that the optimal solution to the problem has been obtained, and the model of the concealed communication and perception integrated system is optimized.
[0187] In some embodiments, in step S4, the agent is trained through continuous interaction with the environment, and after multiple iterations, the policy network of the deep reinforcement learning algorithm converges; the convergence of the policy network specifically refers to the convergence of the overall energy efficiency index of the covert sensory integrated system to a stable value.
[0188] Example 2: Based on Example 1, this example further illustrates the effects of the present invention through simulation.
[0189] (1) Simulation parameters;
[0190] System simulation parameters: Assume the ISAC base station is equipped with The transmitting antenna and The power budget of the base station is set to the receiving antenna. The number of STAR-RIS components is set to noise power variance The number of signal characters is set to The concealment constant is set to Hardware impairment coefficient at the base station transceiver and user end Total power consumption of base station static circuit Power amplification efficiency at the base station Simultaneously, the phase noise lumped parameters at the STAR-RIS smart surface can be reconstructed through reflection and projection. User communication rate threshold ;
[0191] Algorithm training simulation parameters: Batch size Discount factor Experience buffer capacity diffusion steps Soft update weight The learning rate for both the actor network and the critic network is 0.0001, and the penalty term coefficient is... Priority experience replay factor .
[0192] (2) Simulation results and data analysis;
[0193] Figure 3 The proposed PER-GDMTD3 algorithm was compared with two learning-based benchmark methods: TD3 and DDPG. Experimental results show that the proposed method significantly outperforms all competing algorithms in terms of the cumulative average reward. Furthermore, its convergence speed is also faster than other TD3 base schemes. This performance advantage stems from two key innovations: integrating a generative diffusion model and introducing the PER mechanism.
[0194] Figure 4 This indicates that as the sensing CRB threshold increases, the system's energy efficiency performance shows an upward trend. The mechanism behind this phenomenon is that a higher CRB threshold relaxes the sensing accuracy constraints, thereby reducing the power allocated to sensing and increasing the degrees of freedom (DoFs) available for optimizing covert beamforming. These additional DoFs contribute to achieving higher total rate, ultimately driving improvements in system energy efficiency performance.
[0195] Figure 5 Increasing the number of STAR-RIS units (M) improves system energy efficiency, primarily due to the precise phase control achieved through more reflective units, which enhances accurate beam focusing capabilities. Furthermore, system energy efficiency increases with the hardware impairment factor. It increases and then decreases monotonically. An increase in STAR-RIS phase noise exacerbates signal distortion in base stations and user equipment, forcing base stations to increase transmission power to mitigate signal degradation and maintain reliable communication and target awareness. Simultaneously, as STAR-RIS phase noise (a noise intensity index) increases... (The lower the phase noise, the greater the noise), and energy efficiency drops sharply as phase noise increases.
[0196] Figure 6 This indicates that the system's energy efficiency increases with the increase of the o value. This is because when the o value increases, the concealment constraint becomes more relaxed, leading to an increase in the transmission rate, thereby improving energy efficiency.
[0197] Figure 7 This demonstrates that, in this invention, the Non-Orthogonal Multiple Access (NOMA) scheme is superior to the Orthogonal Multiple Access (OMA) scheme in terms of energy efficiency. This is because, compared to the OMA scheme, the NOMA protocol enables simultaneous service for multiple users.
[0198] The simulation results above demonstrate that adding non-orthogonal multiple access (NOMA) assistance to concealed sensing integration can effectively improve system security and user communication quality. Meanwhile, the proposed dual-delay deep deterministic strategy gradient (TD3) priority experience replay algorithm (PER-GDMTD3) based on diffusion model has a good effect on handling the optimization problem constructed in this invention.
[0199] While embodiments of the present invention have been described in detail above, it will be apparent to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations fall within the scope and spirit of the invention as set forth in the claims. Furthermore, the invention described herein may have other embodiments and can be implemented or carried out in various ways.
Claims
1. A STAR-RIS-assisted joint optimization method for concealed synesthesia based on deep reinforcement learning, characterized in that, Includes the following steps: S1: Establish a concealed sensing integrated system model based on the STAR-RIS reconfigurable smart surface that simultaneously reflects and projects. The concealed communication and sensing integrated system model includes a communication and sensing integrated base station model, a reconfigurable smart surface model, a channel model, a concealed communication signal model, and a sensing signal model. The communication and sensing integrated base station model is used to transmit dual-function communication and sensing signal waveforms. The reconfigurable smart surface model is used to overcome the obstruction of surrounding obstacles when the communication signal propagates. The channel model is used to establish the real communication channel in the system scenario. The concealed communication signal model is used to establish the communication signals received by concealed users and public users. The sensing signal model is used to establish the echo signal containing target information received by the base station. S2: Based on the aforementioned integrated covert communication and sensing system model, an optimization problem is established under the constraints of communication, sensing, and non-orthogonal multiple access (NOMA) of the system. Specifically, the system's energy efficiency is used as the objective function, and the beamforming vector of the base station is used as the optimization objective. Reflection and projection phase of reconfigurable smart surfaces and the covariance matrix of the sensed signal To find the optimal solution for the optimized variables; S3: Based on the optimization objective function and optimization variables, a deep reinforcement algorithm is proposed—a dual-delay deep deterministic policy gradient-first empirical replay algorithm based on a diffusion model. The algorithm constructs the optimization problem as a Markov decision process (MDP), takes the hardware-constrained reconfigurable smart surface STAR-RIS-assisted covert wireless network as the environment, and the central controller of the base station as the agent. Through continuous interaction between the agent and the environment, the system gradually learns the optimal policy to maximize the cumulative reward, thereby achieving joint optimization of beamforming, STAR-RIS phase shift, and sensing covariance. S4: Based on the Markov Decision Process (MDP) in step S3, the agent obtains the current state from the environment. , will state The input is fed into a pre-trained policy network, which then calculates and outputs an action that it believes will maximize long-term reward in the current state. After the agent performs this action, it acts on the environment, and the agent enters the next decision cycle until convergence is achieved. When convergence occurs, it means that the optimal solution to the problem has been obtained, and the model of the concealed communication and perception integrated system is optimized.
2. The STAR-RIS-assisted integrated optimization method for hidden synesthesia based on deep reinforcement learning as described in claim 1, characterized in that, In step S1, the integrated communication and sensing base station model has hardware defects at both the transceiver and receiver ends, including those equipped with... One transmitting antenna and A base station with one receiving antenna, a single-antenna public user, a single-antenna covert user, and a potential hostile guardian Willie, who is also a target of detection; the base station transmits a dual-function signal to communicate with the single-antenna public user Carol and the single-antenna covert user Bob via a reconfigurable smart surface, while simultaneously detecting the potential hostile guardian Willie, who intends to monitor the covert user's communications.
3. The STAR-RIS-assisted integrated optimization method for concealed synesthesia based on deep reinforcement learning according to claim 2, characterized in that, In step S1, the specific expressions for the reflection and projection phase matrices of the reconfigurable smart surface model are as follows: ; in, It is the first The reflection transmission amplitude coefficient of each component, Indicates the first Phase offset corresponding to each component The coefficient is used for transmission ( ) or reflection ( ), This indicates the number of components on the reconfigurable smart surface. This indicates that the vector elements inside the parentheses are extracted to form a new diagonal matrix, and j represents the imaginary unit; Considering the practical hardware implementation of reconfigurable smart surfaces, their phase noise matrix is modeled using a von Mises distribution: ,in Represents lumped parameters, Indicates phase noise power; The channel model includes the channel from a dual-function base station to a reconfigurable smart surface. Channel from reconfigurable smart surface to public user Carol Channel from reconfigurable smart surface to covert user Bob A channel from the reconfigurable smart surface to the hostile guardian Willie. and the echo channel from the reconfigurable smart surface to the base station Among them, the channels from the reconfigurable smart surface to the public user Carol, the covert user Bob, and the hostile guardian Willie are... as well as The channel model is a line-of-sight channel, the channel from the dual-function base station to the reconfigurable smart surface. and echo channel The channel is modeled as a Rayleigh channel; The covert communication signal model includes a dual-function signal consisting of N symbols transmitted by a base station, where the base station transmits the signal in the Nth symbol series. The transmitted signal at each symbol is: ; in, Indicates that no covert signal is sent. The transmitted signal at each symbol Indicates the launch of public information The corresponding beamforming vector is given to the public user Carol. Indicates the first time when sending a covert signal The transmitted signal at each symbol Indicates the transmission of covert information The corresponding beamforming vector is given to the covert user Bob. This indicates that the dedicated sensing signal follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution, and parameter... It is the covariance of the complex Gaussian distribution of the dedicated sensing signal. This indicates that the hardware impairments at the base station follow a complex Gaussian distribution when no covert signals are transmitted, with parameter 0 being the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This means extracting the vector elements inside the parentheses to form a new diagonal matrix. This indicates that the hardware impairment at the base station when transmitting covert signals follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This represents the hardware impairment coefficient of the base station; Hidden users in the first The received signal at each symbol is: ; in, This indicates that the integrated communication and sensing base station uses a composite channel between a reconfigurable smart surface and the concealed user Bob, where... This represents the channel from the reconfigurable smart surface to the covert user Bob. The conjugate transpose of . This represents the projected phase matrix of the reconfigurable smart surface. This represents the projected phase noise matrix of the reconfigurable smart surface. This represents the channel from a dual-function base station to a reconfigurable smart surface. Indicates the first The Gaussian white noise at each symbol follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This indicates that the hardware impairment at the location of the hidden user Bob follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. The mathematical expectation of the signal. The covariance matrix represents the transmitted signal; The covariance matrix of the transmitted signal is expressed as: ; in, Indicates the launch of public information The corresponding beamforming vector is given to the public user Carol. Indicates the transmission of covert information The corresponding beamforming vector is given to the covert user Bob. It is the covariance of the complex Gaussian distribution of the dedicated sensing signal. This represents the hardware impairment coefficient of the base station. This means extracting the vector elements inside the parentheses to form a new diagonal matrix; Public users in The received signal at each symbol is: ; in, This indicates that the integrated communication and sensing base station uses a composite channel between a reconfigurable smart surface and the public user Carol, where... This represents the channel from the reconfigurable smart surface to the public user Carol. The conjugate transpose of . This represents the reflection phase matrix of the reconfigurable smart surface. This represents the reflection phase noise matrix of the reconfigurable smart surface. This represents the channel from a dual-function base station to a reconfigurable smart surface. Indicates the first The Gaussian white noise at each symbol follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This indicates that the hardware impairment at the public user Carol's location follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution, where The mathematical expectation of the signal. The covariance matrix represents the transmitted signal; The rival guardian Willie in the The received signal at each symbol and the total error detection probability are: ; ; in, This indicates that the integrated communication and sensing base station uses a composite channel between the reconfigurable smart surface and the hostile guardian Willie, where... This represents the channel from the reconfigurable smart surface to the hostile guardian Willie. The conjugate transpose of . This represents the projected phase matrix of the reconfigurable smart surface. This represents the projected phase noise matrix of the reconfigurable smart surface. This represents the channel from a dual-function base station to a reconfigurable smart surface. Indicates the first The Gaussian white noise at each symbol follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution. This represents the total error detection probability. This represents the probability of a false positive. Indicates the probability of missed detection; The sensing signal model includes the echo signal received by the base station, and its specific expression is as follows: ; in, Represents the echo channel from the reconfigurable smart surface to the base station. transpose, This represents the transpose of the projected phase noise matrix of the reconfigurable smart surface. This represents the channel coefficients related to the target's radar cross-section. This represents the turning vector of the hostile guardian Willie. This indicates Willie's azimuth. This represents the channel from a dual-function base station to a reconfigurable smart surface. Indicates the first time when sending a covert signal The transmitted signal at each symbol This indicates that the Gaussian white noise at the base station follows a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance matrix of a complex Gaussian distribution. This indicates that the hardware impairments at the base station receiver follow a complex Gaussian distribution, where parameter 0 is the mean of the complex Gaussian distribution. It is the covariance of the complex Gaussian distribution, where The mathematical expectation of the signal. This represents the hardware impairment coefficient of the base station.
4. The STAR-RIS-assisted integrated optimization method for concealed synesthesia based on deep reinforcement learning according to claim 3, characterized in that, In step S2, the specific expression of the optimization objective function is: in, This represents the total power consumption of the static circuitry at the base station. This indicates the power amplification efficiency at the base station. This indicates that the hidden user Bob is transmitting covert information. The achievable rate at that time This indicates that the public user Carol transmits hidden information. The achievable rate at that time; The original problem of the optimization problem is specifically described as follows: ; Indicates the system's energy efficiency. Indicates the launch of public information The corresponding beamforming vector is given to the public user Carol. Indicates the transmission of covert information The corresponding beamforming vector is given to the covert user Bob. This indicates the maximum power threshold of the base station. This represents the threshold of the CRB (Crame-Rao Boundary). This represents the conjugate transpose of the composite channel between the integrated communication and sensing base station and the concealed user Bob via a reconfigurable smart surface. This indicates that the hidden user Bob is transmitting covert information. The achievable rate at that time This indicates that the user can decode the signal. , No. The transmission amplitude coefficient of each component Indicates the first The reflection transmission amplitude coefficient of each component, Indicates the first The phase shift corresponding to the transmission of each component Indicates the first The phase shift corresponding to the reflection of each component express The solution.
5. The STAR-RIS-assisted integrated optimization method for concealed synesthesia based on deep reinforcement learning according to claim 4, characterized in that, In step S3, the specific process of the deep enhancement algorithm is as follows: Using the base station central controller as an agent and a covert wireless network assisted by a hardware-constrained reconfigurable smart surface STAR-RIS as the environment, a Markov decision process model is constructed. This model includes states, actions, and immediate rewards. The agent observes current state information from the environment, including dynamic channel state information. The agent inputs this current state information into a pre-trained optimal policy network to obtain the current action, which includes the beamforming vector, STAR-RIS phase shift, and perceptual covariance matrix to be optimized. The agent executes the action and receives an immediate reward from the environment, which is associated with the system's objective function and constraints. Through continuous interaction between the agent and the environment, the policy network is optimized based on the immediate reward to maximize the cumulative reward, thereby achieving joint optimization of the beamforming vector, STAR-RIS phase shift matrix, and perception covariance matrix.
6. The STAR-RIS-assisted integrated optimization method for concealed synesthesia based on deep reinforcement learning according to claim 5, characterized in that, The action space in the Markov decision process Defined as the set of all possible actions that an agent can perform in any given state, the actions ∈ The joint control variables to be optimized include the beamforming vector of the base station, the discrete phase shift values of each unit of the simultaneously transmitted and projected reconfigurable smart surface, and the covariance matrix of the sensed waveform. The expression is as follows: ; in, Indicates the first The agent's actions during step training Indicates the first Public information is launched during training. The corresponding beamforming vector is given to the public user Carol. Indicates the first Launching covert information during training The corresponding beamforming vector is given to the covert user Bob. Indicates the first Dedicated sensing signals during step training Indicates the first The reflection transmission amplitude coefficient of the components during step training. No. The projection phase offset of the components during step training. No. The reflection phase shift of the components during step training; The state space It consists of the set of all possible states of the environment, the states ∈ Satisfying the Markov property, including the system's dynamic channel state information and the agent's action in the previous time step, the expression is: ; in, Indicates the first -1 step of training: the agent's actions. Indicates the first The channel from the dual-function base station to the reconfigurable smart surface during step training. Indicates the first During training, the channel from the smart surface to the public user Carol can be reconstructed. Indicates the first During training, the channel from the smart surface to the hidden user Bob can be reconstructed. This represents the channel from the reconfigurable smart surface to the hostile guardian Willie. Indicates the first The echo channel from the smart surface to the base station can be reconstructed during step training; The instant reward function is a mapping mechanism used to determine the state of an agent. Next action Then, the environment generates a scalar feedback signal. The expression is: ; in, Indicates the system's energy efficiency. This indicates that the hidden user Bob is transmitting covert information. The achievable rate at that time This indicates that the user can decode the signal. This indicates that the integrated communication and sensing base station uses the conjugate transpose of the composite channel between the reconfigurable smart surface and the hostile guardian Willie. express The solution, It is the covariance of the complex Gaussian distribution. It is a step function. The value is 1. The value is 0. This represents the coefficient of the penalty term.
7. The STAR-RIS-assisted integrated optimization method for hidden synesthesia based on deep reinforcement learning according to claim 6, characterized in that, The proposed dual-delay deep deterministic policy gradient-first empirical replay algorithm based on a diffusion model replaces the actor network in the dual-delay deep deterministic policy gradient algorithm with a generative diffusion model, while employing priority empirical replay to improve algorithm performance. The inverse process of the generative diffusion model involves a series of iterative denoising steps, from the noise vector... Reconstructing the original data sample Based on the original data sample distribution Select Action The approximate expression for the reverse process is: ; in, Follows a standard Gaussian distribution. This represents the mean of a standard Gaussian distribution. Indicates the first Noise samples at each diffusion step Indicates a time step. Indicates state, Denotes the predetermined variance coefficients, where , express Factors The cumulative product, Indicates the covariance scheduling factor; Through gradual noise reduction, the original data sample Represented as: ; in, Represents a deep neural network. Indicates the first Noise samples at each diffusion step express Factors The cumulative product, Indicates state; Deep Neural Networks In a given state The denoised noise is generated under the given conditions, and its mean is subsequently approximated indirectly: ; in, Indicates the noise mean. Represents a deep neural network. Indicates the first Noise samples at each diffusion step express Factors The cumulative product, Indicates state; By reparameterizing to achieve differentiable sampling and stable backpropagation, the reverse process proceeds as follows: After noise reduction, it can be expressed as: ; in, Indicates the noise mean. Indicates the first Noise samples at each diffusion step Indicates state, It is the operator for the Hadamard product. This indicates that it follows a standard Gaussian distribution. It is the identity matrix. ,in express Factors The cumulative product, This represents the covariance scheduling factor.
8. The STAR-RIS-assisted integrated optimization method for hidden synesthesia based on deep reinforcement learning according to claim 7, characterized in that, The gradient-first empirical replay algorithm for a dual-delay deep deterministic strategy based on a diffusion model includes the following steps: S31: System initialization; initialize two online commentator networks. , An online actor network based on a generative diffusion model and the corresponding parameters Initialize the target network and its parameters corresponding to the online network. Simultaneously, initialize a priority experience replay buffer. ; S32: Network Training and Motion Generation; Regarding Training Steps =1 to Perform the following sub-steps: S321: Observe the current environmental status and initialize random variables. ,in Follows a standard Gaussian distribution. It is the identity matrix; S322: Diffusion denoising process; for the number of diffusion steps = Step 1, then perform the following sub-steps: S3221: Generating a denoised distribution using a deep neural network ,in Represented as conditional information; S3222: Calculate the mean of a Gaussian distribution. and ; S33: Action selection and execution; calculation And based on this, randomly select an action. Perform the selected action in the environment. Observe the next state and instant rewards ; S34: Experience storage and extraction; storing and extracting experience tuples Store in priority playback buffer And set its priority to the current maximum priority, from the priority experience PER buffer. Extract a batch of empirical data from D; S35: Experience priority update; for experience = Calculate up to D, the first... Sampling probability of a piece of experience ,in Representing experience priority, Indicates the capacity of the experience buffer. This is an introduced positive constant used to ensure that the probability of each experience being sampled is not zero. Prioritize experience replay factors; calculate experience weights and update experience priorities. Each experience tuple is then weighted according to the absolute value of its temporal error (TD). Weights are assigned, and samples are then drawn in order of weight. The weights are defined as follows: Weighted normalization representation , Indicates hyperparameters, Indicates the first Sampling probability of a piece of experience This represents the maximum value among the K weights; Compensation for the high variance caused by importance sampling weights; the first The timing error (TD) of the empirical data. Represented as ,in , This refers to an online network of critics. An online actor network representing a generative diffusion model. This indicates the corresponding online network parameters; This represents the target network parameters corresponding to the online network, and the noise. Indicates noise It follows a Gaussian distribution, where the parameter 0 is the mean. Let Variance be the variance, and its value be [value]. Indicates the discount factor; S36: Update the online critic network parameters; based on the loss function. Update the parameters of the online critic network, including Indicates the normalized weights. , This refers to an online network of critics. An online actor network representing a generative diffusion model. This indicates the corresponding online network parameters; This indicates the target network parameters corresponding to the online network; S37: Update online actor network parameters; based on the loss function Update online actor network parameters; S38: Soft update target network parameters; based on soft update , , Update the target network parameters, where Indicates the soft update weight. This indicates the corresponding online network parameters; This represents the target network parameters corresponding to the online network.
9. The STAR-RIS-assisted integrated optimization method for hidden synesthesia based on deep reinforcement learning according to claim 8, characterized in that, In step S4, the agent is trained through continuous interaction with the environment. After multiple iterations, the policy network of the deep reinforcement learning algorithm converges. Specifically, the convergence of the policy network means that the overall energy efficiency index of the concealed sensory integrated system converges to a stable value.