A multi-robot intelligent metasurface assisted lossy communication system optimization method

By optimizing the phase shift of the IRS reflection unit using multidimensional correlated complex Gaussian variables and Markov decision processes, the channel correlation problem under multi-IRS collaborative deployment was solved, realizing a low-complexity, high-robust lossy communication system and improving the spectrum efficiency and reliability of 6G communication.

CN121151823BActive Publication Date: 2026-06-05DAOKE ZHIXING (XIAN) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DAOKE ZHIXING (XIAN) TECHNOLOGY CO LTD
Filing Date
2025-09-13
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

In lossy communication systems assisted by multi-robot intelligent metasurfaces, existing technologies have failed to effectively address the channel spatial correlation problem caused by the collaborative deployment of multiple IRSs, leading to inaccuracies in traditional outage probability analysis models and affecting system reliability and spectral efficiency.

Method used

The channel is modeled using multidimensional correlated complex Gaussian variables. By combining Shannon's loss theorem and Markov decision processes, the phase shift of the IRS reflection unit is optimized through the CTDE-TD3 algorithm, which is trained centrally and executed in a distributed manner, thus achieving joint control with low complexity and high robustness.

Benefits of technology

It significantly reduces the probability of system outages, provides theoretical support and practical tools for AI-driven 6G communication systems, and improves the system's spectrum efficiency and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121151823B_ABST
    Figure CN121151823B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-robot intelligent metasurface auxiliary lossy communication system optimization methods, first, under the framework of relevant Rayleigh channel, the spatial correlation of base station-IRS-user link is jointly modeled, the coupling channel matrix between the IRS reflecting unit carried by any robot is constructed using multidimensional correlated complex Gaussian variable, the distortion constraint is converted into SNR threshold using Shannon lossy theorem, and the outage probability closed form is derived;Second, the joint optimization is modeled as multi-agent Markov decision process MAMDP, and the IRS carried by each robot is an independent agent, which is centrally trained and distributedly executed using CTDE-TD3 algorithm;Finally, a low-complexity outage lower bound is given to realize real-time prediction.The application can realize multi-IRS cooperative regulation with low complexity and high robustness, significantly reduce the system outage probability, and provide theoretical support and practical tools for AI-oriented 6G communication system design and real-time optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication technology, specifically relating to an optimization method for a multi-robot intelligent metasurface-assisted lossy communication system. Background Technology

[0002] Within the research framework for sixth-generation (6G) mobile communication, intelligent reflecting surface (IRS) technology is considered a key enabling technology for improving network coverage and system capacity due to its low cost, low power consumption, and reconfigurable wireless propagation environment. However, most existing work focuses on the deployment and optimization of a single IRS in an ideal independent Rayleigh channel, failing to fully consider the channel spatial correlation problem caused by the collaborative deployment of multiple IRSs in real-world scenarios. Specifically, when multiple robots carrying IRSs are deployed in the same area, the base station-IRS-user links experienced by each reflecting unit exhibit significant coupling in terms of propagation path, scatterer distribution, and arrival / departure angles, resulting in a high correlation of channel coefficients. This correlation not only weakens the spatial diversity gain of multiple IRSs but also causes deviations in the traditional outage probability analysis model based on the independent and identically distributed Rayleigh assumption, thus affecting the accuracy of system reliability assessment.

[0003] In 6G air-to-ground integrated and emergency communication scenarios, multiple autonomously mobile robots are needed to be dispatched to designated areas to perform on-site reconnaissance, rescue, or relay tasks. To quickly restore or enhance local coverage, each robot can carry a reconfigurable intelligent metasurface and be dynamically deployed at different coordinate points. In this scenario, the reflective units of the IRS carried by each robot and the base station-user link exhibit spatially non-uniform coupling due to differences in robot positions. That is, IRS units within the same robot experience similar scattering environments, while the links between different robots exhibit highly correlated but non-stationary channel statistical characteristics due to position, orientation, and body obstruction. The traditional outage probability model based on the assumption of "fixed location, independent and identically distributed Rayleigh" no longer holds, leading to inaccurate system reliability assessments.

[0004] With the deep integration of 6G networks and artificial intelligence, the communication paradigm is shifting from "lossless transmission" to "decision-oriented lossy communication," offering a solution to the aforementioned problems—the receiver only needs to recover information that meets a specific distortion threshold, thereby improving spectral efficiency and reducing the probability of outages. However, in the framework of multi-robot mobile IRS, channel correlation, robot position uncertainty, distributed IRS phase shift control, and distortion tolerance are deeply coupled: high-dimensional, non-convex joint optimization poses a severe challenge to the robot's limited computing power and battery life; real-time accurate CSI is difficult to obtain; if the robot makes distributed decisions based solely on local observations, it is highly susceptible to a game-theoretic dilemma of local beam optimization and global outage deterioration. Therefore, there is an urgent need for a low-complexity, high-robustness joint optimization framework for multi-robot intelligent metasurface-assisted lossy communication systems. Under the premise of unified quantification of channel correlation, robot position uncertainty, and distortion thresholds, this framework can achieve real-time collaborative control of distributed IRS phase shifts, significantly reducing the probability of system outages, and providing theoretical support and practical tools for AI-driven 6G lossy communication. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, this invention provides an optimization method for a multi-robot intelligent metasurface-assisted lossy communication system. First, within a relevant Rayleigh channel framework, the spatial correlation of the base station-IRS-user link is jointly modeled. A coupling channel matrix between the reflection units of any robot-carried IRS is constructed using multidimensional correlated complex Gaussian variables. The distortion constraint is transformed into an SNR threshold using Shannon's loss theorem, deriving a closed-form expression for the outage probability. Second, the joint optimization is modeled as a multi-agent Markov decision process (MAMDP), with each robot-carried IRS acting as an independent agent. The CTDE-TD3 algorithm is used for centralized training and distributed execution. Finally, a low-complexity lower bound for outages is provided, enabling real-time prediction. This invention achieves multi-IRS collaborative control with low complexity and high robustness, significantly reducing the system outage probability and providing theoretical support and practical tools for the design and real-time optimization of AI-oriented 6G communication systems.

[0006] The technical solution adopted by this invention to solve its technical problem is as follows:

[0007] Step 1: In the base station-IRSk-single user link, characterize the channel of each IRS reflection unit using multidimensional correlated complex Gaussian variables;

[0008] Step 2: Based on Shannon's lossy source-channel separation theorem, the distortion threshold D is transformed into the required minimum SNR, and analytical constraints on the outage probability are given;

[0009] Step 3: The joint optimization problem is modeled as a multi-agent Markov decision process (MAMDP); an independent agent is configured for each IRS, and efficient collaborative distributed control of IRS reflection coefficients is achieved based on the centralized training and distributed execution CTDE framework.

[0010] Furthermore, step 1 specifically includes:

[0011] Step 1-1: Construct the downlink model of base station-K block IRS-single user BS-IRSk-User;

[0012] Assume the base station is located at the origin of the coordinate system and equipped with a single omnidirectional transmitting antenna with a transmission power of P. s K passive IRS blocks are deployed within the visible area of ​​the link, denoted as: Each IRS contains M reflective elements, with the spacing between elements much smaller than the wavelength, and the phase θ of each element is... k,m ∈[0,2π] can be independently controlled; for the k-th IRS, its reflection unit can be represented as

[0013] Consider a scenario where the channel coefficient of a line-of-sight link d is G0 and the path loss is PL0. Assuming the line-of-sight link is slowly variable, its CSI can be used long-term with a single measurement feedback. Each signal beam emitted by the BS is transmitted through the BS-IRSk-User link, and finally, the user end superimposes the reflected signal beams received from K IRS-controlled signals to obtain an equivalent total signal. The signal beam y received through any BS-IRSk-User link... k Represented as:

[0014]

[0015] In the formula, x k,m PL is the signal beam transmitted from the source to the k-th IRS. k G is the equivalent path loss of the k-th BS-IRS-User link. k,m , G′ k,m y represents the transmission channel coefficients from the m-th reflection unit of BS to IRSk and from the m-th reflection unit of IRSk to the user terminal, respectively; k,m This represents the signal transmitted via the IRSk-User link and ultimately received by the user, where M represents the number of IRS reflection units, and P... s θ0 represents the initial phase, where θ represents the base station's transmitted signal power.

[0016] Step 1-2: Generate equivalent channel coefficients;

[0017] The equivalent channel coefficients of the BS-IRSk-User link are defined as follows:

[0018]

[0019] In the formula, Θ k Let G be the reflection coefficient matrix of IRSk. k G′ is the array of incident channel coefficients from BS to the k-th IRS; k This is the set of reflection channel coefficients from the k-th IRS to the user. Indicates the transpose of G′;

[0020] Finally, the user terminal receives and synthesizes the reflected signals from the K IRS:

[0021]

[0022] Where G0 represents the channel coefficient of line-of-sight link d, PL0 represents the path loss, and N0 represents random noise;

[0023] Therefore, the power of the synthesized signal at the user end is:

[0024]

[0025] Define the magnitude of the equivalent channel coefficients at the receiver in the model scenario, denoted as |H T |:

[0026]

[0027] Steps 1-3: Establish relevant Rayleigh channel models;

[0028] For any BS-IRSk-User link, the Rayleigh channel model is used to characterize the correlation of the channels in each reflecting unit; the channel coefficient G of the m-th reflecting unit is... k,m and G′ k,m for:

[0029]

[0030] Where: j is the imaginary unit, λ k,m ,λ′ k,m ∈(-1,1) is, G k,m ,G′ k,m Related factors, X k,0 ,Y k,0 ,X′ k,0 ,Y′ k,0 Used as the baseline variable; Let σ be mutually independent normally distributed random variables; k,m , σ′ k,m G k,m , G′ k,m The variance follows a normal distribution;

[0031] Define the modulus of the incident channel coefficient array and the reflection channel coefficient array for each IRS block:

[0032] |G k |=(|G k,1 |,|G k,2 |,…,|G k,M |)=(r k,1 ,r k,2 ,…,r k,M ),r m ≥0

[0033] |G′ k |=(|G′ k,1 |,|G′ k,2 |,…,|G′ k,M |)=(r′ k,1 ,r′ k,2 ,…,r′ k,M ),r′ m ≥0

[0034] Furthermore:

[0035]

[0036] in: I0(·) is a zeroth-order Bessel function of the first kind; r m Let r′ represent the magnitude of the incident link channel coefficient of the m-th reflecting unit. m Let r denote the magnitude of the reflection link channel coefficient of the m-th reflection unit. k,m Let r′ represent the magnitude of the incident channel coefficient of the m-th reflecting unit of the k-th IRS. k,m Let represent the magnitude of the reflection channel coefficient of the m-th reflection unit in the k-th IRS;

[0037] Furthermore, step 2 specifically includes:

[0038] Step 2-1: Based on Shannon's lossy source-channel separation theorem, the system's distortion threshold D is transformed into the minimum signal-to-noise ratio threshold γ′0 that the channel side needs to satisfy, as well as the constraint on the sum of the magnitudes of the equivalent channel coefficients of each IRS auxiliary link.

[0039] Let K = 2 IRSs used for auxiliary communication, M = 4 IRS reflection units, and C be the channel capacity per unit bandwidth of the system. Assume the system uses Gaussian codebook for channel coding. According to Shannon's theorem, the instantaneous channel capacity of the system is a function of the instantaneous signal-to-noise ratio, C(γ). Furthermore, the incident and reflected channel environments of the IRS are independent, and each signal beam is in phase at the receiver after phase modulation by the IRS. The signal at BS is generated using a binary Bernoulli source U ~ Bern(p), p ∈ [0, 0.5], with a lossy source rate-distortion function R. S (D), the system channel coding rate is R C =1, the threshold signal-to-noise ratio is γ′0, and a communication interruption event will occur when the real-time signal-to-noise ratio of the system is lower than γ′0; from the definition of signal-to-noise ratio, we know:

[0040]

[0041] Among them, P R This indicates the combined signal power received by the user. Indicates noise power;

[0042] The signal-to-noise ratio (SNR) corresponding to the achievable limit bit rate for rate-distortion in a lossy communication system is γ′0. For the scenario where B = 1 / 2, according to the source-channel separation theorem, we have:

[0043]

[0044] The simplified relationship between the threshold signal-to-noise ratio γ0′ and the acceptable distortion D of the lossy communication system is as follows:

[0045]

[0046] The condition for the system not to experience a communication interruption event is that the real-time signal-to-noise ratio of the synthesized signal at the receiving end is not lower than the threshold γ0′:

[0047]

[0048] The calculation yields:

[0049]

[0050] Where PL′ represents the path loss of the auxiliary link;

[0051] The above formula can be rewritten as:

[0052]

[0053] The constraint on the sum of the magnitudes of the equivalent channel coefficients for each IRS auxiliary link is as follows:

[0054]

[0055] The term on the right side of the inequality is denoted as H0′;

[0056] Step 2-2: Calculate the system interruption probability;

[0057] For the BS-IRSk-User auxiliary link, its equivalent channel coefficient magnitude |H k |The multivariate joint probability density function is The channel conditions between IRS auxiliary links are independent, yielding the joint probability density function of the sum of the magnitudes of the equivalent channel coefficients of the two IRS auxiliary links: |H1| + |H2|. Calculate the probability of system outage:

[0058]

[0059] in As given in step 1, Ω is the integration region where the communication interruption event occurs, which includes the case where the sum of the magnitudes of the equivalent channel coefficients of the two IRS auxiliary links is less than the threshold H0′.

[0060] The interruption probability is rewritten as:

[0061]

[0062] By changing the order of integration, the system's interruption probability can be obtained:

[0063]

[0064] Steps 2-3: Modeling the lossy communication optimization problem;

[0065] The equivalent channel coefficient modulus |H at the receiving end T As the optimization objective, the optimization problem can be described as follows: Under the conditions of satisfying the system's lossy communication feasibility requirements, received signal power constraints, phase shift constraints of each IRS reflection unit, and the inability to directly obtain real-time CSI, jointly set the reflection coefficients of K IRS blocks to maximize |H|. T The specific details are as follows:

[0066]

[0067] θ k,m ∈[0,2π],

[0068]

[0069] λ k,m ,λ′ k,m ∈[-1,1],

[0070]

[0071] Where λ k =(λ k,1 ,..,λ k,M ),λ′ k =(λ′) k,1 ,..,λ′ k,M ) are the incident channel coefficient correlation factor array and the reflected channel coefficient correlation factor array of IRSk, respectively. max This indicates that the system's received signal power cannot exceed a certain set maximum power value, Θ k Let E(.) denote the reflection coefficient matrix of IRSk, E(.) denotes the expectation, and d(.) denotes the distortion measure function.

[0072] Furthermore, step 3 specifically includes:

[0073] Step 3-1: Consider the optimization scenario where each IRS has a continuous phase shift, and transform the optimization problem into a MAMDP problem;

[0074] For the links in step 1, the Rayleigh channel communication systems of the K IRSs are considered as the overall environment E, and the NNs deployed at each IRS are considered as agents, unable to directly obtain the real-time CSI of their respective channel environments; based on the above settings, The model is a MAMDP, denoted as<S,A,R> :

[0075] S represents the state space, defined as the set of reflection coefficient matrices of each IRS and the state information of the communication system: for IRSk, its local state space S is defined. k Its real-time reflection coefficient state information and its conditional state information; let Represents the state information of IRSk at time t, where:

[0076]

[0077] The state space S of the entire system can then be represented as the sub-state spaces Si. k The union, namely:

[0078]

[0079] S = {S1,S2,S3,…,S} M}

[0080] in, This represents the conditional state information of IRSk at time t;

[0081] A represents the action space: defined as the space in which each agent adjusts the IRS reflection coefficient; for the k-th agent, its local action space A is defined. kTo provide the space for adjusting the reflection coefficient of IRSk, let This represents the local state of the agent with respect to IRSk at time t. The actions taken utilize a continuous motion space:

[0082]

[0083] Where: Δ m ∈(0,π / 6] is a constant, set before optimization begins; Δθ k,1 ,Δθ k,2 ,…,Δθ k,M This represents the phase adjustment of IRSk on its M reflection units at time t;

[0084] The action space A of the entire system is the set of the action spaces of all agents:

[0085]

[0086] A = {A1, A2, A3, ..., A} M}

[0087] R represents the system's reward space: based on the optimization problem The constraints and settings are defined, with the action reward consisting of a reward term and a penalty term; for the k-th agent, assume that the local state observed at time t is... The action to be performed is The equivalent channel coefficient of the BS-IRSk-User link is The IRSs are in a cooperative relationship. There are two design methods for the reward obtained by agent k. The first method is to set the same global cooperative reward for all agents:

[0088]

[0089] l is a penalty term when the power of the synthesized signal at the receiving end does not meet the minimum requirements for lossy communication or exceeds the maximum power constraint. Its function value is determined by the state information and action information.

[0090]

[0091] When the equivalent channel coefficient magnitude |H T |Rewards for each agent when the system's lossy communication feasibility constraint is not met When the maximum power constraint at the receiver is not met Then set it to |H T | Upper limit constraint value H Tmax ;

[0092] Step 3-2: Assign an independent agent to each IRS, using a centralized training-distributed execution CTDE framework;

[0093] For multi-IRS assisted communication scenarios, the reflection coefficient of multiple IRSs is optimized by distributing intelligent agents to each IRS to independently adjust its reflection coefficient;

[0094] For any agent k, a DDNN based on a two-branch TD3 framework is used as its core; the Critic network uses two identical neural networks to simultaneously learn two Q functions, evaluates the value of the agent's actions, and selects the smaller Q value for network parameter updates; π is used. k This indicates that it evaluates the Actor network, Q k,1 Q k,2 These represent two evaluation Critic networks, where μ k ,ω k,1 and ω k,2 These represent the parameters of the corresponding networks, denoted by π′. k ,q′ k Let μ' and μ' represent the target network, respectively. k ,ω′ k,1 and ω′ k,2 ;

[0095] Each agent employs a CTDE interaction strategy: a Critic network is deployed for each agent, Q... k,1 Q k,2 All are centralized evaluation networks, meaning that during training, they can access the global action 'a', which contains all the local actions and states of all agents. t and global state s t The Actor network deployed by each agent adopts a distributed execution strategy, with each agent relying solely on its own local observation information. Generate Actions Without relying on the action information of other agents; all agents share a common experience replay pool. Record global state s t Action a t and reward information r t This is so that each agent can better learn the strategies that exist in environments with other agents.

[0096] An electronic device includes: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to enable the electronic device to perform the above-described communication system optimization method.

[0097] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described communication system optimization method.

[0098] A chip includes a processor for retrieving and running a computer program from a memory, causing a device equipped with the chip to perform the aforementioned communication system optimization method.

[0099] A computer program product includes a computer storage medium storing a computer program, the computer program including instructions executable by at least one processor, which, when executed by the at least one processor, implement the above-described communication system optimization method.

[0100] The beneficial effects of this invention are as follows:

[0101] This invention constructs an accurate collaborative optimization model for a multi-IRS-assisted lossy communication system under a highly correlated Rayleigh channel based on the MADRL framework of centralized training and distributed execution. On this basis, a closed lower bound for the system outage probability satisfying arbitrary distortion thresholds is derived, achieving low-complexity and highly robust real-time performance prediction. The method proposed in this invention can provide theoretical support and practical tools for the design and real-time optimization of AI-oriented 6G communication systems in complex multi-IRS scenarios. Attached Figure Description

[0102] Figure 1 It is a lossy communication system scenario model with multiple IRS assistance in the context of related channel scenarios.

[0103] Figure 2 This is a schematic diagram of a lossy communication system with multiple IRS assistance in a correlated channel scenario.

[0104] Figure 3 It is an algorithm framework based on MADRL.

[0105] Figure 4 These are curves showing the change in system interruption probability with signal transmission power in embodiments of the present invention, (a) correlation of different channel coefficients, and (b) acceptable distortion of different systems.

[0106] Figure 5 This is the curve showing the change of reward with the number of iterations when ν = 0 in an embodiment of the present invention.

[0107] Figure 6 These are the curves showing the change of rewards for each agent with the number of iterations when ν = 0.3 in this embodiment of the invention, a) reward of agent 1, b) reward of agent 2.

[0108] Figure 7These are the curves showing the change of rewards for each agent with the number of iterations when ν = 0.7 in this embodiment of the invention, a) reward of agent 1, b) reward of agent 2.

[0109] Figure 8 This is the curve showing the change of system reward with the number of iterations when K = 1, 2, 3 in the embodiments of the present invention.

[0110] Figure 9 These are curves showing the change of system reward with the number of iterations under different multi-agent frameworks in embodiments of the present invention.

[0111] Figure 10 The curves showing the average optimization performance of the algorithm in this embodiment of the invention as a function of the number of deployed IRS K are: (a) average convergence reward, and (b) average number of convergence passes.

[0112] Figure 11 The average optimization performance of the algorithm in this embodiment of the invention is related to the channel coefficient λ. k ,λ′ k Value variation curves, (a) average convergence reward, (b) average number of convergence passes.

[0113] Figure 12 These are the performance curves of the reflection coefficient optimization algorithm based on MADRL in this embodiment of the invention, (a) correlation of different channels, and (b) receiveable distortion of different systems.

[0114] Figure 13 This is the curve showing the change in average interruption probability of the system in this embodiment of the invention with acceptable distortion. Detailed Implementation

[0115] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0116] This invention utilizes the CTDE-MADRL framework to achieve collaborative phase shift optimization of multiple IRSs in highly correlated channels, and provides a lower bound for the closed-loop interruption probability that satisfies the distortion threshold. It has low complexity, strong real-time performance, and can directly serve the design and optimization of AI-driven 6G systems.

[0117] Step S1: In the base station-IRSk-single user link, characterize the channel of each IRS reflection unit using multidimensional correlated complex Gaussian variables;

[0118] Step S101: Construct the base station-K-block IRS-single user downlink model. The lossy communication system scenario model and schematic diagram with multiple IRS assistance under relevant channel scenarios are as follows: Figure 1 , Figure 2 As shown. Assume the base station is located at the origin and equipped with a single omnidirectional transmitting antenna with a transmit power of P. s K passive IRS blocks are deployed within the visible area of ​​the link, denoted as: Each IRS contains M reflective elements, with the spacing between elements much smaller than the wavelength, and the phase θ of each element is... k,m ∈[0,2π] can be independently adjusted. For the k-th IRS, its reflection unit can be represented as

[0119] Considering a scenario where the channel coefficient of a line-of-sight link d is G0 and the path loss is PL0, and assuming the line-of-sight link is slowly variable, the CSI of this link can be used for a long time with a single measurement feedback. For the proposed communication system, the signal beams emitted by the BS are transmitted through the BS-IRSk-User link, and finally, the user end superimposes the reflected signal beams received from K IRS-controlled signals to obtain an equivalent total signal. The signal beam y received through any BS-IRSk-User link... k It can be represented as:

[0120]

[0121] In the formula, x k,m PL is the signal beam transmitted from the source to the k-th IRS. k G is the equivalent path loss of the k-th BS-IRS-User link. k,m , G′ k,m y represents the transmission channel coefficients from the m-th reflection unit of BS to IRSk and from the m-th reflection unit of IRSk to the user terminal, respectively; k,m This represents the signal transmitted via the IRSk-User link and ultimately received by the user, where M represents the number of IRS reflection units, and P... S θ0 represents the initial phase, where θ represents the base station's transmitted signal power.

[0122] Step S102: Generate equivalent channel coefficients. Accordingly, the equivalent channel coefficients of the BS-IRSk-User link can be defined as follows:

[0123]

[0124] In the formula, Θ k Let G be the reflection coefficient matrix of IRSk. k G′ is the array of incident channel coefficients from BS to the k-th IRS; k This is the set of reflection channel coefficients from the k-th IRS to the user. G′ k Transpose of;

[0125] Finally, the user terminal will receive and synthesize the reflected signals from K IRS:

[0126]

[0127] Where G0 represents the channel coefficient of line-of-sight link d, PL0 represents the path loss, and N0 represents random noise;

[0128] Therefore, the power of the synthesized signal at the user end can be obtained as follows:

[0129]

[0130] Similarly, the equivalent channel coefficient modulus of the receiver in this system model scenario can be defined by the above formula, denoted as |H T |:

[0131]

[0132] Step S103: Establish the relevant Rayleigh channel model. For any BS-IRSk-User link, the Rayleigh channel model is also used to characterize the correlation of the channels of each reflecting unit. The channel coefficient G of the m-th reflecting unit is... k,m and G′ k,m for:

[0133]

[0134] Where: j is the imaginary unit, λ k,m ,λ′ k,m ∈(-1,1) is, G k,m ,G′ k,m Related factors, X k,0 ,Y k,0 ,X′ k,0 ,Y′ k,0 As the benchmark variable, σ are mutually independent normally distributed random variables (local variables). k,m , σ′ k,m G k,m , G′ k,m The variance follows a normal distribution;

[0135] It is important to note that, since it is assumed that each IRS block is independent and cannot communicate with each other, the reference variables on which the channel coefficients of each IRS reflection unit depended during generation need to be resampled. For the convenience of subsequent analysis, the modulus values ​​of the incident channel coefficient array and the reflection channel coefficient array for each IRS block are defined here as follows:

[0136] |G k |=(|G k,1 |,|G k,2 |,…,|G k,M |)=(r k,1 ,r k,2 ,…,r k,M ),r m ≥0

[0137] |G′ k |=(|G′ k,1 |,|G′ k,2 |,…,|G′ k,M |)=(r′ k,1 ,r′ k,2 ,…,r′ k,M ),r′ m ≥0

[0138] Furthermore:

[0139]

[0140] in: I0(·) is a zeroth-order Bessel function of the first kind, r m Let r′ represent the magnitude of the incident link channel coefficient of the m-th reflecting unit. m Let r denote the magnitude of the reflection link channel coefficient of the m-th reflection unit. k,m Let r′ represent the magnitude of the incident channel coefficient of the m-th reflecting unit of the k-th IRS. k,m Let represent the magnitude of the reflection channel coefficient of the m-th reflection unit in the k-th IRS;

[0141] Step S2: Based on Shannon's lossy source-channel separation theorem, the distortion threshold D is transformed into the required minimum SNR, and an analytical constraint on the outage probability is given;

[0142] Step S201: Based on Shannon's lossy source-channel separation theorem, the system's distortion threshold D is transformed into the minimum signal-to-noise ratio threshold γ′0 that the channel side needs to satisfy, as well as the constraint on the sum of the magnitudes of the equivalent channel coefficients of each IRS auxiliary link.

[0143] Assume the number of IRSs used for auxiliary communication is K = 2, the number of IRS reflection units is M = 4, the channel capacity per unit bandwidth is C, and the system uses Gaussian codebook for channel coding. According to Shannon's theorem, the instantaneous channel capacity of the system is a function of the instantaneous signal-to-noise ratio C(γ). Furthermore, it is assumed that the incident and reflected channel environments of the IRS are independent, and that the signal beams are phase-coordinated at the receiver after phase modulation by the IRS to analyze the theoretical performance of the system. Here, the signal at BS is considered to be generated by a binary Bernoulli source U ~ Bern(p), p ∈ [0, 0.5], with a lossy source rate-distortion function R. S (D), the system channel coding rate is R. C =1, the threshold signal-to-noise ratio is γ′0, and a communication interruption event will occur when the real-time signal-to-noise ratio of the system is lower than γ′0. From the definition of signal-to-noise ratio, we can obtain:

[0144]

[0145] Among them, P R This indicates the combined signal power received by the user. Indicates noise power;

[0146] The signal-to-noise ratio (SNR) corresponding to the achievable limit bit rate for rate-distortion in a lossy communication system is γ′0. Considering the scenario where B = 1 / 2, according to the source-channel separation theorem, we can obtain:

[0147]

[0148] Simplification yields the relationship between the threshold signal-to-noise ratio γ′0 and the acceptable distortion D of the lossy communication system:

[0149]

[0150] The condition for the system not to experience a communication interruption event is that the real-time signal-to-noise ratio of the synthesized signal at the receiving end is not lower than the threshold γ′0:

[0151]

[0152] Substituting the values ​​into the calculation, we get:

[0153]

[0154] At any given time, the optimal optimization direction for the IRS reflection coefficient will tend towards the direction where the signal beams of each link are in phase. Here, we consider the case where each IRS is in optimal condition, and the equivalent channel coefficient H of each IRS link after optimization. k As the phase of G0 approaches the same, the above equation can be rewritten as:

[0155]

[0156] The constraint on the sum of the magnitudes of the equivalent channel coefficients for each IRS auxiliary link is as follows:

[0157]

[0158] The term on the right side of the inequality is denoted as H′0 to facilitate the subsequent calculation of the theoretical interruption probability of the system.

[0159] Step S202: Calculate the system interruption probability;

[0160] For the BS-IRSk-User auxiliary link, its equivalent channel coefficient magnitude |H k |The multivariate joint probability density function is Here, the channel conditions between IRS auxiliary links are assumed to be independent. Therefore, the joint probability density function of the sum of the magnitudes of the equivalent channel coefficients of the two IRS auxiliary links, |H1|+|H2|, can be obtained. Calculate the probability of system outage:

[0161]

[0162] in As given in step S1, A is the integration region where the communication interruption event occurs, encompassing cases where the sum of the magnitudes of the equivalent channel coefficients of the two IRS auxiliary links is less than the threshold H0′. The interruption function can be rewritten as:

[0163]

[0164] By changing the order of integration, the system's interruption probability can be obtained:

[0165]

[0166] Step S203: Modeling the lossy communication optimization problem;

[0167] The equivalent channel coefficient modulus |H at the receiving end T As the optimization objective, the optimization problem can be described as follows: Under the conditions of satisfying the system's lossy communication feasibility requirements, received signal power constraints, phase shift constraints of each IRS reflection unit, and the inability to directly obtain real-time CSI, jointly set the reflection coefficients of K IRS blocks to maximize |H|. T The specific details are as follows:

[0168]

[0169] θ k,m ∈[0,2π],

[0170]

[0171] λ k,m ,λ′ k,m ∈[-1,1],

[0172]

[0173] Where λ k =(λ k,1 ,..,λ k,M ),λ′ k =(λ′) k,1 ,..,λ′ k,M ) are the incident channel coefficient correlation factor array and the reflected channel coefficient correlation factor array of IRSk, respectively, Θ k Let E(.) denote the reflection coefficient matrix of IRSk, E(.) denotes the expectation, and d(.) denotes the distortion measure function.

[0174] Step S3: Model the joint optimization problem as a multi-agent Markov decision process (MAMDP); configure an independent agent for each IRS, and achieve efficient collaborative distributed control of IRS reflection coefficients based on the centralized training and distributed execution CTDE framework.

[0175] Step S301: Consider the optimization scenario where each IRS has a continuous phase shift, and transform the optimization problem into a MAMDP problem;

[0176] MADRL-based algorithm frameworks such as Figure 3 As shown. For the links in step S1, the Rayleigh channel communication systems of the K IRSs are considered as the overall environment E, and the NNs deployed at each IRS are considered as agents, unable to directly obtain the real-time CSI of their respective channel environments. Based on the above settings, to... The model is a MAMDP, denoted as<S,A,R> :

[0177] S represents the state space, defined as the set of reflection coefficient matrices of each IRS and the state information of the communication system: for IRSk, its local state space S is defined. k This provides its real-time reflection coefficient state information and its conditional state information. Let... This represents the state information of IRSk at time t, where:

[0178]

[0179] The state space S of the entire system can then be represented by its sub-state spaces Si. k The union, namely:

[0180]

[0181] in, This represents the conditional state information of IRSk at time t;

[0182] A represents the action space: defined as the space in which each agent adjusts the IRS reflection coefficient. For the k-th agent, its local action space A is defined. k To provide the space for adjusting the reflection coefficient of IRSk, let This represents the local state of the agent with respect to IRSk at time t. The actions taken utilize a continuous motion space:

[0183]

[0184] Where: Δ m ∈(0,π / 6] is a constant, set before optimization begins. Therefore, the action space A of the entire system is the set of action spaces for all agents:

[0185]

[0186] R represents the system's reward space: based on the optimization problem The constraints are defined here, where the action reward consists of a reward term and a penalty term. For the k-th agent, assume that its observed local state at time t is... The action to be performed is The equivalent channel coefficient of the BS-IRSk-User link is In the communication scenario studied in this chapter, the various IRSs are in a cooperative relationship. There are two design methods for the reward obtained by agent k. The first method is to set the same global cooperative reward for all agents:

[0187]

[0188] l is a penalty term when the power of the synthesized signal at the receiving end does not meet the minimum requirements for lossy communication or exceeds the maximum power constraint. Its function value is determined by the state information and action information.

[0189]

[0190] When the equivalent channel coefficient magnitude |H T |Rewards for each agent when the system's lossy communication feasibility constraint is not met When the maximum power constraint at the receiver is not met Then set it to |H T | Upper limit constraint value H Tmax .

[0191] Step S302: Assign an independent agent to each IRS, using a centralized training-distributed execution CTDE framework;

[0192] For multi-IRS assisted communication scenarios, the reflection coefficient of multiple IRSs is optimized by distributing agents to each IRS and independently adjusting their reflection coefficients. For any agent k, a DDNN based on a dual-branch TD3 framework is used as its core. The Critic network uses two identical neural networks to simultaneously learn two Q functions, evaluates the value of the agent's actions, and selects the smaller Q value for network parameter updates. This reduces the initial significant overestimation of Q values ​​by the agents in the DDPG framework, thus avoiding problems such as policy violation and insufficient exploration. π is used. k This indicates that it evaluates the Actor network, Q k,1 Q k,2 These represent two evaluation Critic networks, where μ k ,ω k,1 and ω k,2These represent the parameters of the corresponding networks, also denoted by π′. k ,Q′ k Let μ' and μ' represent the target network, respectively. k ,ω′ k,1 and ω′ k,2 .

[0193] In this algorithm, each agent will adopt a CTDE interaction strategy: a Critic network, Q... k,1 Q k,2 All are centralized evaluation networks, meaning that during training, they can access the global action 'a', which contains all the local actions and states of all agents. t and global state s t This allows for more efficient learning of collaborative strategies among multiple agents; the Actor network deployed by each agent adopts a distributed execution strategy, with each agent relying solely on its own local observation information. Generate Actions This algorithm maintains the independence of agent interactions without relying on the action information of other agents, which aligns with the scenario where IRSs in a communication system cannot communicate with each other. It also avoids the need for frequent access to global information during multi-step interactions with the environment, reducing the computational complexity of each interaction. Finally, in this algorithm, all agents share a common experience replay pool. Record global state s t Action a t and reward information r t This is so that each agent can better learn the strategies that exist in environments with other agents.

[0194] Example 1:

[0195] The Monte Carlo method was used to calculate the theoretical outage probability in K=2 and M=4 channel scenarios and compared with simulation values. The study investigated the changes in system outage probability under four scenarios (corresponding to two BS-IRS channels and two IRS-User channels) with different values ​​of correlation factor arrays (including channel independence) and four scenarios with different values ​​of acceptable distortion D (including lossless communication). Similarly, when studying scenarios with different values ​​of correlation factor arrays, the signal-to-noise ratio threshold γ0 at the receiver was set to 10dB, λ1=[0.95,0.9,0.9,0.85], λ′1=[0.9,0.95,0.85,0.9], and D=[0,0.1,0.2,0.35]. When studying scenarios with different acceptable distortions, the correlation channel coefficients of each IRS were set to strongly correlated scenarios. The sampling times for the four sets of correlation channel coefficient arrays in the Monte Carlo experiment were 2×10⁻⁶. 7 Second-rate.

[0196] Figure 4 The diagram illustrates the variation of the theoretical outage probability with the BS transmit signal power under different channel coefficient correlations and acceptable system distortions when the number of deployed IRSs K=2. It shows that, under the same communication system parameters, the overall theoretical outage probability decreases significantly after introducing new IRS auxiliary links. This indicates that lossy communication systems with multiple IRS assistance under correlated Rayleigh channels have a higher optimization upper limit.

[0197] Example 2:

[0198] The performance of the MADRL-based multi-IRS reflection coefficient optimization algorithm was verified through simulation experiments. The dual-branch structure of the two centralized Critic networks deployed by each agent requires a global state vector s with a dimension of (3M+1)·K. t After extracting a large number of state features, and combining them with a global action a of dimension M·K. t In this concatenation, the number of neurons in the fully connected layer of the dual-branch network is increased to enhance its ability to process high-dimensional data and better capture the interactions between multiple IRSs: the fully connected layer neurons in the dynamic branch are set to 256, and the fully connected layer neurons in the conditional branch are set to 128; correspondingly, the fully connected layer neurons in fusion layers 1 and 2 are also set to 128.

[0199] Simulation experiments were conducted for scenarios with K=1, 2, and 3 agents. K=1 represents a single IRS-assisted communication scenario (line-of-sight link exists and channel coefficient G0=2). The global experience pool was used. The size is increased to 10,000 to provide a sufficiently rich global sample pool for centralized training of each agent. Correspondingly, the batch size of samples extracted during network training is increased to 256 to ensure that the centralized Critic network, with its significantly increased network capacity, can be adequately trained.

[0200] Figure 5 The curves showing the reward variation with the number of iterations are presented when the number of agents K=2 and ν=0. This is achieved by using a single global feedback reward formula where each agent shares the same reward. The reward for v=0. You can see... Figure 5 After initial oscillations, the system reward curve tends to converge after about 180 iterations, which proves that the algorithm designed in this chapter can enable each agent to effectively learn the cooperative reflection coefficient optimization strategy in a multi-agent environment.

[0201] Figure 6 The reward settings for each agent are based on the local reward system when v = 0.3. It can be seen that when v = 0.3, the local rewards for each agent are relatively... Figure 5The convergence number of passes is reduced and the final reward oscillates within a local range. This is because the local reward prompts each IRS to prioritize increasing its own gain, but this disrupts phase alignment, forcing the global reward to be reverted, resulting in policy oscillation.

[0202] Figure 7 The curves showing the change of each agent's reward with the number of iterations are displayed when ν = 0.7. It can be seen that the convergence value of the final reward curve of each agent is relatively stable at this time, but the convergence value is lower than that of the previous two cases. This is because when the local weight is too high, each IRS only focuses on increasing its own link gain and ignores phase alignment. As a result, the superposition loss with the Loss component is large, and the global reward drops sharply.

[0203] Figure 8 The diagram illustrates the system reward curves after each agent performs the final step in each iteration, under different IRS deployment numbers K. It shows that as the number of agents K increases, the fluctuation of the system reward value during training increases, and the number of convergence iterations increases. This is because as the number of agents increases, the network dimension and the complexity of the cooperative strategy increase simultaneously, leading to a tendency for the system state to become unbalanced during training, resulting in significant fluctuations in the reward curve and slower convergence. However, overall, the algorithm still enables the system reward to converge to a relatively stable state, and the final convergence reward value significantly improves with increasing K. This verifies the effectiveness of the MADRL algorithm in this optimization scenario and also demonstrates that using multiple IRS assisted communication systems can achieve a higher optimization ceiling compared to single IRS assisted communication systems.

[0204] Figure 9 The same set of related channel coefficient samples G were shown. k ,G′ k The curves showing the change of system reward with the number of iterations when the above algorithms are used for optimization are shown below. It can be seen that the MADRL algorithm proposed in this invention can obtain better convergence reward than the MADQN and MAPPO algorithms, and has a faster convergence speed and more stable system reward change during training, which verifies the effectiveness of this invention.

[0205] Figure 10 , Figure 11 The figures show the average optimization performance of the system after convergence using multi-agent optimization algorithms with different policy frameworks in scenarios with multiple channel coefficient sampling, varying with the number of deployed IRS K and different channel coefficient correlation factors λ. k ,λ′ k A graph showing the relationship between the changing values.

[0206] Figure 10 Channel coefficient correlation factor array λ for each IRS k ,λ′ kAll are set as highly relevant scenarios, and the vertical axis represents the application of each optimization algorithm to N. s =The average convergence reward and average number of convergence passes of the system are calculated after the optimization of the scenario with 1000 sets of related channel coefficients to convergence. Overall, as K increases, the number of auxiliary links in the communication system increases, and the average convergence reward of the algorithm shows an upward trend; at the same time, the increase in the number of agents will lead to higher network complexity, further increasing the difficulty of finding cooperative strategies among agents, and the average number of convergence passes of the algorithm will increase accordingly.

[0207] Figure 11 The optimization performance curves of each algorithm under different channel coefficient correlation scenarios are further presented when K=2. As the channel correlation increases, the average convergence reward of the algorithm decreases and the number of steps required for convergence increases. The MADRL method designed in this invention has better optimization performance than the comparison algorithm in various correlation scenarios, and the decrease in average reward and the increase in average convergence steps are the smallest. This shows that the algorithm has good adaptability to optimization problems under correlated channel scenarios.

[0208] Figure 12 The diagram illustrates the variation of system outage probability with base station signal transmit power under different channel correlation and acceptable system distortion scenarios when K=2. The solid line represents the system outage probability calculated after convergence of the reflection coefficients under 1000 sets of correlated channel coefficient samples optimized by this algorithm, while the dashed line represents the theoretical outage probability calculated using the Monte Carlo method. The system outage probability optimized by this invention is very close to the theoretical outage probability obtained by the Monte Carlo method. This demonstrates that the algorithm designed in this invention can effectively jointly optimize the reflection coefficients of various IRSs under multiple communication scenarios with different channel coefficient correlations and acceptable system distortion, thereby improving the system's transmission reliability.

[0209] Figure 13 Show P S Under the sampling conditions of 0dB, K=2, and 1000 sets of highly correlated channel coefficients, four different algorithms were used to optimize the reflection coefficients of each IRS until convergence. The relationship between the system's average outage probability and the system's acceptable distortion D was then calculated. As the system's acceptable distortion D decreases, the average outage probability of the optimized system increases. Compared with the comparison algorithms, this invention has the best optimization effect and a more stable trend under different values ​​of acceptable distortion D.

Claims

1. An optimization method for a multi-robot intelligent metasurface-assisted lossy communication system, characterized in that, Includes the following steps: Step 1: In the base station-IRSk-single user link, characterize the channel of each IRS reflection unit using multidimensional correlated complex Gaussian variables; Step 1-1: Construct the downlink model of base station-K block IRS-single user BS-IRSk-User; Assume the base station is located at the origin of the coordinate system and equipped with a single omnidirectional transmitting antenna with a transmission power of . K passive IRS blocks are deployed within the visible area of ​​the link, denoted as: Each IRS contains M reflective elements, with the spacing between elements much smaller than the wavelength, and the phase of each element is... Independently adjustable; for the k-th IRS, its reflective unit can be represented as ; Considering the channel coefficients of the line-of-sight link d, Path loss is In this scenario, the line-of-sight link is considered to be slowly variable, allowing for long-term use of the CSI via a single measurement feedback. Each signal beam emitted by the BS is transmitted through the BS-IRSk-User link, and finally, the user end superimposes the reflected signal beams received from K IRS-controlled signals to obtain an equivalent total signal. The signal beam received through any BS-IRSk-User link... Represented as: ; In the formula, The signal beam transmitted from the source to the k-th IRS, The equivalent path loss for the k-th BS-IRS-User link is... , These are the transmission channel coefficients from the m-th reflection unit of BS to IRSk and from the m-th reflection unit of IRSk to the user terminal, respectively. This indicates the signal transmitted by the IRSk-User link and ultimately received by the user. This indicates the number of reflective elements in the IRS. This indicates the base station's transmit signal power. Indicates the initial phase; Step 1-2: Generate equivalent channel coefficients; The equivalent channel coefficients of the BS-IRSk-User link are defined as follows: ; In the formula, Here is the reflection coefficient matrix of IRSk. This is the array of incident channel coefficients from BS to the nth IRS; This is the set of reflection channel coefficients from the nth IRS to the user. express Transpose of; Finally, the user terminal combines the reflected signals from the received K IRS: ; in, This represents the channel coefficient of the line-of-sight link d. Indicates path loss. Indicates random noise; Therefore, the power of the synthesized signal at the user end is: ; Define the magnitude of the equivalent channel coefficients at the receiver in the model scenario, denoted as . : ; Steps 1-3: Establish relevant Rayleigh channel models; For any BS-IRSk-User link, the Rayleigh channel model is used to characterize the correlation of the channels in each reflection unit; the channel coefficient of the m-th reflection unit is... and for: ; Where: j is the imaginary unit. yes , Related factors, Used as the baseline variable; , are mutually independent normally distributed random variables; , They represent , The variance follows a normal distribution; Define the modulus of the incident channel coefficient array and the reflection channel coefficient array for each IRS block: ; Furthermore: ; in: It is a zeroth-order Bessel function of the first kind; This represents the magnitude of the incident link channel coefficient of the m-th reflecting unit. This represents the magnitude of the reflection link channel coefficient of the m-th reflection unit. This represents the magnitude of the incident channel coefficient of the m-th reflecting unit in the k-th IRS. This represents the magnitude of the reflection channel coefficient of the m-th reflection unit in the k-th IRS block; Step 2: Based on Shannon's lossy source-channel separation theorem, the distortion threshold D is transformed into the required minimum SNR, and analytical constraints on the outage probability are given; Step 2-1: Based on Shannon's lossy source-channel separation theorem, the system's distortion threshold D is transformed into the minimum signal-to-noise ratio threshold that the channel side must satisfy. And constraints on the sum of the magnitudes of the equivalent channel coefficients for each IRS auxiliary link; Let the number of IRSs used for auxiliary communication be... Number of IRS reflective units Let the channel capacity per unit bandwidth of the system be C, and assume that the system uses Gaussian codebook for channel coding. According to Shannon's theorem, the instantaneous channel capacity of the system is a function of the instantaneous signal-to-noise ratio. Furthermore, the incident and reflected channel environments of the IRS are independent of each other, and the signal beams are in phase at the receiver after phase modulation by the IRS; the signal at BS is a binary Bernoulli source. The generated, lossy source rate-distortion function is: The system channel coding rate is The threshold signal-to-noise ratio is When the real-time signal-to-noise ratio of the system is lower than Communication interruption events may occur; according to the definition of signal-to-noise ratio: ; in, This indicates the combined signal power received by the user. Indicates noise power; The signal-to-noise ratio corresponding to the achievable limit bit rate for a lossy communication system is the rate-distortion ratio. ,for In this scenario, according to the source-channel separation theorem: ; Simplifying, we get the threshold signal-to-noise ratio. Relationship with acceptable distortion D in lossy communication systems: ; The condition for the system not to experience a communication interruption event is that the real-time signal-to-noise ratio of the synthesized signal at the receiving end is not lower than a threshold. : ; The calculation yields: ; in, This represents the path loss of the auxiliary link; The above formula can be rewritten as: ; The constraint on the sum of the magnitudes of the equivalent channel coefficients for each IRS auxiliary link is as follows: ; The terms on the right side of the inequality are denoted as ; Step 2-2: Calculate the system interruption probability; For the BS-IRSk-User auxiliary link, its equivalent channel coefficient magnitude is The multivariate joint probability density function is The channel conditions between IRS auxiliary links are independent, and the sum of the magnitudes of the equivalent channel coefficients of the two IRS auxiliary links is obtained. joint probability density function ; Calculate the system's interruption probability: ; in As given in step 1, The integration region where the communication interruption event occurred includes the sum of the magnitudes of the equivalent channel coefficients of the two IRS auxiliary links being less than a threshold. The situation; The interruption probability is rewritten as: ; By changing the order of integration, the system's interruption probability can be obtained: ; Steps 2-3: Modeling the lossy communication optimization problem; The equivalent channel coefficient magnitude at the receiving end The optimization objective is to jointly set the reflection coefficients of K IRS blocks to maximize the signal strength while satisfying the system's lossy communication feasibility requirements, received signal power constraints, phase shift constraints of each IRS reflection unit, and the inability to directly obtain real-time CSI. The specific description is as follows: ; in These are the incident channel coefficient correlation factor array and the reflected channel coefficient correlation factor array of IRSk, respectively. This indicates that the system's received signal power cannot exceed a certain set maximum power value. This represents the reflection coefficient matrix of IRSk. Indicates the expectation. This represents a distortion measurement function; Step 3: The joint optimization problem is modeled as a multi-agent Markov decision process (MAMDP); an independent agent is configured for each IRS, and efficient collaborative distributed control of IRS reflection coefficients is achieved based on the centralized training and distributed execution CTDE framework.

2. The optimization method for a multi-robot intelligent metasurface-assisted lossy communication system according to claim 1, characterized in that, Step 3 specifically involves: Step 3-1: Consider the optimization scenario where each IRS has a continuous phase shift, and transform the optimization problem into a MAMDP problem; For the links in step 1, the relevant Rayleigh channel communication systems of the K IRSs are considered as the overall environment E, and the NNs deployed at each IRS are considered as agents, and cannot directly obtain the real-time CSI of their respective channel environments; The model is a MAMDP, denoted as : S represents the state space, defined as the set of reflection coefficient matrices of each IRS and the state information of the communication system: for IRSk, its local state space is defined. Its real-time reflection coefficient state information and its conditional state information; let Represents the state information of IRSk at time t, where: ; Then the state space of the entire system Represented as sub-state spaces The union, namely: ; in, This represents the conditional state information of IRSk at time t; A represents the action space: defined as the space in which each agent adjusts the IRS reflection coefficient; for the k-th agent, its local action space is defined. To provide the space for adjusting the reflection coefficient of IRSk, let This represents the local state of the agent with respect to IRSk at time t. The actions taken utilize a continuous motion space: ; in: It is a constant and should be set in advance before optimization begins; This represents the phase adjustment of IRSk on its M reflection units at time t; The action space A of the entire system is the set of the action spaces of all agents: ; R represents the system's reward space: based on the optimization problem The constraints and settings are defined, with the action reward consisting of a reward term and a penalty term; for the k-th agent, assume that the local state observed at time t is... The action performed is The equivalent channel coefficient of the BS-IRSk-User link is The IRSs are in a cooperative relationship. There are two design methods for the reward obtained by agent k. The first method is to set the same global cooperative reward for all agents: ; This is a penalty term for when the power of the synthesized signal at the receiving end does not meet the minimum requirements for lossy communication or exceeds the maximum power constraint. Its function value is determined by the state information and action information. ; When the equivalent channel coefficient modulus Rewards for each agent when the system's lossy communication feasibility constraint is not met When the maximum power constraint at the receiver is not met Then set as Upper limit constraint value ; Step 3-2: Assign an independent agent to each IRS, using a centralized training-distributed execution CTDE framework; For multi-IRS assisted communication scenarios, the reflection coefficient of multiple IRSs is optimized by distributing intelligent agents to each IRS to independently adjust its reflection coefficient; For any agent k, a DDNN based on a two-branch TD3 framework is used as its core; the Critic network uses two identical neural networks to simultaneously learn two Q functions, evaluates the value of the agent's actions, and selects the smaller Q value for network parameter updates; using This indicates that they are evaluating the Actor network. These represent two evaluation Critic networks, where... and These represent the parameters of the corresponding networks, respectively. Let each represent the target network, and the network parameters be expressed as follows: and ; Each agent employs a CTDE interaction strategy: a Critic network is deployed for each agent. , All are centralized evaluation networks, meaning that during training, they can access the global action database, which contains all the local actions and states of all agents. and global state The Actor network deployed by each agent adopts a distributed execution strategy, with each agent relying solely on its own local observation information. Generate Actions It does not rely on the action information of other agents; all agents share a common experience replay pool. Record global state ,action and reward information This is so that each agent can better learn the strategies that exist in environments with other agents.

3. An electronic device, characterized in that, include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the method as described in any one of claims 1 to 2.

4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 2.

5. A chip, characterized in that, include: A processor for retrieving and running a computer program from memory, causing a device on which the chip is mounted to perform the method as described in any one of claims 1 to 2.

6. A computer program product, characterized in that, The computer program product includes a computer storage medium storing a computer program, the computer program including instructions executable by at least one processor, which, when executed by the at least one processor, implement the method as described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Air-ground non-orthogonal multiple access uplink transmission method based on intelligent reflecting surface

    CN114422056A

  • Intelligent metasurface auxiliary underwater acoustic data transmission method based on non-orthogonal multiple access

    CN117439673A