An Intelligent Reflecting Surface Communication Method Based on Distributed Reinforcement Learning

Through distributed reinforcement learning, optimize beamforming of base stations and intelligent reflection surfaces, the problems of interval interference and eavesdropper impact in 5G networks are solved, and the spectrum efficiency of main users and the improvement of communication security is achieved.

CN115802379BActive Publication Date: 2025-07-22NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211347278.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-07-22
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

In the prior art, the ultra-intensive deployment of 5G networks makes it difficult to eliminate interval interference, the fairness difference between users is large, and the eavesdropper affects communication security, especially when the direct link in millimeter wave communication is blocked, the communication quality of the receiving end is difficult to ensure.

Method used

The distributed reinforcement learning method is adopted, through the joint regulation of the base station and the intelligent reflection surface, the active beamforming vector of the base station and the phase shift matrix of the intelligent reflection surface is optimized, and the communication system is built to maximize the spectrum efficiency of the main user and suppress the received power of the eavesdropper. The deep gradient MADDPG model is used for solving it using the multi-agent deterministic strategy.

Benefits of technology

On the premise of ensuring the communication quality of the main user, it effectively suppresses the receiving power of the eavesdropper, improves the spectrum efficiency and optimizes the coordinated allocation capability of the electromagnetic environment, and solves the impact of the eavesdropper on communication security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115802379B_ABST
    Figure CN115802379B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent reflecting surface communication method based on distributed reinforcement learning, which obtains the initial environmental parameters of the base station and each intelligent reflecting surface. The initial environmental parameters include the initial active beamforming vector of the base station and the initial phase shift matrix of each intelligent reflecting surface. An optimization problem of the communication system is constructed and solved with the maximum spectral efficiency of each primary user, and the final active beamforming vector of the base station and the final phase shift matrix of each intelligent reflecting surface are obtained. Data transmission is carried out by replacing the initial active beamforming vector with the final active beamforming vector and the initial phase shift matrix with the final phase shift matrix. The present invention makes full use of the intelligent reflecting surface and effectively suppresses the received power of eavesdroppers while ensuring the communication quality of primary users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of Internet of Things communication, and particularly relates to an intelligent reflecting surface communication method based on distributed reinforcement learning. Background Art

[0002] Currently, with the implementation of 5G key technologies such as ultra-high density network, massive MIMO system, millimeter wave and other technologies, the capacity of 5G network has been increased by at least 1000 times, and at the same time, the development of innovative applications such as VR and AR has been promoted. According to the ITU report, it is expected that from 2020 to 2032, the mobile network traffic will increase at a rate of 55% per year, and it is expected to reach 5016 EB (Exabyte, 260) / month in 2030. However, the ultra-dense network has led to problems such as a sharp increase in the cost of cell deployment and maintenance, and difficulty in eliminating inter-cell interference, which have also become problems that need to be further solved in B5G (Beyond-5G) and 6G.

[0003] With the development of metamaterial technology, a possible solution is to reconstruct the electromagnetic environment through IRS (Intelligent Reflecting Surface). An array of surface oscillators composed of a large number of low-cost metamaterial units is adopted. Each metamaterial oscillator can be controlled to change its electromagnetic reflectivity characteristics (such as polarization direction, amplitude, phase) through special designs such as geometric shape and crystal arrangement at the molecular scale on the sub-wavelength scale. Through the independent real-time control of each unit by an embedded low-cost controller, passive beamforming of the backscattered signal in a specific direction can be realized, local holes can be solved, electromagnetic pollution can be reduced, edge users can be supported, thereby improving the coordinated allocation ability of the spectrum in the edge space-time. Since IRS is transparent to the user side, IRS can also be regarded as a medium with the ability of spatial reconstruction.

[0004] In the past two years, the research on improving the capacity and quality of mobile communication networks by deploying intelligent reflecting surfaces has attracted extensive attention from researchers and the industry. Especially in compensating for the performance attenuation caused by the occlusion of ultra-high frequency communications such as millimeter wave and terahertz by dense buildings, in the existing general technologies, by simply adjusting the transmitting end array, due to the limited energy of the base station and the fact that the direct link to the user is often blocked, the communication quality of the primary user at the receiving end cannot be guaranteed, and due to the different positions, the fairness among users is different, and there are potential eavesdroppers affecting the security of communication. Summary of the Invention

[0005] The object of the present invention is to provide an intelligent reflecting surface communication method based on distributed reinforcement learning, which restricts the receiving power of eavesdroppers by simultaneously regulating the base station and the intelligent reflecting surface, and avoids potential eavesdroppers from affecting the security of communication.

[0006] The present invention adopts the following technical solution: An intelligent reflecting surface communication method based on distributed reinforcement learning, which is applied to a communication system. The communication system includes a base station and several intelligent reflecting surfaces, and also includes several primary users and several eavesdroppers. The primary users and the eavesdroppers are both wirelessly connected to the base station or wirelessly connected to the base station through the intelligent reflecting surfaces. The method includes the following steps:

[0007] Obtain the initial environmental parameters of the base station and each intelligent reflecting surface. The initial environmental parameters include the initial active beamforming vector of the base station and the initial phase shift matrix of each intelligent reflecting surface.

[0008] Construct and solve an optimization problem for the communication system with the maximum spectral efficiency of each primary user to obtain the final active beamforming vector of the base station and the final phase shift matrix of each intelligent reflecting surface.

[0009] Use the final active beamforming vector to replace the initial active beamforming vector and the final phase shift matrix to replace the initial phase shift matrix for data transmission.

[0010] Furthermore, the optimization problem is:

[0011]

[0012] where p k is the active beamforming vector of the base station for the k-th primary user, Θ i is the phase shift matrix of the i-th intelligent reflecting surface, K is the number of primary users, SINR k is the signal-to-noise ratio of the k-th user, θ g.M is the phase shift of the M-th reflecting unit in the g-th intelligent reflecting surface, P max is the maximum transmission power of the base station, IT e is the received power of the e-th eavesdropper, Γ is the upper bound of the received power of the eavesdropper, and E is the number of eavesdroppers.

[0013] Furthermore, a multi-agent deterministic policy deep gradient MADDPG model is used to solve the optimization problem;

[0014] The reward function of the multi-agent deterministic policy deep gradient MADDPG model is:

[0015]

[0016] where R is the reward value and p is the adjustment parameter.

[0017] Furthermore, the calculation method of IT e is:

[0018]

[0019] where h i,kis the channel state information from the i-th intelligent reflecting surface to the k-th primary user, g i is the channel state information from the base station to the i-th intelligent reflecting surface, I is the number of intelligent reflecting surfaces, W k is the direct transmission link from the base station to the k-th primary user.

[0020] Furthermore,, SINR k is calculated as follows:

[0021]

[0022] where is the variance of the ambient noise.

[0023] Furthermore,, in the multi-agent deterministic policy deep gradient MADDPG model:

[0024] The spectral efficiency combination of K primary users and E eavesdroppers is used as the combined state;

[0025] The combination of the active beamforming vector of the base station and the phase shift matrix of the intelligent reflecting surface is used as the combined action.

[0026] Another technical solution of the present invention: An intelligent reflecting surface communication device based on distributed reinforcement learning, which is applied to a communication system. The communication system includes a base station and several intelligent reflecting surfaces, and also includes several primary users and several eavesdroppers. The primary users and the eavesdroppers are both wirelessly connected to the base station or wirelessly connected to the base station through the intelligent reflecting surface; the device includes:

[0027] An acquisition module, configured to acquire the initial environmental parameters of the base station and each intelligent reflecting surface. The initial environmental parameters include the initial active beamforming vector of the base station and the initial phase shift matrix of each intelligent reflecting surface;

[0028] A solving module, configured to construct and solve an optimization problem of the communication system with the maximum spectral efficiency of each primary user, and obtain the final active beamforming vector of the base station and the final phase shift matrix of each intelligent reflecting surface;

[0029] A replacement module, configured to perform data transmission by replacing the initial active beamforming vector with the final active beamforming vector and the initial phase shift matrix with the final phase shift matrix.

[0030] Another technical solution of the present invention: An intelligent reflecting surface communication device based on distributed reinforcement learning, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor implements the above-mentioned intelligent reflecting surface communication method based on distributed reinforcement learning when executing the computer program.

[0031] Another technical solution of the present invention: An intelligent reflecting surface communication system based on distributed reinforcement learning, including a base station and several intelligent reflecting surfaces of the communication system, and also including several primary users and several eavesdroppers. Both the primary users and the eavesdroppers are wirelessly connected to the base station or wirelessly connected to the base station through the intelligent reflecting surfaces;

[0032] The communication system further includes the above-mentioned intelligent reflecting surface communication device based on distributed reinforcement learning.

[0033] The beneficial effects of the present invention are as follows: By constructing an optimization problem with the maximum spectral efficiency of the primary users and solving it to obtain the active beamforming vector of the base station and the phase shift matrix of each intelligent reflecting surface, the active beamforming and passive beamforming of the communication system are realized with the active beamforming vector and the phase shift matrix, so that the intelligent reflecting surfaces are fully utilized, and the received power of the eavesdroppers is effectively suppressed while ensuring the communication quality of the primary users. Description of the Drawings

[0034] Figure 1 It is a schematic diagram of the architecture of the communication system in the embodiment of the present invention;

[0035] Figure 2 It is a flowchart of the communication method in the embodiment of the present invention;

[0036] Figure 3 It is a schematic diagram of the intelligent reflecting surface unit and the phase shift matrix in the embodiment of the present invention;

[0037] Figure 4 It is a schematic diagram of the structure of the communication device in the embodiment of the present invention;

[0038] Figure 5 It is a schematic diagram of the structure of the communication device in another embodiment of the present invention;

[0039] Figure 6 It is a comparison diagram of the phase shift bits of the reflecting surface in the embodiment of the present invention;

[0040] Figure 7 In the embodiment of the present invention, different P max and the influence comparison diagram of the reflecting surface unit array on the spectral efficiency;

[0041] Figure 8 It is a comparison diagram of the effects of different beamforming methods in the embodiment of the present invention. Detailed Embodiments

[0042] The present invention will be described in detail below in conjunction with the drawings and specific embodiments.

[0043] The present invention discloses an intelligent reflecting surface communication method based on distributed reinforcement learning, which is applied to a communication system. The communication system includes a base station and several intelligent reflecting surfaces, and also includes several primary users and several eavesdroppers. The primary users and the eavesdroppers are both wirelessly connected to the base station or wirelessly connected to the base station through the intelligent reflecting surfaces; as Figure 2 shown, the method includes the following steps: S110, obtaining the initial environmental parameters of the base station and each intelligent reflecting surface, where the initial environmental parameters include the initial active beamforming vector of the base station and the initial phase shift matrix of each intelligent reflecting surface; S120, constructing and solving an optimization problem of the communication system with the maximum spectral efficiency of each primary user to obtain the final active beamforming vector of the base station and the final phase shift matrix of each intelligent reflecting surface; S130, using the final active beamforming vector to replace the initial active beamforming vector and the final phase shift matrix to replace the initial phase shift matrix for data transmission.

[0044] The present invention constructs an optimization problem with the maximum spectral efficiency of the primary users, and solves it to obtain the active beamforming vector of the base station and the phase shift matrix of each intelligent reflecting surface. Then, the active beamforming and passive beamforming of the communication system are realized with the active beamforming vector and the phase shift matrix, so that the intelligent reflecting surfaces are fully utilized, and the received power of the eavesdroppers is effectively suppressed while ensuring the communication quality of the primary users.

[0045] In an embodiment of the present invention, the communication system includes a base station (Base Station), I intelligent reflecting surfaces (Intelligent Reflective Surface), K primary users (Primary User), and E eavesdroppers (Eavesdropper). In the downlink, the base station transmits signals s1, s2,..., s with the same frequency required by K primary users through N antennas. k The signal sources of the primary users and the eavesdroppers at the receiving end consist of two parts: the direct link part and the reflected link part. The direct link part is from the base station to the primary users and the eavesdroppers, and the reflected link part is from the base station through the reflecting surface to the primary users and the eavesdroppers.

[0046] Since the intelligent reflecting surface can modulate the signal phase of the reflected link part, the signal beamforming of the reflected link can be achieved, and it is superimposed with the direct link signal to achieve targeted enhancement and suppression. The base station is equipped with N antennas, which adjust the amplitude and phase of the transmitted signal to form the beamforming at the active end. Each intelligent reflecting surface has M reflection units and is designed as a square planar array to modulate the amplitude and phase of the reflected link signal. All the primary users and eavesdroppers at the receiving end are only equipped with one antenna to receive the signals in the environment. The signal sources at the receiving end include two parts:

[0047] 1) Direct link: This part of the signal is sent from the base station to the user. Since the user and the transmitting base station are often blocked by obstacles, there is no direct line-of-sight (LoS) signal in the direct link BS-User. The Rayleigh fading channel model is adopted.

[0048]

[0049] where ρ0 is the reference loss at a distance of 1 meter, represents the distance between the base station and the k-th user (this user refers to the primary user or the eavesdropper), and α BU is the path loss coefficient from the base station to the user. It is a constant between 2 and 6, that is, the large-scale fading exponent. The values from the base station to the user, from the reflecting surface to the user, and from the base station to the reflecting surface are all different. W k represents the direct link, that is, the wireless transmission link from the base station to the user. obeys a complex Gaussian distribution with zero mean and unit variance, that is, the random scattering component.

[0050] 2) Reflection link: The reflection link is a cascaded channel, consisting of two parts: BS-IRS and IRS-User. The channel state information is represented by g i and h i,k respectively. The BS-IRS channel includes a direct path and a non-direct path, and the Rice fading channel model is adopted, where K is the Rice factor.

[0051]

[0052] where g i is the channel state information from the base station to the i-th intelligent reflecting surface. K1 is the Rice fading coefficient or the Rice factor, which is a constant used to form the Rice fading. The Rice fading is a small-scale fading, and the result of multiplying it by the large-scale fading caused by distance represents the overall fading information. is the direct path information from the base station to the i-th intelligent reflecting surface. Los is the direct path, and nlos is the non-direct path.

[0053]

[0054] v i,0 represents the large-scale fading caused by distance. is the distance from the base station to the i-th intelligent reflecting surface, and α BI represents the large-scale fading coefficient from the base station to the intelligent reflecting surface.

[0055]

[0056] includes the outgoing information of the base station's transmitting antenna and the incoming information of the reflecting surface. is the transmit beamforming vector, represents the incident vector of the reflecting surface, including the azimuth angle and elevation angle of incidence.

[0057]

[0058] where h i,k is the channel state information from the reflecting surface to the user, and v i,k is the large-scale fading from the i-th BIS to the k-th user, is the outgoing vector of the reflecting surface, including the outgoing angle information.

[0059] Therefore, the signal expression at the receiving end can be obtained:

[0060]

[0061] It can be seen that the signal at the receiving end consists of three parts, represents the signal s received by user k k , represents the signals of other users received by user k, and ω represents the noise in the environment, which is complex Gaussian noise satisfying CN(0, δ ω 2 ).

[0062] Therefore, the signal-to-noise ratio (SINR k ) of the k-th user at the receiving end is:

[0063]

[0064] where is the variance of the environmental noise.

[0065] Since there is an eavesdropper in the scenario (the eavesdropper is an illegal user within the coverage of the communication system), the concept of interference temperature is introduced. As the transmit power of the base station increases, there is an upper bound on the received power of the eavesdropper, which is the interference temperature. Therefore, the interference temperature is defined as:

[0066]

[0067] where h i,k is the channel state information from the i-th intelligent reflecting surface to the k-th primary user, g i is the channel state information from the base station to the i-th intelligent reflecting surface, I is the number of intelligent reflecting surfaces, and W k is the direct transmission link from the base station to the k-th primary user.

[0068] Generally, joint beamforming of the base station and the reflecting surface is utilized to enhance the signal quality of users. Usually, when solving such problems, the scenario setting is relatively simple, the number of reflecting surfaces and users is small, and the problem to be solved is a convex problem, so the convex optimization method is often used to solve it. However, in the present invention, due to the relatively complex scenario and the large number of reflecting units, the phase shift of each unit of the reflecting surface and the coefficient of the transmitting antenna of the base station are set by the greedy algorithm, and then the active and passive beamforming is adjusted, so that the direct signal and the non-direct signal form the active and passive joint beamforming, in order to achieve the purpose of reconstructing the electromagnetic propagation environment.

[0069] The transmitting power of the base station end of the present invention has conditional restrictions and shall not exceed the maximum transmitting power P max , assuming that there are N transmitting antennas, and the power on the transmitting antenna corresponding to the active beamforming vector is ||p k || 2 , p k is the active beamforming vector required for the base station to transmit the k-th user, and the restriction condition is that the transmitting power cannot exceed the maximum transmitting power:

[0070]

[0071] Such as Figure 3 shown, is the phase shift matrix of the reflecting surface, and θ is adjusted within the value range g.M , on the premise of not exceeding the maximum transmitting power of P2 (i.e., the constraint condition P2 in Equation 10), the active beamforming vector is adjusted so that all legitimate users can achieve the maximum spectral efficiency, that is, the optimization problem is:

[0072]

[0073] Among them, p k is the active beamforming vector of the base station for the k-th primary user, Θ i is the phase shift matrix of the i-th intelligent reflecting surface, K is the number of primary users, SINR k is the signal-to-noise ratio of the k-th user, θ g.M is the phase shift of the M-th reflecting unit in the g-th intelligent reflecting surface, P max is the maximum transmitting power of the base station, IT e is the received power of the e-th eavesdropper, Γ is the upper bound of the received power of the eavesdropper, and E is the number of eavesdroppers.

[0074] In the present invention, an auxiliary communication device can be added, and the auxiliary communication device can directly send the strategy of the phase shift and amplitude change of the reflecting surface to the controller of each reflecting surface, so that each reflecting surface adjusts its own phase shift matrix according to the corresponding strategy.

[0075] In addition, in order to solve the above optimization problem, the optimization problem can be solved by inputting the above state, action, and reward into MADDPG, thereby determining the best strategy. In this embodiment, the spectrum efficiency combination of K main users and E eavesdroppers is used as the combined state in the multi-agent deterministic strategy deep gradient MADDPG model; the active beamforming vector of the base station and the phase shift matrix combination of the smart reflection surface are used as the combined action.

[0076] The specific solution strategy can be found in Table 1.

[0077] Table 1

[0078]

[0079]

[0080] In summary, the method of this embodiment can be summarized as follows:

[0081] Get the base station and each IRS in I IRS i Environmental parameters in the current scenario, where i = 1, 2, ..., I, environmental parameters include the base station and each IRS i This observation is a local observation, which is the phase shift matrix of each agent at the current moment. The agent makes a strategy based on its own observation and changes to the next phase shift matrix.

[0082] Input the environmental parameters of each intelligent reflective surface in the current scene into the multi-agent deterministic policy deep gradient MADDPG model. In some embodiments, the method further includes: training the MADDPG model with a tuple (S, A, R, S') consisting of state, action and reward.

[0083] Among them, S is the global state, which is the global state formed by the fusion of the results of the current K legitimate users and E potential eavesdroppers. S' is the next state. Action A is the joint action, that is, the actions of the BS and all IRSs are spliced together to form a global action. The reward is the global reward. According to the results obtained at the user end, if the user's spectrum efficiency increases, a positive reward is obtained. If the efficiency decreases or the interference temperature of the potential eavesdropper exceeds the upper bound, a negative reward is obtained.

[0084] The reward function of the multi-agent deterministic policy deep gradient MADDPG model is:

[0085]

[0086] Among them, R is the reward value and p is the adjustment parameter.

[0087] Next, the multi-agent deterministic policy deep gradient MADDPG model outputs the base station and each IRSi The beamforming strategy in the current scenario, where the beamforming strategy includes the active beamforming vector p at the base station side k and the passive beamforming vector Θ of each reflecting surface i .

[0088] Finally, the global state S(t), the global action A(t), the reward R(t), and the state S(t + 1) at the next moment are sent to the experience replay buffer of the deterministic policy gradient MADDPG model to train the model.

[0089] Through the above MADDPG-based optimization algorithm, the best action can be searched in the continuous space, and the fairness among devices and the differences among their served objects are considered simultaneously.

[0090] There are a total of four networks in MADDPG:

[0091] Actor current network: Responsible for the update iteration of the network parameters θ, selects the current action A according to the current state S, interacts with the environment to generate the next state S' and the reward R;

[0092] Actor target network: Responsible for selecting the next action A' from the next state S' sampled from the experience replay buffer. The network parameters θ of this network are periodically copied from the actor current network for θ update;

[0093] Critic current network: Responsible for the update iteration of the value network parameters θ, calculates the current Q value, Q(S, A|θ), that is: y i = R + γQ'(S', A', θ');

[0094] Critic target network: Responsible for calculating the Q'(S', A', θ') part in the target Q value. The network parameters θ' in this network are periodically copied from the Critic current network for θ update.

[0095] MADDPG adopts a "soft" update method of updating only a little bit each time, that is:

[0096] μ' k+1 = τμ' k + (1 - τ)μ' k (12)

[0097] θ' k+1 = τθ' k + (1 - τ)θ' k (13)

[0098] Among them, τ is the update coefficient, and this update method can greatly improve the stability of learning.

[0099] The current network of the Actor uses a deterministic policy to generate deterministic actions, and the loss gradient is:

[0100]

[0101] The loss function of the current network of the Critic uses the mean squared error:

[0102] J(θ) = E[(y k -Q(S, A|θ) 2 ) (15)

[0103] It can be seen from this that the embodiment of the present invention uses the MADDPG method to learn the active and passive beamforming vectors, and each agent has an actor and a critic. Among them, the actor maps the local observation state s n to an appropriate action a n according to the policy network π n , and the critic evaluates the quality of this policy according to its value network Q n . Both the actor and the critic have an online network and a target network to ensure the stability of learning and overcome over-optimism.

[0104] The actor network is trained with its own local observations and actions, while the critic network needs to use global observations and global actions. During the training process, Q n outputs the policy gradient of the policy π n according to the actions and states of other agents. During the execution process, the well-trained π n can independently select the optimal action according to its own state without considering other agents, resulting in less synchronization and communication overhead.

[0105] MADDPG further adds noise N to explore better policies during the training process. Another basic technique is the experience replay buffer (RB). Each agent is equipped with an RB to store (S(t), A(t), r(t), s(t + 1)), which will be randomly extracted to update the weights. In addition, experience replay can effectively avoid highly correlated actions of consecutive updates. MADDPG inherits the DDPG method into the multi-agent domain. It not only eliminates the non-stationary characteristics of DQN and policy gradients, but also retains the great advantages of DDPG. It can search the action space in a continuous rather than discrete manner.

[0106] The assisted communication method based on the reflecting surface provided in this example, while considering the distance and angle of each reflecting surface from the base station, determines the optimal adjustment strategy for the reflecting and base station active beamforming vectors by adopting the MADDPG model, optimizes each agent at the reflecting and base station ends, maximizes the reward at the current time, ensures that all reflecting surfaces can jointly complete the assisted communication task, and at the same time improves the energy consumption efficiency of the system to maximize the utilization of the transmission power of the base station.

[0107] To demonstrate the advantages of the collaborative sensing method based on MADDPG, it was further compared with the DQN method, and the results are as Figure 8 shown (where DQN: an algorithm of reinforcement learning; deep Q network MADDPG: multi-agent reinforcement learning; Attention: attention mechanism; MLP: multi-layer perceptron; LSTM: long short-term memory network; DQN+attention: a method combining the DQN algorithm with the attention mechanism).

[0108] According to Figure 8 it can be seen that the curves from top to bottom are DQN+Attention, DQN+LSTM, DQN+MLP, MADDPG+Attention, MADDPG+LSTM, and MADDPG+MLP in turn. To achieve effective beamforming, the phase shift of each element in the IRS should change in turn. In the MADDPG+Attention method, a self-attention layer is added between the input layer and the output layer of the critic instead of the MLP. Since attention can learn the deep relationship of the input sequence, the learning process is significantly accelerated. LSTM can selectively control the gradient propagation in the neural network to achieve the purpose of accelerating the learning process. Therefore, compared with MADDPG+MLP, MADDPG+LSTM can also converge in less training time. However, regardless of the type of neural network, the reward of MADDPG is better than that of DQN. Although the convergence speed of DQN is much faster, due to the lack of a critic network, the reward is lower than that of MADDPG, resulting in the PBF of IRSs and the ABF of BS not being well optimized.

[0109] In addition, in Figure 6 , the influence of different quantization orders of the phase shift matrix on the results is compared. Regarding the quantization order, the values in the phase shift matrix are between 0 and 2π. The quantization order is 1, that is, 1 bit, that is, 2 values, and the values of all units are two, π and π / 2. The quantization order is 2, that is, 2 bits, that is, 4 values, and the values of all units are four, π / 4, π / 2, 3π / 4, π. The quantization order will affect the situation of taking discrete values in this interval. The larger the quantization order, the more values are taken and the better the effect.

[0110] The value of the continuous phase shift (i.e., θ) can be any value within [0, 2π), the value of the 1-bit quantized phase shift is 0 and π, and the values of the quantized phase shift are 0, π / 2, π, and 3π / 2. As the quantization order increases, the assisting effect of the reflecting surface on user communication and the suppressing effect on eavesdroppers approach the optimal continuous phase case.

[0111] In Figure 7 as the transmission power P of the BS max changes (specifically, the value of P max ranges from 10 dB to 50 dB), and as the unit array of the reflecting surface gradually increases from 9×9 to 18×18, a comparison is made with the case of only active beamforming at the base station (ABF only). It can be seen that the performance of joint beamforming is better than that of ABF only, and as the size of the reflecting surface increases, the performance gap further expands, which once again proves that passive beamforming of the reflecting surface is an effective method to enhance SE and suppress interference.

[0112] The present invention also discloses an intelligent reflecting surface communication device based on distributed reinforcement learning, which is applied to a communication system. The communication system includes a base station and several intelligent reflecting surfaces, and also includes several primary users and several eavesdroppers. The primary users and the eavesdroppers are both wirelessly connected to the base station or wirelessly connected to the base station through the intelligent reflecting surfaces; as Figure 4 shown, the device includes: an acquisition module 210, configured to acquire the initial environmental parameters of the base station and each intelligent reflecting surface, where the initial environmental parameters include the initial active beamforming vector of the base station and the initial phase shift matrix of each intelligent reflecting surface; a solving module 220, configured to construct and solve an optimization problem of the communication system with the maximum spectral efficiency of each primary user, and obtain the final active beamforming vector of the base station and the final phase shift matrix of each intelligent reflecting surface; a replacement module 230, configured to perform data transmission by replacing the initial active beamforming vector with the final active beamforming vector and the initial phase shift matrix with the final phase shift matrix.

[0113] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / modules, due to being based on the same concept as the method embodiment of the present application, for their specific functions and the technical effects brought, please refer to the method embodiment part specifically, and will not be elaborated here.

[0114] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is used as an example. In practical applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional modules are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.

[0115] The present invention also discloses an intelligent reflecting surface communication device 300 based on distributed reinforcement learning, as Figure 5 shown, which includes a memory 310, a processor 320, and a computer program 330 stored in the memory 310 and executable on the processor. When the processor 320 executes the computer program, it implements the above-mentioned intelligent reflecting surface communication method based on distributed reinforcement learning.

[0116] The device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The device may include but is not limited to a processor and a memory. Those skilled in the art can understand that it may include more or fewer components, or combine certain components, or different components. For example, it may also include input and output devices, network access devices, etc.

[0117] The so-called processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0118] In some embodiments, the memory may be an internal storage unit of the extraction device, such as the hard disk or memory of the extraction device. In other embodiments, the memory may also be an external storage device of the extraction device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the extraction device. Further, the memory may also include both the internal storage unit of the extraction device and the external storage device. The memory is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program. The memory may also be used to temporarily store data that has been output or will be output.

[0119] The present invention discloses an intelligent reflecting surface communication system based on distributed reinforcement learning. As Figure 1 shown, the communication system includes a base station and a plurality of intelligent reflecting surfaces, and also includes a plurality of primary users and a plurality of eavesdroppers. The primary users and the eavesdroppers are both wirelessly connected to the base station or wirelessly connected to the base station through the intelligent reflecting surfaces.

Claims

1. An intelligent reflecting surface communication method based on distributed reinforcement learning, characterized in that Applied to a communication system, the communication system includes a base station and a number of intelligent reflecting surfaces, and also includes a number of primary users and a number of eavesdroppers. The primary users and eavesdroppers are both wirelessly connected to the base station or wirelessly connected to the base station through the intelligent reflecting surfaces. The method includes the following steps: Obtain the initial environmental parameters of the base station and each intelligent reflecting surface. The initial environmental parameters include the initial active beamforming vector of the base station and the initial phase shift matrix of each intelligent reflecting surface. Construct and solve an optimization problem of the communication system with the maximum spectral efficiency of each primary user to obtain the final active beamforming vector of the base station and the final phase shift matrix of each intelligent reflecting surface. Use the final active beamforming vector to replace the initial active beamforming vector and the final phase shift matrix to replace the initial phase shift matrix for data transmission. The optimization problem is: , Among them, is the active beamforming vector of the base station for the k-th primary user, is the phase shift matrix of the i-th intelligent reflecting surface, K is the number of primary users, is the signal-to-noise ratio of the k-th user, is the phase shift of the M-th reflecting element in the g-th intelligent reflecting surface, is the maximum transmit power of the base station, is the received power of the e-th eavesdropper, is the upper bound of the received power of the eavesdropper, and E is the number of eavesdroppers; Use the multi-agent deterministic policy deep gradient MADDPG model to solve the optimization problem. The reward function of the multi-agent deterministic policy deep gradient MADDPG model is: , Among them, is the reward value, is the adjustment parameter.

2. The intelligent reflecting surface communication method based on distributed reinforcement learning according to claim 1, wherein The said calculation method is as follows: , Among them, is the channel state information from the i-th intelligent reflecting surface to the k-th primary user, is the channel state information from the base station to the i-th intelligent reflecting surface, and I is the number of intelligent reflecting surfaces, is the direct transmission link from the base station to the k-th primary user.

3. The intelligent reflecting surface communication method based on distributed reinforcement learning according to claim 2, characterized in that, The calculation method is as follows: , Among them, is the variance of the environmental noise.

4. The intelligent reflecting surface communication method based on distributed reinforcement learning according to claim 3, characterized in that In the multi-agent deterministic policy deep gradient MADDPG model: Use the spectral efficiency combination of K primary users and E eavesdroppers as the combined state. Use the combination of the active beamforming vector of the base station and the phase shift matrix of the intelligent reflecting surface as the combined action.

5. An intelligent reflecting surface communication device based on distributed reinforcement learning, characterized in that, Applied to a communication system, the communication system includes a base station and a number of intelligent reflecting surfaces, and also includes a number of primary users and a number of eavesdroppers. The primary users and eavesdroppers are both wirelessly connected to the base station or wirelessly connected to the base station through the intelligent reflecting surfaces. The device includes: An acquisition module for obtaining the initial environmental parameters of the base station and each intelligent reflecting surface. The initial environmental parameters include the initial active beamforming vector of the base station and the initial phase shift matrix of each intelligent reflecting surface. A solving module for constructing and solving an optimization problem of the communication system with the maximum spectral efficiency of each primary user to obtain the final active beamforming vector of the base station and the final phase shift matrix of each intelligent reflecting surface. A replacement module for using the final active beamforming vector to replace the initial active beamforming vector and the final phase shift matrix to replace the initial phase shift matrix for data transmission. The optimization problem is: , Among them, is the active beamforming vector of the base station for the k-th primary user, is the phase shift matrix of the i-th intelligent reflecting surface, K is the number of primary users, is the signal-to-noise ratio of the k-th user, is the phase shift of the M-th reflecting element in the g-th intelligent reflecting surface, is the maximum transmit power of the base station, is the received power of the e-th eavesdropper, is the upper bound of the received power of the eavesdropper, and E is the number of eavesdroppers; Use the multi-agent deterministic policy deep gradient MADDPG model to solve the optimization problem. The reward function of the multi-agent deterministic policy deep gradient MADDPG model is: , Among them, is the reward value, is the adjustment parameter.

6. An intelligent reflecting surface communication device based on distributed reinforcement learning, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements a method for intelligent reflecting surface communication based on distributed reinforcement learning according to any one of claims 1-4.

7. An intelligent reflecting surface communication system based on distributed reinforcement learning, characterized in that, The base station of the communication system and a number of intelligent reflecting surfaces also include a number of primary users and a number of eavesdroppers. The primary users and eavesdroppers are both wirelessly connected to the base station or wirelessly connected to the base station through the intelligent reflecting surfaces. The communication system also includes an intelligent reflecting surface communication device based on distributed reinforcement learning according to claim 5 or 6.