Internet of Things security resource optimization method and system based on multi-agent reinforcement learning
By optimizing the reflection coefficient of controllable reflective elements and base station beamforming through multi-agent reinforcement learning, the problem of insufficient power-phase coupling optimization in IoT systems is solved, and resource allocation optimization and secure transmission rate improvement are achieved in multi-user NOMA systems.
Patent Information
- Application Number
- CN202511428325.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-09-30
AI Technical Summary
Existing technologies have failed to effectively address the coupling optimization relationship between power and phase in IoT systems, resulting in non-globally optimal resource allocation and a lack of a resource intelligent decision-making architecture for multi-agent collaborative optimization. This makes it difficult to achieve synergistic enhancement of channel quality and interference management in a high-dimensional hybrid action space.
A method for optimizing IoT security resources based on multi-agent reinforcement learning is constructed. By cooperating with multiple agents, the reflection coefficient of controllable reflective elements, base station beamforming and power allocation coefficients are optimized. The MADDPG and D3QN dual-agent collaborative optimization framework is adopted to achieve programmable control of signal propagation paths and interference cancellation. Combined with power superposition and beamforming, the signal-to-interference ratio and the sensing gain of leakage direction are optimized.
This method achieves optimal allocation of power resources in a multi-user NOMA system, enhances the channel gain of legitimate users, suppresses eavesdropping link reflections, improves the overall secure transmission rate and spectrum utilization efficiency of the system, and effectively solves the problem of non-globally optimal resource allocation in traditional methods.
Smart Images

Figure CN120916144A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of data processing, and particularly relates to an Internet of Things security resource optimization method and system based on multi-agent reinforcement learning. BACKGROUND
[0002] The exponential growth of the number of Internet of Things terminals requires the system to support the simultaneous access of a large number of devices, and the traditional communication architecture faces multiple challenges such as spectrum resource shortage, serious interference and insufficient transmission security.
[0003] On the one hand, in terms of spectrum efficiency, orthogonal multiple access (OMA) technology is difficult to meet the high-density device connection demand due to resource allocation isolation. Non-orthogonal multiple access (NOMA) realizes multiple user sharing of the same frequency resource through power domain multiplexing, which has a significant advantage in improving system capacity, but at the same time introduces serious intra-group and inter-group co-channel interference problems, affecting system stability. On the other hand, in terms of security performance, the traditional network relies on encryption protocols to protect information security, but the open nature of the wireless channel makes it easy for passive eavesdroppers to listen to the communication content, especially in the IoT environment with limited edge perception and terminal computing capability, the traditional cryptographic mechanism is difficult to resist physical layer attacks. It is an urgent need to build a security enhancement mechanism centered on the physical layer.
[0004] To further improve the transmission performance, in recent years, a reconfigurable intelligent surface is introduced as an auxiliary node to intelligently control the signal reflection path, enhance the link quality and weaken the interference. However, this also brings complexity to resource scheduling. In the RIS-NOMA system, the continuous power allocation at the base station side and the discrete phase matrix at the RIS side need to be jointly optimized, forming a high-dimensional mixed action space (i.e., a joint optimization problem containing continuous and discrete variables). Traditional reinforcement learning algorithms (such as DDPG, PPO) have difficulty in effectively converging in high-dimensional discrete action space, and there are problems such as unstable training and degradation of strategy performance.
[0005] In the prior art, although some patents have tried to introduce deep reinforcement learning (Deep Reinforcement Learning, DRL) mechanism to improve the performance of RIS-NOMA system, there are still many deficiencies. For example, patent CN120151897A proposes a RIS phase control method based on single-agent DRL, which only optimizes the phase modulation and does not consider the coordination relationship with the base station power control, resulting in incomplete system resource configuration and limited performance improvement. Patent CN119149200A adopts a stage-wise separate optimization strategy, modeling and solving power allocation and phase control separately, which reduces the complexity of model design and training, but due to the lack of global coordination mechanism, the overall spectrum utilization efficiency and security performance of the system are not optimal.
[0006] Furthermore, patent CN117320043A introduces a double-RIS structure in the millimeter wave communication scenario and optimizes the user's position information scheduling combined with the DRL method, but does not involve the resource allocation strategy at the system level, especially lacks the modeling and implementation of the power and phase joint control mechanism. Although patent CN119363779A proposes a MADDPG-based offloading power optimization method, it realizes intelligent allocation of power resources, but does not introduce the RIS phase shift optimization part, so it is difficult to realize the synergistic enhancement of channel quality and interference management in the spatial dimension.
[0007] In summary, the current technical solutions either ignore the coupling optimization relationship between power and phase, or use separate modeling to lead to suboptimal system performance, or only focus on a single control dimension without establishing a complete perception-control closed-loop mechanism. Therefore, how to face the large-scale Internet of Things scenario and build a resource intelligent decision-making architecture that adapts to high-dimensional mixed action space and supports multi-agent collaborative optimization is still a key problem that needs to be solved in the current technical field. SUMMARY
[0008] In view of the deficiencies of the prior art, the present application proposes an Internet of Things security resource optimization method based on multi-agent reinforcement learning, comprising: S1: constructing a downlink scenario, the downlink scenario specifically comprising: a base station equipped with an antenna and a user group distributed near the base station communicate with each other; each user group consists of a user, and a controllable reflecting element with controllable reflecting elements are communicatively deployed between the base station BS and the users; S2: constructing a signal transmission mechanism, assuming that the channel state information is in an available state, the controllable reflecting elements adjust the reflection coefficient of the controllable reflecting elements based on an intelligent controller, and the base station generates a corresponding beamforming vector; S3: obtaining the signal-to-interference ratio of the users, including eavesdroppers and legitimate users; based on the difference in signal-to-interference ratio between legitimate users and eavesdroppers, obtaining the security transmission rate of each user; S4: constructing a joint optimization framework MAD2RA, which optimizes the reflection coefficient of the controllable reflecting elements, the base station beamforming and the power allocation coefficient through multi-agent collaboration to maximize the total security transmission rate.
[0009] Further, in S1: the base station uses power superposition coding in each non-orthogonal multiple access group, and combines beamforming, specifically: the base station superimposes the power of the first Within the user group Multiplexing user signals in the power domain is represented as follows: ; in, Indicates the first Within the user group Normalized data symbols for each user , Indicates assignment to the first Total power of each user group Indicates the first The user group Power allocation factor for each user Indicates the first Within the user group The signal after power domain multiplexing for each user; In S2, the base station applies a corresponding beamforming vector to the signal of each user group, as follows: ; ; in, Indicates that the base station is targeting the first The signal vector output after applying beamforming vectors to each user group; Represents a beamforming set. It is the first Beamforming vectors corresponding to users in each user group; Indicates the first The number of users contained in a user group, i.e., the user index within the group. The upper limit; Indicates user group index The upper limit; The base station simultaneously transmits superimposed signals to all user groups. , represented as: ; The user receives signals, which include target signals, intra-group interference, and inter-group interference. The user's received signal corresponds to a user link that includes a direct link from the base station to the user and a reflected link through the base station to a controllable reflection unit and then to the user, represented as: ; in, This represents the channel matrix from the base station (BS) to the controllable reflective element (RIS), which is also the reflection link from the controllable reflective element to the user. This indicates the reflection link from the controllable reflective element to the user. It is a direct link from the base station to the user; Indicates the reflection link at the user's receiving end; This represents the reflection coefficient of a controllable reflective element; The reflection link at the user receiver satisfies the SIC decoding order of non-orthogonal multiple access; Based on this, the first Within the group Signal received by a user Represented as: ; in, For the first Within the group Gaussian white noise received by each user; Represents the set of distractors; Indicates that the base station is for the first Within the group Signals transmitted by individual users; Based on user location distribution, user groups are grouped through a two-dimensional joint optimization. The objective function for dynamic grouping is as follows: ; in, This indicates a maximize operation; This represents the set of grouped indexes, which is also the optimization variable; They represent the first The first in the group The channel matrix corresponding to the user and the first user The first in the group Channel matrix corresponding to each user; This represents the conjugate transpose operation; Assuming the user's location is The location of the base station is ; in, The first The x-axis coordinates, y-axis coordinates, and z-axis coordinates of each user in three-dimensional space; These are the x-axis, y-axis, and z-axis coordinates in the three-dimensional spatial coordinates of the base station, respectively. The channel response, including path loss and Ricean fading, is obtained as follows: ;
[0010] ; in, Rice factor; The direct path loss index; Indicates the first The three-dimensional spatial distance between each user and the base station; denotes a first reference path loss factor; denotes a non-line-of-sight channel component vector; denotes a line-of-sight channel component vector; The channel of the controllable reflecting element reflecting link obeys a Rician distribution, denoted as: ; wherein, denotes the conjugate transpose of the channel matrix observed by the th user in the controllable reflecting element reflecting link; denotes a second reference path loss factor; denotes the distance from the base station to the controllable reflecting element; denotes the distance from the controllable reflecting element to the th user, denotes the conjugate transpose of the channel matrix from the base station to the controllable reflecting element.
[0011] Further, S3 is specifically: obtaining the signal-to-interference ratio of the legitimate user denoted as: ; wherein, denotes the transmit power of the th legitimate user; denotes the channel gain of the th legitimate user in the th group; denotes the conjugate transpose of the channel matrix of the th legitimate user in the th group; denotes the effective gain of the expected signal; denotes the transmit power of the th legitimate user; denotes the beamforming vector corresponding to the legitimate user in the th user group; obtaining the received signal of the eavesdropper, denoted as: ; wherein, denotes the received signal of the eavesdropper, denotes the total number of eavesdroppers; denotes the channel matrix of the eavesdropper; denotes the total number of users in the th user group; denotes the transmit signal of the th user in the th group; denotes the Gaussian white noise at the receiving end of the eavesdropper; Obtain the eavesdropper's equivalent channel and minimize the equivalent gain of the eavesdropper's equivalent channel by optimizing the reflection coefficient of the controllable reflection element; Obtain the signal-to-interference ratio corresponding to the eavesdropper. , represented as: ; The difference in capabilities between legitimate users and eavesdroppers is quantified by secure transmission rate, and expressed as follows: ; in, Indicates the first The first in the group Secure transmission rate for each user; Indicates the reachable rate for legitimate users; Indicates the reachability of the eavesdropper; This indicates the non-negation operation.
[0012] Furthermore, in S3, a sensing gain constraint in the leakage direction is introduced. ,in, and These represent the polar angle and azimuth angle, respectively, indicating the direction of the leaked potential target. The leakage beam gain of the system in the direction of a potential eavesdropper is expressed as: ; in, For ISAC, a leak detection threshold is required. Introducing a leakage control weight coefficient λ to ensure secure communication assisted by the sensing function, the optimization objective is expressed as: ; in, This represents the sum of all valid user precoded vectors; This represents the conjugate transpose of the eavesdropping channel in the corresponding direction; This indicates the leakage gain threshold.
[0013] Furthermore, S4 specifically refers to: A joint optimization framework, MAD2RA, is constructed to optimize the reflection coefficient matrix, beamforming vector, and power allocation coefficient of controllable reflective elements through multi-agent collaboration to maximize the overall system safety rate. The optimization task is divided into optimization objective functions corresponding to the base station side and the RIS side.
[0014] Furthermore, the optimization process on the base station side is represented as follows: S401: Constructing an environment that includes mutually communicating base stations, users, and controllable reflective elements; S402: Construct a multi-agent architecture MADDPG based on the environment, set each user's channel in the environment as a user channel agent; the multi-agent architecture includes a policy network and a target network; S403: The policy network of the multi-agent architecture obtains the global state of the user channel agent in the time slot , the global state including the equivalent channel from the base station to the user and the intra-group interference and inter-group interference in the time slot ; S404: The policy network of the multi-agent architecture generates the power allocation coefficient of the user channel agent based on the global state; S405: Calculate the total safe transmission rate based on the power allocation coefficient, input the total safe transmission rate as a reward to the target network of the multi-agent architecture, the target network of the multi-agent architecture calculates the time difference error, and reversely updates the weight parameters of the policy network; and re-execute step S3.
[0015] Further, in S4, the optimization process on the RIS side is represented as: S406: Construct an environment shared with the base station side, the environment including the base station, the legitimate user, the controllable reflecting element and the potential eavesdropper in mutual communication; the optimization target of the RIS side is to adjust the phase shift matrix of the reflecting element to maximize the equivalent channel gain of the legitimate user, reduce the eavesdropping link gain and construct a null in the potential leakage direction; S407: Construct a single-agent architecture D3QN, set the controllable reflecting element as a central control agent, that is, an RIS agent; the single-agent architecture includes an evaluation network and a target network; S408: The evaluation network of the single-agent architecture obtains the state of the RIS agent in the time slot t, the state including the set of power allocation coefficients output by S4 on the base station side, the phase shift matrix of the current controllable reflecting element, the equivalent channel from the base station to the legitimate user and the SINR proportion of the eavesdropping link; S409: The evaluation network of the single-agent architecture generates the phase shift action of the controllable reflecting element based on the state of the RIS agent in the time slot t obtained in S408, the phase shift action uniformly quantizes the phase shift of each reflecting element into Q levels, and outputs an N-dimensional phase level vector to correspond to the phase shift matrix of different reflecting elements; S410: Based on the phase shift matrix, adjust the reflection coefficient and beamforming vector of the controllable reflection element; calculate the total secure transmission rate based on this and use it as a reward input to the target network of the single-agent architecture; simultaneously record experience samples including the current state, phase shift action, reward, and next time slot state of the RIS agent, and adjust the sampling priority of the experience samples in the priority experience replay pool according to the change of the signal-to-interference ratio of the eavesdropping link; the target network calculates the target Q value by combining the experience samples, compares it with the current Q value of the evaluation network to construct a loss function, minimizes the loss function to update the weight parameters of the evaluation network in reverse; then repeat step S408.
[0016] Furthermore, the optimization objective function on the base station side is specifically: given the reflection coefficient of the controllable reflective element, optimize the power allocation coefficient and the beamforming vector; formally, it is: ; in, Indicates the total secure transmission rate; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The transmission rate of each legitimate user; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The transmission rate of an eavesdropper; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The interference ratio of a legitimate user; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The interference ratio of a single eavesdropper's signal; This indicates the maximum allocated power.
[0017] Furthermore, the optimization objective function on the RIS side is specifically: given the power allocation coefficient and beamforming vector, optimize the reflection coefficient of the controllable reflective element, formalized as: ; in, This indicates that, given the power allocation coefficients and beamforming vector, the first... Group 1 The transmission rate of each legitimate user; This indicates that, given the power allocation coefficients and beamforming vector, the first... Group 1 The transmission rate of an eavesdropper; This indicates that, given the power allocation coefficients and beamforming vector, the first... Group 1 SINR of a legitimate user; SINR of the i-th legitimate user given the power allocation factor and the beamforming vector; SINR of the i-th legitimate user given the power allocation factor and the beamforming vector; SINR of the i-th legitimate user given the power allocation factor and the beamforming vector; SINR of the i-th legitimate user given the power allocation factor and the beamforming vector;
[0018] The application discloses a multi-agent reinforcement learning-based security resource optimization system for Internet of Things, which is used for implementing the multi-agent reinforcement learning-based security resource optimization method for Internet of Things. The scene construction module is used for constructing a downlink scene. The signal transmission module is used for constructing a signal transmission mechanism. The data acquisition module is used for acquiring the SINR of users, wherein the users include legitimate users and eavesdroppers. The joint optimization module is MAD2RA, which is used for optimizing the reflection coefficient of the controllable reflecting element, the base station beamforming vector and the user power allocation factor through multi-agent cooperation to maximize the total security transmission rate of the system.
[0019] Compared with the prior art, the application has the following beneficial effects:
[0020] 1. The application deploys RIS with multiple controllable reflecting units between the base station and the user, realizes programmable control of the signal propagation path.
[0021] 2. The application constructs a collaborative optimization framework based on MADDPG and D3QN double agents.
[0022] 3. The application jointly considers the base station power coefficient, RIS phase shift configuration, channel CSI and eavesdropping link characteristics, and constructs a hybrid state space; the power action is a continuous variable, and the phase shift is a multi-level discrete variable, forming a high-dimensional hybrid action space. For this space, the application realizes effective modeling and rapid convergence through a double network architecture and an experience replay mechanism. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 A system framework diagram of an Internet of Things security resource optimization system based on multi-agent reinforcement learning. DETAILED DESCRIPTION
[0024] REFERENCE Figure 1 A method for optimizing security resources of an Internet of Things based on multi-agent reinforcement learning, comprising: S1: Constructing a downlink scenario, the downlink scenario specifically being: A base station equipped with an antenna and a group of users distributed near the base station communicate with each other; each user group is composed of users, and a controllable reflecting element is communicatively deployed between the base station BS and the users; Obtain the power allocation coefficient of each user; S2: Constructing a signal transmission mechanism, assuming that the channel state information is available, the controllable reflecting element adjusts the reflection coefficient of the controllable reflecting element based on an intelligent controller, and the base station generates a corresponding beamforming vector; S3: Obtain the signal-to-interference ratio of the users, including eavesdroppers and legitimate users; based on the difference in signal-to-interference ratio between legitimate users and eavesdroppers, obtain the security transmission rate of each user; S4: Constructing a joint optimization framework MAD2RA, which optimizes the reflection coefficient of the controllable reflecting element, the base station beamforming and the power allocation coefficient through multi-agent cooperation to maximize the total security transmission rate.
[0025] Further, in S1: The base station uses power superposition coding in each non-orthogonal multiple access group, combined with beamforming, specifically: The base station performs power domain multiplexing on the signal of the user in the user group, represented as: ; Wherein, represents the normalized data symbol of the user in the user group, , Indicates assignment to the first Total power of each user group Indicates the first The user group Power allocation factor for each user Indicates the first Within the user group The signal after power domain multiplexing for each user; In S2, the base station applies a corresponding beamforming vector to the signal of each user group, as follows: ; ; in, Indicates that the base station is targeting the first The signal vector output after applying beamforming vectors to each user group; Represents a beamforming set. It is the first Beamforming vectors corresponding to users in each user group; Indicates the first The number of users contained in a user group, i.e., the user index within the group. The upper limit; Indicates user group index The upper limit; The base station simultaneously transmits superimposed signals to all user groups. , represented as: ; The user receives signals, which include target signals, intra-group interference, and inter-group interference. The user's received signal corresponds to a user link that includes a direct link from the base station to the user and a reflected link through the base station to the controllable reflection unit and then to the user, represented as: ; in, This represents the channel matrix from the base station (BS) to the controllable reflective element (RIS), which is also the reflection link from the controllable reflective element to the user. This indicates the reflection link from the controllable reflective element to the user. It is a direct link from the base station to the user; Indicates the reflection link at the user's receiving end; This represents the reflection coefficient of a controllable reflective element; The reflection link at the user receiver satisfies the SIC decoding order of non-orthogonal multiple access; Based on this, the first Within the group Signal received by a user Represented as: ; wherein, is the Gaussian white noise received by the th user in the th group; denotes a set of interference terms; denotes a signal transmitted by the base station to the th user in the th group; grouping users into groups by two-dimensional joint optimization based on user location distribution; The dynamic grouping objective function is as follows: ; wherein, denotes a maximization operation; denotes a set of group indices, i.e., optimization variables; denote the channel matrix corresponding to the th user in the th group and the channel matrix corresponding to the th user in the th group, respectively; denotes a conjugate transpose operation; Assume that the user location is and the base station location is ; wherein, are the x-axis coordinate, y-axis coordinate and z-axis coordinate in the three-dimensional space of the th user, respectively; are the x-axis coordinate, y-axis coordinate and z-axis coordinate in the three-dimensional space of the base station, respectively; The channel response containing path loss and Rician fading is obtained and denoted as: ;
[0026] ; wherein, is a Rician factor; is a direct path loss exponent; denotes the three-dimensional spatial distance between the th user and the base station; denotes a first reference path loss factor; denotes a non-line-of-sight channel component vector; denotes a line-of-sight channel component vector; The channel of the controllable reflecting element reflection link is subject to Rician distribution and denoted as: ; wherein, Indicating the first step in the reflection link of a controllable reflective element The conjugate transpose of the channel matrix observed by each user; Indicates the loss factor of the second reference path; This indicates the distance from the base station to the controllable reflective element; Indicates the controllable reflective element up to the first Distance between users This represents the conjugate transpose of the channel matrix from the base station to the controllable reflective element.
[0027] Furthermore, S3 specifically refers to: Obtain the signal interference ratio of legitimate users , represented as: ; in, Indicates the first The transmit power of each legitimate user; Indicates the first The first in the group Channel gain for each legitimate user; Indicates the first The first in the group The conjugate transpose of a valid user channel matrix; Indicates the effective gain of the desired signal; Indicates the first The transmit power of each legitimate user; Indicates the first Beamforming vectors corresponding to legitimate users in a user group; The received signal from the eavesdropper is represented as: ; in, This indicates the signal received by the eavesdropper. Indicates the total number of eavesdroppers; The channel matrix represents the eavesdropper; Indicates the first The total number of users within a user group; Indicates the first Within the group The transmitted signal of each user; This represents the Gaussian white noise at the receiver's end; Obtain the eavesdropper's equivalent channel and minimize the equivalent gain of the eavesdropper's equivalent channel by optimizing the reflection coefficient of the controllable reflection element; Obtain the signal-to-interference ratio corresponding to the eavesdropper. , represented as: ; The difference in capabilities between legitimate users and eavesdroppers is quantified by secure transmission rate, and expressed as follows: ; wherein, denotes the secure transmission rate of the th user in the th group; denotes the achievable rate of a legitimate user; denotes the achievable rate of an eavesdropper; denotes the take non-negative operation.
[0028] Further, in S3, the perception gain constraint of the eavesdropping direction is introduced wherein, and denote the polar angle and the azimuth angle of the eavesdropping potential target direction, respectively, denotes the eavesdropping beam gain of the system in a certain potential eavesdropping direction, and is expressed as: ; wherein, is the ISAC perception eavesdropping threshold; The eavesdropping control weight coefficient λ is introduced to guarantee the perception function to assist the secure communication, and the optimization objective is expressed as: ; wherein, denotes the sum of the precoding vectors of all legitimate users; denotes the conjugate transpose of the eavesdropping channel in the corresponding direction; denotes the eavesdropping gain threshold.
[0029] Further, S4 is specifically: A joint optimization framework MAD2RA is constructed, and the reflection coefficient matrix of the controllable reflecting element, the beamforming vector and the power allocation coefficient are optimized by multi-agent cooperation to maximize the total secure rate of the system; the optimization task is divided into optimization objective functions corresponding to the base station side and the RIS side.
[0030] Further, the optimization process at the base station side is expressed as: S401: An environment is constructed, which includes a base station, a user and a controllable reflecting element in communication with each other; S402: A multi-agent architecture MADDPG is constructed based on the environment, and the channel of each user in the environment is set as a user channel agent; the multi-agent architecture includes a policy network and a target network; S403: The policy network of the multi-agent architecture obtains the global state of the user channel agent in the time slot , which includes the equivalent channel from the base station to the user and the intra-group interference and inter-group interference in the time slot ; S404: The policy network of the multi-agent architecture generates a power allocation coefficient of the user channel agent based on the global state; S405: Based on the power allocation coefficient, the total safe transmission rate is calculated, and the total safe transmission rate is input as a reward to the target network of the multi-agent architecture. The target network of the multi-agent architecture calculates the time difference error, and reversely updates the weight parameters of the policy network; and re-executes step S3.
[0031] Further, in S4, the optimization process on the RIS side is represented as: S406: An environment shared with the base station side is constructed, and the environment includes a base station, a legitimate user, a controllable reflecting element, and a potential eavesdropper in communication with each other. The optimization target on the RIS side is to adjust the phase shift matrix of the reflecting element to maximize the equivalent channel gain of the legitimate user, reduce the link gain of the eavesdropper, and construct a null in the potential leakage direction; S407: A single-agent architecture D3QN is constructed, and the controllable reflecting element is set as a central control agent, that is, an RIS agent. The single-agent architecture includes an evaluation network and a target network; S408: The evaluation network of the single-agent architecture obtains the state of the RIS agent at time slot t, which includes a set of power allocation coefficients output by the base station side S4, a phase shift matrix of the current controllable reflecting element, an equivalent channel from the base station to the legitimate user, and a SINR ratio of the eavesdropping link; S409: The evaluation network of the single-agent architecture generates a phase shift action of the controllable reflecting element based on the state of the RIS agent at time slot t obtained in S408. The phase shift action uniformly quantizes the phase shift of each reflecting element into Q levels, and outputs an N-dimensional phase level vector to correspond to the phase shift matrix of different reflecting elements; S410: Based on the phase shift matrix, the reflection coefficient and the beamforming vector of the controllable reflecting element are adjusted. Based on this, the total safe transmission rate is calculated and input as a reward to the target network of the single-agent architecture. Experience samples including the current state of the RIS agent, the phase shift action, the reward, and the next time slot state are recorded, and the sampling priority of the experience samples in the priority experience replay pool is adjusted according to the change of the signal-to-interference ratio of the eavesdropping link. The target network calculates the target Q value in combination with the experience samples, compares it with the current Q value of the evaluation network to build a loss function, minimizes the loss function to reversely update the weight parameters of the evaluation network, and then re-executes step S408.
[0032] Further, the optimization objective function on the base station side is specifically: given the reflection coefficient of the controllable reflecting element, the power allocation coefficient and the beamforming vector are optimized; which is formalized as: ; Wherein, represents the total safe transmission rate; represents the transmission rate of the i-th legitimate user given the reflection coefficient of the controllable reflecting element; represents the transmission rate of the i-th legitimate user given the reflection coefficient of the controllable reflecting element; represents the transmission rate of the i-th eavesdropper given the reflection coefficient of the controllable reflecting element; represents the transmission rate of the i-th eavesdropper given the reflection coefficient of the controllable reflecting element; represents the transmission rate of the i-th legitimate user given the reflection coefficient of the controllable reflecting element; represents the transmission rate of the i-th eavesdropper given the reflection coefficient of the controllable reflecting element; represents the transmission rate of the i-th legitimate user given the reflection coefficient of the controllable reflecting element; represents the transmission rate of the i-th eavesdropper given the reflection coefficient of the controllable reflecting element; represents the transmission rate of the i-th legitimate user given the reflection coefficient of the controllable reflecting element; represents the transmission rate of the i-th eavesdropper given the reflection coefficient of the controllable reflecting element; represents the transmission rate of the i-th legitimate user given the reflection coefficient of the controllable reflecting element; represents the transmission rate of the i-th eavesdropper given the reflection coefficient of the controllable reflecting element; represents the maximum allocated power.
[0033] Further, the optimization objective function on the RIS side is specifically: optimizing the reflection coefficient of the controllable reflecting element given the power allocation coefficient and the beamforming vector, which is formalized as: ; wherein, represents the transmission rate of the i-th legitimate user given the power allocation coefficient and the beamforming vector; represents the transmission rate of the i-th legitimate user given the power allocation coefficient and the beamforming vector; represents the transmission rate of the i-th eavesdropper given the power allocation coefficient and the beamforming vector; represents the transmission rate of the i-th eavesdropper given the power allocation coefficient and the beamforming vector; represents the transmission rate of the i-th legitimate user given the power allocation coefficient and the beamforming vector; represents the transmission rate of the i-th eavesdropper given the power allocation coefficient and the beamforming vector; represents the transmission rate of the i-th legitimate user given the power allocation coefficient and the beamforming vector; represents the transmission rate of the i-th eavesdropper given the power allocation coefficient and the beamforming vector; represents the transmission rate of the i-th legitimate user given the power allocation coefficient and the beamforming vector; represents the transmission rate of the i-th eavesdropper given the power allocation coefficient and the beamforming vector; represents the transmission rate of the i-th legitimate user given the power allocation coefficient and the beamforming vector; represents the transmission rate of the i-th eavesdropper given the power allocation coefficient and the beamforming vector; represents the phase shift matrix of the controllable reflecting element.
[0034] The Internet of Things security resource optimization system based on multi-agent reinforcement learning is used to implement the above-mentioned Internet of Things security resource optimization method based on multi-agent reinforcement learning, and is characterized by comprising: a scene construction module for constructing a downlink scene; a signal transmission module for constructing a signal transmission mechanism, assuming that channel state information (CSI) is available, and controlling the controllable reflecting element (RIS) to adjust the phase shift based on the intelligent controller to make the base station generate a corresponding beamforming vector; A data acquisition module is configured to acquire the signal-to-interference-and-noise ratios of users, including eavesdroppers and legitimate users, and acquire the secure transmission rate of each user based on the difference between the signal-to-interference-and-noise ratios of the legitimate users and the eavesdroppers. A joint optimization module, which is MAD2RA, is configured to maximize the total secure transmission rate of the system by optimizing the reflection coefficients of the controllable reflecting elements, the base station beamforming vectors and the user power allocation coefficients through multi-agent cooperation.
[0035] The specific embodiments described above are intended to be illustrative of the present application and should not be construed as limiting the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A method for optimizing security resources of an Internet of Things based on multi-agent reinforcement learning, characterized in that, Comprise: S1: Constructing a downlink scenario, the downlink scenario is specifically: A base station equipped with an antenna and a group of users distributed around the base station communicate with each other; each group of users consists of one user, and a controllable reflecting element is communicatively disposed between the base station BS and the user ; a power allocation coefficient of each user is obtained; S2: Construct a signal transmission mechanism, assuming that the channel state information is available, and the controllable reflecting element Adjust the reflection coefficient of the controllable reflecting element based on the intelligent controller, and the base station generates the corresponding beamforming vector; S3: Obtaining the signal-to-interference ratio of users, including eavesdroppers and legitimate users; Based on the difference in signal-to-interference ratio between legitimate users and eavesdroppers, the safe transmission rate of each user is obtained; S4: Constructing a joint optimization framework MAD2RA, the joint optimization framework optimizes the reflection coefficient of the controllable reflecting element, the beamforming of the base station and the power allocation coefficient through multi-agent cooperation to maximize the total safe transmission rate.
2. The method of claim 1, wherein, In S1: The base station uses power superposition coding in each non-orthogonal multiple access group, combined with beamforming, specifically: Base station to the first Within the user group Multiplexing user signals in the power domain is represented as follows: ; wherein, denotes the normalized data symbol of the user in the user group, , denotes the total power allocated to the user group, denotes the power allocation factor for the user in the user group, denotes the signal after power domain multiplexing for the user in the user group; In S2, the base station applies a corresponding beamforming vector to the signal of each user group, denoted as: ; ; wherein, represents a signal vector outputted by the base station after applying a beamforming vector to the th user group; represents a beamforming set, is a beamforming vector corresponding to a user in the th user group; represents the number of users contained in the th user group, i.e., the upper limit of the user index in the group; represents the upper limit of the user group index ; The base station simultaneously transmits superimposed signals for all user groups is expressed as: ; The user receives a signal, which includes a target signal, intra-group interference and inter-group interference; The user link corresponding to the user's received signal contains a direct link from the base station to the user and a reflection link from the base station to the controllable reflecting element and then to the user, denoted as: ; wherein, HBS→RISdenotes the channel matrix from the base station BS to the controllable reflecting element RIS, i.e. the reflecting link from the controllable reflecting element to the user; HBS→RISdenotes the channel matrix from the base station BS to the controllable reflecting element RIS, i.e. the reflecting link from the controllable reflecting element to the user; HBS→RISdenotes the channel matrix from the base station BS to the controllable reflecting element RIS, i.e. the reflecting link from the controllable reflecting element to the user; HBS→RISdenotes the channel matrix from the base station BS to the controllable reflecting element RIS, i.e. the reflecting link from the controllable reflecting element to the user; HBS→RISdenotes the channel matrix from the base station BS to the controllable reflecting element RIS, i.e. the reflecting link from the controllable reflecting element to the user; The reflection link at the user receiving end satisfies the SIC decoding order of non-orthogonal multiple access; Based on this, the first user in the first group receives a signal represented as: ; in, For the first Within the group Gaussian white noise received by each user; Represents the set of distractors; Indicates that the base station is for the first Within the group Signals transmitted by individual users; Based on the user location distribution, the user groups are grouped through two-dimensional joint optimization; The dynamic grouping objective function is as follows: ; wherein, denotes a maximization operation; denotes a set of group indices, i.e. optimization variables; denotes a channel matrix corresponding to the k-th user in the i-th group and denotes a channel matrix corresponding to the k-th user in the i-th group, respectively; denotes a channel matrix corresponding to the k-th user in the i-th group and denotes a channel matrix corresponding to the k-th user in the i-th group, respectively; denotes a channel matrix corresponding to the k-th user in the i-th group and denotes a conjugate transpose operation; Assume the user position is , and the base station position is ; wherein, are respectively an x-axis coordinate, a y-axis coordinate and a z-axis coordinate in three-dimensional space coordinates of the i-th user; are respectively an x-axis coordinate, a y-axis coordinate and a z-axis coordinate in three-dimensional space coordinates of the i-th user; are respectively an x-axis coordinate, a y-axis coordinate and a z-axis coordinate in three-dimensional space coordinates of the base station; Obtain the channel response including path loss and Rice fading, denoted as: ; ; wherein is a Rice factor; is a direct path loss exponent; denotes the three-dimensional spatial distance of the user to the base station; denotes a first reference path loss factor; denotes a non-line-of-sight channel component vector; denotes a line-of-sight channel component vector; The channel of the controllable reflecting element reflection link obeys the Rice distribution, denoted as: ; wherein denotes the conjugate transpose of the channel matrix observed by the user in the reflected link of the controllable reflecting element; denotes a second reference path loss factor; denotes the distance from the base station to the controllable reflecting element; denotes the distance from the controllable reflecting element to the user, denotes the conjugate transpose of the channel matrix from the base station to the controllable reflecting element.
3. The method of claim 2, wherein, S3 specifically: Acquiring a signal-to-interference ratio for a legitimate user is expressed as: ; wherein, denotes the transmit power of the th legitimate user; denotes the transmit power of the th legitimate user in the th group; denotes the channel gain of the th legitimate user in the th group; denotes the effective gain of the desired signal; denotes the transmit power of the th legitimate user; denotes the transmit power of the th legitimate user in the user group; Obtain the eavesdropper's received signal, denoted as: ; wherein, represents the received signal of the eavesdropper, represents the total number of eavesdroppers; represents the channel matrix of the eavesdropper; represents the total number of users in the group of users; represents the transmitted signal of the user in the group; represents the Gaussian white noise at the receiving end of the eavesdropper; Obtain the eavesdropper's equivalent channel, and minimize the equivalent gain of the eavesdropper's equivalent channel by optimizing the reflection coefficient of the controllable reflecting element; acquiring a signal-to-interference ratio corresponding to the eavesdropper is expressed as: ; Quantify the difference in ability between legitimate users and eavesdroppers through safe transmission rate, denoted as: ; wherein, denotes the user in the group; denotes the achievable rate of a legitimate user; denotes the achievable rate of an eavesdropper; denotes the take non-negative operation.
4. The method of claim 3, wherein a leakage control weight coefficient λ is introduced to guarantee the perception function assisted secure communication, and the optimization objective is denoted as: In S3, the perception gain constraint of the direction of leakage is introduced where, and denote the polar and azimuth angles of the potential target direction of leakage, respectively, denotes the leakage beam gain of the system in the direction of a certain potential eavesdropper, and is expressed as: ; wherein, ISAC sensing leakage threshold; 5. The method of claim 4, wherein S4 specifically: ; wherein, denotes the sum of all legitimate user precoding vectors; denotes the eavesdropping channel conjugate transpose in the corresponding direction; denotes the leakage gain threshold. Construct a joint optimization framework MAD2RA, and optimize the reflection coefficient matrix of the controllable reflecting element, the beamforming vector and the power allocation coefficient through multi-agent cooperation to maximize the total safe rate of the system; Divide the optimization task into optimization objective functions corresponding to the base station side and the RIS side.
6. The method of claim 5, wherein The optimization process at the base station side is denoted as: S401: Constructing an environment, the environment includes a base station, a user and a controllable reflecting element in communication with each other; S402: Based on the environment, a multi-agent architecture MADDPG is constructed, and the channel of each user in the environment is set as a user channel agent; the multi-agent architecture includes a policy network and a target network; S404: The policy network of the multi-agent architecture generates the power allocation coefficient of the user channel agent based on the global state; S403: The policy network of the multi-agent architecture obtains the global state of the user channel agent at the time slot , which includes the equivalent channel from the base station to the user and the intra-group interference and inter-group interference in the time slot ; S405: Based on the power allocation coefficient, the total safe transmission rate is calculated, and the total safe transmission rate is input as a reward into the target network of the multi-agent architecture. The target network of the multi-agent architecture calculates the timing difference error, and reversely updates the weight parameters of the policy network. Then, step S3 is re-executed.
7. The method of claim 6, wherein the method further comprises: In S4, the optimization process on the RIS side is represented as: S406: An environment shared with the base station side is constructed, and the environment includes a base station, a legitimate user, a controllable reflecting element, and a potential eavesdropper in communication with each other. The optimization goal on the RIS side is to adjust the phase shift matrix of the reflecting element to maximize the equivalent channel gain of the legitimate user, reduce the link gain of the eavesdropper, and construct a null in the potential leakage direction. S407: A single-agent architecture D3QN is constructed, and the controllable reflecting element is set as a central control agent, i.e., an RIS agent. The single-agent architecture includes an evaluation network and a target network. S408: The evaluation network of the single-agent architecture obtains the state of the RIS agent at time slot t, and the state includes a set of power allocation coefficients output by the base station side S4, a phase shift matrix of the controllable reflecting element, an equivalent channel from the base station to the legitimate user, and a SINR ratio of the eavesdropping link. S409: Based on the state of the RIS agent at time slot t obtained in S408, the evaluation network of the single-agent architecture generates a phase shift action of the controllable reflecting element. The phase shift action uniformly quantizes the phase shift of each reflecting element into Q levels, and outputs an N-dimensional phase level vector to correspond to the phase shift matrix of different reflecting elements. S410: Based on the phase shift matrix, the reflection coefficient and the beamforming vector of the controllable reflecting element are adjusted. Based on this, the total safe transmission rate is calculated and input as a reward into the target network of the single-agent architecture. Experience samples including the current state of the RIS agent, the phase shift action, the reward, and the next time slot state are recorded. The sampling priority of the experience samples in the priority experience replay pool is adjusted according to the change of the signal-to-interference ratio of the eavesdropping link. The target network calculates the target Q value in combination with the experience samples, compares the target Q value with the current Q value of the evaluation network to construct a loss function, minimizes the loss function to reversely update the weight parameters of the evaluation network, and then re-executes step S408.
8. The method of claim 7, wherein the optimization objective function on the base station side is specifically: given the reflection coefficient of the controllable reflecting element, the power allocation coefficient and the beamforming vector are optimized; and the optimization objective function is formalized as:
9. The method of claim 8, wherein the optimization objective function on the RIS side is specifically: given the power allocation coefficient and the beamforming vector, the reflection coefficient of the controllable reflecting element is optimized; and the optimization objective function is formalized as: ; in, Indicates the total secure transmission rate; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The transmission rate of each legitimate user; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The transmission rate of an eavesdropper; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The interference ratio of a legitimate user; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The interference ratio of a single eavesdropper's signal; This indicates the maximum distributed power. The scene construction module is configured to construct a downlink scene. The scene construction module is configured to construct a downlink scene. ; wherein, denotes the transmission rate of the kth legitimate user of the mth group given the power allocation factor and the beamforming vector; denotes the transmission rate of the kth eavesdropper of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth legitimate user of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth eavesdropper of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth legitimate user of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth eavesdropper of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth legitimate user of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth eavesdropper of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth legitimate user of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth eavesdropper of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth legitimate user of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth eavesdropper of the mth group given the power allocation factor and the beamforming vector; denotes the phase shift matrix of the controllable reflecting element.
10. A multi-agent reinforcement learning based security resource optimization system for Internet of Things, for implementing the multi-agent reinforcement learning based security resource optimization method according to any one of claims 1-9. The signal transmission module is configured to construct a signal transmission mechanism, and control a reconfigurable intelligent surface (RIS) to adjust a phase shift based on an intelligent controller, so that a base station generates a corresponding beamforming vector, assuming that channel state information (CSI) is available. The data acquisition module is configured to acquire a signal-to-interference-and-noise ratio (SINR) of users, including eavesdroppers and legitimate users, and acquire a secure transmission rate of each user based on a difference in the SINR between the legitimate users and the eavesdroppers. The joint optimization module is MAD2RA, which is configured to optimize a reflection coefficient of the RIS, a base station beamforming vector and a user power allocation coefficient through multi-agent cooperation, so as to maximize a total secure transmission rate of the system.
Citation Information
Patent Citations
DRL-based dual-RIS-position-assisted millimeter wave communication system optimal configuration method
CN117320043A
Industrial operating system resource instance scheduling optimization method
CN119149200A
RIS auxiliary vehicle edge calculation method based on improved DRL
CN119363779A
DRL-based active RIS-assisted MISO communication system joint optimization method
CN120151897A
Multi-RIS communication network rate improving method based on MADDPG
CN114727318A
Cited By
D2D communication resource allocation method
CN122227419A