A Method and System for Optimizing IoT Security Resources Based on Multi-Agent Reinforcement Learning

By optimizing the reflection coefficient of controllable reflective elements and base station beamforming through multi-agent reinforcement learning, the power-phase coupling optimization problem in IoT systems is solved, achieving efficient resource allocation and secure transmission in multi-user NOMA systems and improving system performance.

CN120916144BActive Publication Date: 2026-01-06EAST CHINA JIAOTONG UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511428325.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-01-06
Estimated Expiration
2045-09-30

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively address the coupling optimization relationship between power and phase in IoT systems, resulting in non-globally optimal system performance. Furthermore, the lack of a resource-intelligent decision-making architecture for multi-agent collaborative optimization makes it difficult to achieve synergistic enhancement of channel quality and interference management in a high-dimensional hybrid action space.

Method used

A method for optimizing IoT security resources based on multi-agent reinforcement learning is constructed. By cooperating among multiple agents, the reflection coefficient of controllable reflective elements, base station beamforming and power allocation coefficients are optimized to form a high-dimensional hybrid action space. A dual-agent architecture of MADDPG and D3QN is adopted to achieve programmable control of signal propagation paths and interference cancellation.

Benefits of technology

It achieves optimal power resource allocation in multi-user NOMA systems, enhances channel gain for legitimate users, suppresses eavesdropping link reflections, and improves the overall secure transmission rate and spectrum utilization efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120916144B_ABST
    Figure CN120916144B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-agent reinforcement learning's Internet of Things security resource optimization method and system, comprising: S1: constructing downlink MISO scene;S2: constructing signal transmission mechanism, assuming that channel state information is available state, controllable reflecting element is adjusted phase shift based on intelligent controller, base station BS generates corresponding beam;S3: obtaining the signal-to-interference-plus-noise ratio of user, user includes eavesdropper and legal user;Based on the signal-to-interference-plus-noise ratio difference of legal user and eavesdropper, the safety transmission rate is obtained;S4: constructing joint optimization framework, to maximize safety transmission rate.The application is deployed with multiple controllable reflecting units between base station and user RIS, realizes the programmable control to signal propagation path.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing technology, specifically a method and system for optimizing IoT security resources based on multi-agent reinforcement learning. Background Technology

[0002] With the exponential growth in the number of IoT terminals, systems need to support the simultaneous access of massive numbers of devices. Traditional communication architectures face multiple challenges, including scarce spectrum resources, severe interference, and insufficient transmission security.

[0003] On the one hand, regarding spectral efficiency, Orthogonal Multiple Access (OMA) technology, due to resource allocation isolation, struggles to meet the demands of high-density device connectivity. Non-Orthogonal Multiple Access (NOMA), through power domain multiplexing, allows multiple users to share the same frequency resource, offering significant advantages in improving system capacity. However, it also introduces severe intra- and inter-group co-channel interference, impacting system stability. On the other hand, regarding security, traditional networks rely on encryption protocols to ensure information security. However, the open nature of wireless channels makes it easy for passive eavesdroppers to intercept communications, especially in IoT environments with limited edge sensing and terminal computing capabilities. Traditional cryptographic mechanisms are ill-equipped to defend against physical layer attacks. Therefore, constructing a security enhancement mechanism centered on the physical layer has become an urgent need.

[0004] To further improve transmission performance, reconfigurable smart surfaces have been introduced as auxiliary nodes in recent years to intelligently regulate signal reflection paths, enhance link quality, and reduce interference. However, this also introduces complexity to resource scheduling. In RIS-NOMA systems, it is necessary to jointly optimize the continuous power allocation on the base station side and the discrete phase matrix on the RIS side, forming a high-dimensional hybrid action space (i.e., a joint optimization problem involving continuous and discrete variables). Traditional reinforcement learning algorithms (such as DDPG and PPO) struggle to converge effectively in high-dimensional discrete action spaces, exhibiting problems such as training instability and policy performance degradation.

[0005] Regarding existing technologies, although some patents have attempted to introduce deep reinforcement learning (DRL) mechanisms to improve the performance of RIS-NOMA systems, many shortcomings remain. For example, patent CN120151897A proposes a RIS phase control method based on single-agent DRL, which only optimizes phase modulation and does not consider the coordination relationship with base station power control, resulting in incomplete system resource allocation and limited performance improvement. Patent CN119149200A adopts a staged separation optimization strategy, modeling and solving power allocation and phase control separately. Although this reduces the complexity of model design and training, the lack of a global coordination mechanism means that the overall spectrum utilization efficiency and security performance of the system do not reach the optimal level.

[0006] Furthermore, while patent CN117320043A introduces a dual RIS structure in millimeter-wave communication scenarios and combines it with the DRL method to optimize user location information scheduling, it does not address system-level resource allocation strategies, particularly lacking modeling and implementation of a joint power and phase control mechanism. Although patent CN119363779A proposes a MADDPG-based offloading power optimization method, achieving intelligent allocation of power resources, it does not introduce a RIS phase shift optimization component, thus making it difficult to achieve coordinated enhancement of channel quality and interference management in the spatial dimension.

[0007] In summary, current technical solutions either neglect the coupling optimization relationship between power and phase, employ separate modeling leading to suboptimal system performance, or focus only on a single control dimension without establishing a complete perception-control closed-loop mechanism. Therefore, how to construct a resource-intelligent decision-making architecture that adapts to high-dimensional hybrid action spaces and supports multi-agent collaborative optimization for large-scale IoT scenarios remains a critical issue that urgently needs to be addressed in this technological field. Summary of the Invention

[0008] To address the shortcomings of existing technologies, this invention proposes an IoT security resource optimization method based on multi-agent reinforcement learning, comprising:

[0009] S1: Constructing the downlink scenario, specifically the downlink scenario as follows:

[0010] Base stations equipped with antennas and user groups distributed near the base stations Mutual communication; each user group All by Composed of several users, a network with communicative capabilities is deployed between the base station (BS) and the users. One controllable reflective element Obtain the power allocation coefficient for each user;

[0011] S2: Construct a signal transmission mechanism, assuming the channel state information is available, and a controllable reflection element. Based on the intelligent controller adjusting the reflection coefficient of the controllable reflective element, the base station generates the corresponding beamforming vector;

[0012] S3: Obtain the signal-to-interference ratio (SIR) of users, including both eavesdroppers and legitimate users; obtain the secure transmission rate for each user based on the difference in SIR between legitimate users and eavesdroppers.

[0013] S4: Construct the joint optimization framework MAD2RA. The joint optimization framework optimizes the reflection coefficient of controllable reflective elements, base station beamforming and power allocation coefficient through multi-agent collaboration to maximize the total secure transmission rate.

[0014] Furthermore, in S1:

[0015] The base station employs power superposition coding within each non-orthogonal multiple access group, combined with beamforming, specifically as follows:

[0016] Base station to the first Within the user group Multiplexing user signals in the power domain is represented as follows:

[0017] ;

[0018] in, Indicates the first Within the user group Normalized data symbols for each user, , Indicates assignment to the first Total power of each user group Indicates the first The user group Power allocation factor for each user Indicates the first Within the user group The signal after power domain multiplexing for each user;

[0019] In S2, the base station applies a corresponding beamforming vector to the signal of each user group, as follows:

[0020] ;

[0021] ;

[0022] in, Indicates that the base station is targeting the first The signal vector output after applying beamforming vectors to each user group; Represents a beamforming set. It is the first Beamforming vectors corresponding to users in each user group; Indicates the first The number of users contained in a user group, i.e., the user index within the group. The upper limit; Indicates user group index The upper limit;

[0023] The base station simultaneously transmits superimposed signals to all user groups. , represented as:

[0024] ;

[0025] The user receives signals, which include target signals, intra-group interference, and inter-group interference.

[0026] The user's received signal corresponds to a user link that includes a direct link from the base station to the user and a reflected link through the base station to the controllable reflection unit and then to the user, represented as:

[0027] ;

[0028] in, This represents the channel matrix from the base station (BS) to the controllable reflective element (RIS), which is also the reflection link from the controllable reflective element to the user. This indicates the reflection link from the controllable reflective element to the user. It is a direct link from the base station to the user; Indicates the reflection link at the user's receiving end; This represents the reflection coefficient of a controllable reflective element;

[0029] The reflection link at the user receiver satisfies the SIC decoding order of non-orthogonal multiple access;

[0030] Based on this, the first Within the group Signal received by a user Represented as:

[0031] ;

[0032] in, For the first Within the group Gaussian white noise received by each user; Represents the set of distractors; Indicates that the base station is for the first Within the group Signals transmitted by individual users;

[0033] Based on user location distribution, user groups are grouped through a two-dimensional joint optimization.

[0034] The objective function for dynamic grouping is as follows:

[0035] ;

[0036] in, This indicates a maximize operation; This represents the set of grouped indexes, which is also the optimization variable; They represent the first The first in the group The channel matrix corresponding to the user and the first user The first in the group Channel matrix corresponding to each user; This represents the conjugate transpose operation;

[0037] Assuming the user's location is The location of the base station is ;

[0038] in, The first The x-axis coordinates, y-axis coordinates, and z-axis coordinates of each user in three-dimensional space; These are the x-axis, y-axis, and z-axis coordinates in the three-dimensional spatial coordinates of the base station, respectively.

[0039] The channel response, including path loss and Ricean fading, is obtained as follows:

[0040] ;

[0041] ;

[0042] in, Rice factor; The direct path loss index; Indicates the first The three-dimensional spatial distance between each user and the base station; Indicates the first reference path loss factor; Represents the non-line-of-sight channel component vector; Represents the line-of-sight channel component vector;

[0043] The channel of the reflective link of the controllable reflective element follows a Ricean distribution, expressed as:

[0044] ;

[0045] in, This indicates the first step in the reflection link of the controllable reflective element. The conjugate transpose of the channel matrix observed by each user; Indicates the loss factor of the second reference path; This indicates the distance from the base station to the controllable reflective element; Indicates the controllable reflective element up to the first Distance between users This represents the conjugate transpose of the channel matrix from the base station to the controllable reflective element.

[0046] Furthermore, S3 specifically refers to:

[0047] Obtain the signal interference ratio of legitimate users , represented as:

[0048] ;

[0049] in, Indicates the first The transmit power of each legitimate user; Indicates the first The first in the group Channel gain for each legitimate user; Indicates the first The first in the group The conjugate transpose of a valid user channel matrix; Indicates the effective gain of the desired signal; Indicates the first The transmit power of each legitimate user; Indicates the first Beamforming vectors corresponding to legitimate users in a user group;

[0050] The received signal from the eavesdropper is represented as:

[0051] ;

[0052] in, This indicates the signal received by the eavesdropper. Indicates the total number of eavesdroppers; The channel matrix represents the eavesdropper; Indicates the first The total number of users within a user group; Indicates the first Within the group The transmitted signal of each user; This represents the Gaussian white noise at the receiver's end;

[0053] Obtain the eavesdropper's equivalent channel and minimize the equivalent gain of the eavesdropper's equivalent channel by optimizing the reflection coefficient of the controllable reflection element;

[0054] Obtain the signal-to-interference ratio corresponding to the eavesdropper. , represented as:

[0055] ;

[0056] The difference in capabilities between legitimate users and eavesdroppers is quantified by secure transmission rate, and expressed as follows:

[0057] ;

[0058] in, Indicates the first The first in the group Secure transmission rate for each user; Indicates the reachable rate for legitimate users; Indicates the reachability of the eavesdropper; This indicates the non-negation operation.

[0059] Furthermore, in S3, a sensing gain constraint for the leakage direction is introduced. ,in, and These represent the polar angle and azimuth angle, respectively, indicating the direction of the leaked potential target. The leakage beam gain of the system in the direction of a potential eavesdropper is expressed as:

[0060] ;

[0061] in, For ISAC, a leak detection threshold is required.

[0062] Introducing a leakage control weight coefficient λ to ensure secure communication assisted by the sensing function, the optimization objective is expressed as:

[0063] ;

[0064] in, This represents the sum of all valid user precoded vectors; This represents the conjugate transpose of the eavesdropping channel in the corresponding direction; This indicates the leakage gain threshold.

[0065] Furthermore, S4 specifically refers to:

[0066] A joint optimization framework, MAD2RA, is constructed to optimize the reflection coefficient matrix, beamforming vector, and power allocation coefficient of controllable reflective elements through multi-agent collaboration to maximize the overall system safety rate. The optimization task is divided into optimization objective functions corresponding to the base station side and the RIS side.

[0067] Furthermore, the optimization process on the base station side is represented as follows:

[0068] S401: Constructing an environment that includes mutually communicating base stations, users, and controllable reflective elements;

[0069] S402: The environment-based multi-agent architecture MADDPG is constructed, which sets the channel of each user in the environment as a user channel agent; the multi-agent architecture includes a policy network and a target network.

[0070] S403: Multi-agent architecture policy network acquires user channel agents in time slots The global state, which includes the equivalent channel from the base station to the user and the time slots. Intra-group interference and inter-group interference;

[0071] S404: The policy network of the multi-agent architecture generates power allocation coefficients for user channel agents based on the global state;

[0072] S405: Based on the power allocation coefficient, calculate the total safe transmission rate, use the total safe transmission rate as a reward input to the target network of the multi-agent architecture, calculate the temporal differential error of the target network of the multi-agent architecture, update the weight parameters of the policy network in reverse, and re-execute step S3.

[0073] Furthermore, in S4, the optimization process on the RIS side is expressed as follows:

[0074] S406: Construct an environment shared with the base station side, which includes mutually communicating base stations, legitimate users, controllable reflective elements, and potential eavesdroppers; the optimization objective of the RIS side is to adjust the phase shift matrix of the reflective elements to maximize the equivalent channel gain of legitimate users, reduce the gain of eavesdropping links, and construct null traps in potential leakage directions;

[0075] S407: Construct a single-agent architecture D3QN, setting the controllable reflective element as the central control agent, i.e., the RIS agent; the single-agent architecture includes an evaluation network and a target network;

[0076] S408: The evaluation network of the single agent architecture obtains the state of the RIS agent in time slot t. The state includes the set of power allocation coefficients output by the S4 on the base station side, the phase shift matrix of the currently controllable reflective element, the equivalent channel from the base station to the legitimate user, and the SINR ratio of the eavesdropping link.

[0077] S409: The evaluation network of the single agent architecture is based on the state of the RIS agent in time slot t obtained by S408. It generates the phase shift action of the controllable reflection element. The phase shift action uniformly quantizes the phase shift of each reflection unit into Q levels and outputs an N-dimensional phase level vector to correspond to the phase shift matrix of different reflection units.

[0078] S410: Based on the phase shift matrix, adjust the reflection coefficient and beamforming vector of the controllable reflection element; calculate the total secure transmission rate based on this and use it as a reward input to the target network of the single-agent architecture; simultaneously record experience samples including the current state, phase shift action, reward, and next time slot state of the RIS agent, and adjust the sampling priority of the experience samples in the priority experience replay pool according to the change of the signal-to-interference ratio of the eavesdropping link; the target network calculates the target Q value by combining the experience samples, compares it with the current Q value of the evaluation network to construct a loss function, minimizes the loss function to update the weight parameters of the evaluation network in reverse; then repeat step S408.

[0079] Furthermore, the optimization objective function on the base station side is specifically: given the reflection coefficient of the controllable reflective element, optimize the power allocation coefficient and the beamforming vector; formally, it is:

[0080] ;

[0081] in, Indicates the total secure transmission rate; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The transmission rate of each legitimate user; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The transmission rate of an eavesdropper; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The interference ratio of a legitimate user; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The interference ratio of a single eavesdropper's signal; This indicates the maximum allocated power.

[0082] Furthermore, the optimization objective function on the RIS side is specifically: given the power allocation coefficient and beamforming vector, optimize the reflection coefficient of the controllable reflective element, formalized as:

[0083] ;

[0084] in, This indicates that, given the power allocation coefficients and beamforming vector, the first... Group 1 The transmission rate of each legitimate user; This indicates that, given the power allocation coefficients and beamforming vector, the first... Group 1 The transmission rate of an eavesdropper; This indicates that, given the power allocation coefficients and beamforming vector, the first... Group 1 The interference ratio of a legitimate user; This indicates that, given the power allocation coefficients and beamforming vector, the first... Group 1 The interference ratio of a single eavesdropper's signal; This represents the phase shift matrix of a controllable reflective element.

[0085] A multi-agent reinforcement learning-based IoT security resource optimization system, used to implement the aforementioned multi-agent reinforcement learning-based IoT security resource optimization method, is characterized by comprising:

[0086] The scenario building module is used to build downlink scenarios;

[0087] The signal transmission module is used to construct the signal transmission mechanism. Assuming that the channel state information (CSI) is available, it controls the controllable reflective element (RIS) to adjust the phase shift based on the intelligent controller, so that the base station generates the corresponding beamforming vector.

[0088] The data acquisition module is used to acquire the signal-to-interference-plus-noise ratio (SIR) of users, including eavesdroppers and legitimate users; and to acquire the secure transmission rate of each user based on the difference in SIR between legitimate users and eavesdroppers.

[0089] The joint optimization module, namely MAD2RA, is used to optimize the reflection coefficient of controllable reflective elements, base station beamforming vector and user power allocation coefficient through multi-agent cooperation, so as to maximize the total secure transmission rate of the system.

[0090] Compared with the prior art, the beneficial effects of the present invention are:

[0091] 1. This invention deploys a RIS (Reflection System) with multiple controllable reflection units between the base station and the user, enabling programmable control of the signal propagation path. The system supports simultaneous access by multiple NOMA user groups, achieving non-orthogonal communication through power superposition and interference cancellation. However, the possibility of eavesdroppers passively monitoring legitimate communication links necessitates physical layer protection.

[0092] 2. This invention constructs a collaborative optimization framework based on MADDPG and D3QN dual agents. An independent agent is constructed for each user channel. Continuous power allocation actions are generated through a deep deterministic policy gradient algorithm to achieve optimal allocation of power resources among multiple users. The D3QN agent is located at the RIS end and acts as a centralized agent. It uses a duel-based dual deep Q network method to learn the optimal phase shift matrix control strategy in a high-dimensional discrete action space, effectively enhancing the gain of legitimate user channels and suppressing eavesdropping link reflections.

[0093] 3. This invention jointly considers base station power coefficient, RIS phase shift configuration, channel CSI, and eavesdropping link characteristics to construct a hybrid state space; power action is a continuous variable, while phase shift is a multi-level discrete variable, forming a high-dimensional hybrid action space. For this space, this invention achieves effective modeling and rapid convergence through a dual-network architecture and an experience replay mechanism. Attached Figure Description

[0094] Figure 1 This is a system framework diagram of an IoT security resource optimization system based on multi-agent reinforcement learning. Detailed Implementation

[0095] Reference Figure 1 A method for optimizing IoT security resources based on multi-agent reinforcement learning, comprising:

[0096] S1: Constructing the downlink scenario, specifically the downlink scenario as follows:

[0097] Base stations equipped with antennas and user groups distributed near the base stations Mutual communication; each user group All by Composed of several users, a network with communicative capabilities is deployed between the base station (BS) and the users. One controllable reflective element Obtain the power allocation coefficient for each user;

[0098] S2: Construct a signal transmission mechanism, assuming the channel state information is available, and a controllable reflection element. Based on the intelligent controller adjusting the reflection coefficient of the controllable reflective element, the base station generates the corresponding beamforming vector;

[0099] S3: Obtain the signal-to-interference ratio (SIR) of users, including both eavesdroppers and legitimate users; obtain the secure transmission rate for each user based on the difference in SIR between legitimate users and eavesdroppers.

[0100] S4: Construct the joint optimization framework MAD2RA. The joint optimization framework optimizes the reflection coefficient of controllable reflective elements, base station beamforming and power allocation coefficient through multi-agent collaboration to maximize the total secure transmission rate.

[0101] Furthermore, in S1:

[0102] The base station employs power superposition coding within each non-orthogonal multiple access group, combined with beamforming, specifically as follows:

[0103] Base station to the first Within the user group Multiplexing user signals in the power domain is represented as follows:

[0104] ;

[0105] in, Indicates the first Within the user group Normalized data symbols for each user, , Indicates assignment to the first Total power of each user group Indicates the first The user group Power allocation factor for each user Indicates the first Within the user group The signal after power domain multiplexing for each user;

[0106] In S2, the base station applies a corresponding beamforming vector to the signal of each user group, as follows:

[0107] ;

[0108] ;

[0109] in, Indicates that the base station is targeting the first The signal vector output after applying beamforming vectors to each user group; Represents a beamforming set. It is the first Beamforming vectors corresponding to users in each user group; Indicates the first The number of users contained in a user group, i.e., the user index within the group. The upper limit; Indicates user group index The upper limit;

[0110] The base station simultaneously transmits superimposed signals to all user groups. , represented as:

[0111] ;

[0112] The user receives signals, which include target signals, intra-group interference, and inter-group interference.

[0113] The user's received signal corresponds to a user link that includes a direct link from the base station to the user and a reflected link through the base station to the controllable reflection unit and then to the user, represented as:

[0114] ;

[0115] in, This represents the channel matrix from the base station (BS) to the controllable reflective element (RIS), which is also the reflection link from the controllable reflective element to the user. This indicates the reflection link from the controllable reflective element to the user. It is a direct link from the base station to the user; Indicates the reflection link at the user's receiving end; This represents the reflection coefficient of a controllable reflective element;

[0116] The reflection link at the user receiver satisfies the SIC decoding order of non-orthogonal multiple access;

[0117] Based on this, the first Within the group Signal received by a user Represented as:

[0118] ;

[0119] in, For the first Within the group Gaussian white noise received by each user; Represents the set of distractors; Indicates that the base station is for the first Within the group Signals transmitted by individual users;

[0120] Based on user location distribution, user groups are grouped through a two-dimensional joint optimization.

[0121] The objective function for dynamic grouping is as follows:

[0122] ;

[0123] in, This indicates a maximize operation; This represents the set of grouped indexes, which is also the optimization variable; They represent the first The first in the group The channel matrix corresponding to the user and the first user The first in the group Channel matrix corresponding to each user; This represents the conjugate transpose operation;

[0124] Assuming the user's location is The location of the base station is ;

[0125] in, The first The x-axis coordinates, y-axis coordinates, and z-axis coordinates of each user in three-dimensional space; These are the x-axis, y-axis, and z-axis coordinates in the three-dimensional spatial coordinates of the base station, respectively.

[0126] The channel response, including path loss and Ricean fading, is obtained as follows:

[0127] ;

[0128] ;

[0129] in, Rice factor; The direct path loss index; Indicates the first The three-dimensional spatial distance between each user and the base station; Indicates the first reference path loss factor; Represents the non-line-of-sight channel component vector; Represents the line-of-sight channel component vector;

[0130] The channel of the reflective link of the controllable reflective element follows a Ricean distribution, expressed as:

[0131] ;

[0132] in, This indicates the first step in the reflection link of the controllable reflective element. The conjugate transpose of the channel matrix observed by each user; Indicates the loss factor of the second reference path; This indicates the distance from the base station to the controllable reflective element; Indicates the controllable reflective element up to the first Distance between users This represents the conjugate transpose of the channel matrix from the base station to the controllable reflective element.

[0133] Furthermore, S3 specifically refers to:

[0134] Obtain the signal interference ratio of legitimate users , represented as:

[0135] ;

[0136] in, Indicates the first The transmit power of each legitimate user; Indicates the first The first in the group Channel gain for each legitimate user; Indicates the first The first in the group The conjugate transpose of a valid user channel matrix; Indicates the effective gain of the desired signal; Indicates the first The transmit power of each legitimate user; Indicates the first Beamforming vectors corresponding to legitimate users in a user group;

[0137] The received signal from the eavesdropper is represented as:

[0138] ;

[0139] in, This indicates the signal received by the eavesdropper. Indicates the total number of eavesdroppers; The channel matrix represents the eavesdropper; Indicates the first The total number of users within a user group; Indicates the first Within the group The transmitted signal of each user; This represents the Gaussian white noise at the receiver's end;

[0140] Obtain the eavesdropper's equivalent channel and minimize the equivalent gain of the eavesdropper's equivalent channel by optimizing the reflection coefficient of the controllable reflection element;

[0141] Obtain the signal-to-interference ratio corresponding to the eavesdropper. , represented as:

[0142] ;

[0143] The difference in capabilities between legitimate users and eavesdroppers is quantified by secure transmission rate, and expressed as follows:

[0144] ;

[0145] in, Indicates the first The first in the group Secure transmission rate for each user; Indicates the reachable rate for legitimate users; Indicates the reachability of the eavesdropper; This indicates the non-negation operation.

[0146] Furthermore, in S3, a sensing gain constraint for the leakage direction is introduced. ,in, and These represent the polar angle and azimuth angle, respectively, indicating the direction of the leaked potential target. The leakage beam gain of the system in the direction of a potential eavesdropper is expressed as:

[0147] ;

[0148] in, For ISAC, a leak detection threshold is required.

[0149] Introducing a leakage control weight coefficient λ to ensure secure communication assisted by the sensing function, the optimization objective is expressed as:

[0150] ;

[0151] in, This represents the sum of all valid user precoded vectors; This represents the conjugate transpose of the eavesdropping channel in the corresponding direction; This indicates the leakage gain threshold.

[0152] Furthermore, S4 specifically refers to:

[0153] A joint optimization framework, MAD2RA, is constructed to optimize the reflection coefficient matrix, beamforming vector, and power allocation coefficient of controllable reflective elements through multi-agent collaboration to maximize the overall system safety rate. The optimization task is divided into optimization objective functions corresponding to the base station side and the RIS side.

[0154] Furthermore, the optimization process on the base station side is represented as follows:

[0155] S401: Constructing an environment that includes mutually communicating base stations, users, and controllable reflective elements;

[0156] S402: The environment-based multi-agent architecture MADDPG is constructed, which sets the channel of each user in the environment as a user channel agent; the multi-agent architecture includes a policy network and a target network.

[0157] S403: Multi-agent architecture policy network acquires user channel agents in time slots The global state, which includes the equivalent channel from the base station to the user and the time slots. Intra-group interference and inter-group interference;

[0158] S404: The policy network of the multi-agent architecture generates power allocation coefficients for user channel agents based on the global state;

[0159] S405: Based on the power allocation coefficient, calculate the total safe transmission rate, use the total safe transmission rate as a reward input to the target network of the multi-agent architecture, calculate the temporal differential error of the target network of the multi-agent architecture, update the weight parameters of the policy network in reverse, and re-execute step S3.

[0160] Furthermore, in S4, the optimization process on the RIS side is expressed as follows:

[0161] S406: Construct an environment shared with the base station side, which includes mutually communicating base stations, legitimate users, controllable reflective elements, and potential eavesdroppers; the optimization objective of the RIS side is to adjust the phase shift matrix of the reflective elements to maximize the equivalent channel gain of legitimate users, reduce the gain of eavesdropping links, and construct null traps in potential leakage directions;

[0162] S407: Construct a single-agent architecture D3QN, setting the controllable reflective element as the central control agent, i.e., the RIS agent; the single-agent architecture includes an evaluation network and a target network;

[0163] S408: The evaluation network of the single agent architecture obtains the state of the RIS agent in time slot t. The state includes the set of power allocation coefficients output by the S4 on the base station side, the phase shift matrix of the currently controllable reflective element, the equivalent channel from the base station to the legitimate user, and the SINR ratio of the eavesdropping link.

[0164] S409: The evaluation network of the single agent architecture is based on the state of the RIS agent in time slot t obtained by S408. It generates the phase shift action of the controllable reflection element. The phase shift action uniformly quantizes the phase shift of each reflection unit into Q levels and outputs an N-dimensional phase level vector to correspond to the phase shift matrix of different reflection units.

[0165] S410: Based on the phase shift matrix, adjust the reflection coefficient and beamforming vector of the controllable reflection element; calculate the total secure transmission rate based on this and use it as a reward input to the target network of the single-agent architecture; simultaneously record experience samples including the current state, phase shift action, reward, and next time slot state of the RIS agent, and adjust the sampling priority of the experience samples in the priority experience replay pool according to the change of the signal-to-interference ratio of the eavesdropping link; the target network calculates the target Q value by combining the experience samples, compares it with the current Q value of the evaluation network to construct a loss function, minimizes the loss function to update the weight parameters of the evaluation network in reverse; then repeat step S408.

[0166] Furthermore, the optimization objective function on the base station side is specifically: given the reflection coefficient of the controllable reflective element, optimize the power allocation coefficient and the beamforming vector; formally, it is:

[0167] ;

[0168] in, Indicates the total secure transmission rate; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The transmission rate of each legitimate user; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The transmission rate of an eavesdropper; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The interference ratio of a legitimate user; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The interference ratio of a single eavesdropper's signal; This indicates the maximum allocated power.

[0169] Furthermore, the optimization objective function on the RIS side is specifically: given the power allocation coefficient and beamforming vector, optimize the reflection coefficient of the controllable reflective element, formalized as:

[0170] ;

[0171] in, This indicates that, given the power allocation coefficients and beamforming vector, the first... Group 1 The transmission rate of each legitimate user; This indicates that, given the power allocation coefficients and beamforming vector, the first... Group 1 The transmission rate of an eavesdropper; This indicates that, given the power allocation coefficients and beamforming vector, the first... Group 1 The interference ratio of a legitimate user; This indicates that, given the power allocation coefficients and beamforming vector, the first... Group 1 The interference ratio of a single eavesdropper's signal; This represents the phase shift matrix of a controllable reflective element.

[0172] A multi-agent reinforcement learning-based IoT security resource optimization system, used to implement the aforementioned multi-agent reinforcement learning-based IoT security resource optimization method, is characterized by comprising:

[0173] The scenario building module is used to build downlink scenarios;

[0174] The signal transmission module is used to construct the signal transmission mechanism. Assuming that the channel state information (CSI) is available, it controls the controllable reflective element (RIS) to adjust the phase shift based on the intelligent controller, so that the base station generates the corresponding beamforming vector.

[0175] The data acquisition module is used to acquire the signal-to-interference-plus-noise ratio (SIR) of users, including eavesdroppers and legitimate users; and to acquire the secure transmission rate of each user based on the difference in SIR between legitimate users and eavesdroppers.

[0176] The joint optimization module, namely MAD2RA, is used to optimize the reflection coefficient of controllable reflective elements, base station beamforming vector and user power allocation coefficient through multi-agent cooperation, so as to maximize the total secure transmission rate of the system.

[0177] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for optimizing security resources of an Internet of Things based on multi-agent reinforcement learning, characterized in that, Comprise: S1: Constructing a downlink scene, the downlink scene is specifically: A base station equipped with an antenna and a group of users distributed around the base station communicate with each other; each group of users consists of one user, and a controllable reflecting element is communicatively disposed between the base station BS and the user ; a power allocation coefficient of each user is obtained; S2: Construct a signal transmission mechanism, assuming that the channel state information is available, and the controllable reflecting element Adjust the reflection coefficient of the controllable reflecting element based on the intelligent controller, and the base station generates the corresponding beamforming vector; S3: Obtaining the signal-to-interference ratio of users, including eavesdroppers and legitimate users; Based on the difference in signal-to-interference ratio between legitimate users and eavesdroppers, the safe transmission rate of each user is obtained; S4: Constructing a joint optimization framework MAD2RA, the joint optimization framework optimizes the reflection coefficient of the controllable reflecting element, the beamforming of the base station and the power allocation coefficient through multi-agent cooperation to maximize the total safe transmission rate; In S4, the optimization process on the RIS side is represented as: S406: Constructing an environment shared with the base station side, the environment contains base stations, legitimate users, controllable reflecting elements and potential eavesdroppers in communication; The optimization goal of RIS side is to adjust the phase shift matrix of reflecting unit to maximize the equivalent channel gain of legitimate users, reduce the gain of eavesdropping link and construct null in potential leakage direction; S407: Constructing a single-agent architecture D3QN, setting the controllable reflecting element as the central control agent, namely the RIS agent; The single-agent architecture includes an evaluation network and a target network; S408: The evaluation network of the single-agent architecture obtains the state of the RIS agent at time slot t, including the set of power allocation coefficients output by S4 on the base station side, the phase shift matrix of the current controllable reflecting element, the equivalent channel from the base station to the legitimate user and the SINR proportion of the eavesdropping link; S409: Based on the state of the RIS agent at time slot t obtained in S408, the evaluation network of the single-agent architecture generates a phase shift action of the controllable reflecting element, which uniformly quantizes the phase shift of each reflecting unit into Q levels, outputs an N-dimensional phase level vector to correspond to the phase shift matrix of different reflecting units; S410: Based on the phase shift matrix, adjust the reflection coefficient and beamforming vector of the controllable reflecting element; Based on this, the total safe transmission rate is calculated and input as a reward to the target network of the single-agent architecture; At the same time, record the experience sample including the current state of the RIS agent, the phase shift action, the reward and the next time slot state, adjust the sampling priority of the experience sample in the priority experience replay pool according to the change of the eavesdropping link signal-to-interference ratio; The target network combines the experience sample to calculate the target Q value, compares it with the current Q value of the evaluation network to build a loss function, minimizes the loss function to update the evaluation network weight parameters in reverse; Then re-execute step S408.

2. The method of claim 1, wherein, In S1: The base station uses power superposition coding in each non-orthogonal multiple access group, combined with beamforming, specifically: Base station to the first Within the user group Multiplexing user signals in the power domain is represented as follows: ; wherein, denotes the normalized data symbol of the th user in the th user group, , denotes the total power allocated to the th user group, denotes the power allocation factor for the th user in the th user group, denotes the signal after power domain multiplexing for the th user in the th user group; In S2, the base station applies a corresponding beamforming vector to the signal of each user group, represented as: ; ; wherein, represents a signal vector outputted by the base station after applying a beamforming vector to the th user group; represents a beamforming set, is a beamforming vector corresponding to a user in the th user group; represents the number of users contained in the th user group, i.e., the upper limit of the user index in the group; represents the upper limit of the user group index ; The base station simultaneously transmits superimposed signals for all user groups is expressed as: ; The user receives the signal, which includes the target signal, intra-group interference and inter-group interference; The user link corresponding to the user's received signal contains the direct link from the base station to the user and the reflected link from the base station to the controllable reflecting element and then to the user, represented as: ; wherein, HBS→RISdenotes the channel matrix from the base station BS to the controllable reflecting element RIS, i.e. the reflecting link from the controllable reflecting element to the user; HBS→RISdenotes the channel matrix from the base station BS to the controllable reflecting element RIS, i.e. the reflecting link from the controllable reflecting element to the user; HBS→RISdenotes the channel matrix from the base station BS to the controllable reflecting element RIS, i.e. the reflecting link from the controllable reflecting element to the user; HBS→RISdenotes the channel matrix from the base station BS to the controllable reflecting element RIS, i.e. the reflecting link from the controllable reflecting element to the user; HBS→RISdenotes the channel matrix from the base station BS to the controllable reflecting element RIS, i.e. the reflecting link from the controllable reflecting element to the user; The reflected link at the user receiving end satisfies the SIC decoding order of non-orthogonal multiple access; Based on this, the first user in the first group receives a signal representing: representing:​ ; wherein is the th user in the th group receives a Gaussian white noise; denotes a set of interference terms; denotes a signal transmitted by the base station to the th user in the th group; Based on the user position distribution, the user groups are grouped through two-dimensional joint optimization; The dynamic grouping objective function is as follows: ; wherein, denotes a maximization operation; denotes a set of group indices, i.e., optimization variables; denotes a channel matrix corresponding to the k-th user in the i-th group and denotes a channel matrix corresponding to the k-th user in the i-th group, respectively; denotes a channel matrix corresponding to the k-th user in the i-th group and denotes a channel matrix corresponding to the k-th user in the i-th group, respectively; denotes a channel matrix corresponding to the k-th user in the i-th group and denotes a conjugate transpose operation; Assume the user position is , and the base station position is ; wherein, are respectively the x-axis coordinate, the y-axis coordinate and the z-axis coordinate in the three-dimensional space coordinates of the are respectively the x-axis coordinate, the y-axis coordinate and the z-axis coordinate in the three-dimensional space coordinates of the base station;​ Obtain the channel response containing path loss and Rayleigh fading, represented as: ; ; wherein, is a Rice factor; is a direct path loss exponent; denotes the three-dimensional spatial distance of the user to the base station; denotes a first reference path loss factor; denotes a non-line-of-sight channel component vector; denotes a line-of-sight channel component vector; The channel of the controllable reflecting element reflected link obeys Rayleigh distribution, represented as: ; wherein denotes the conjugate transpose of the channel matrix observed by the user in the reflected link of the controllable reflecting element; denotes a second reference path loss factor; denotes the distance from the base station to the controllable reflecting element; denotes the distance from the controllable reflecting element to the user, denotes the conjugate transpose of the channel matrix from the base station to the controllable reflecting element.

3. The method of claim 2, wherein, S3 is specifically: Acquiring a signal-to-interference ratio for a legitimate user is expressed as: ; wherein, denotes the transmit power of the th legitimate user; denotes the transmit power of the th legitimate user in the th group; denotes the channel gain of the th legitimate user in the th group; denotes the effective gain of the desired signal; denotes the transmit power of the th legitimate user; denotes the transmit power of the th legitimate user in the user group; Obtain the receiving signal of the eavesdropper, denoted as: ; wherein, represents the received signal of the eavesdropper, represents the total number of eavesdroppers; represents the channel matrix of the eavesdropper; represents the total number of users within the group of users; represents the transmitted signal of the user within the group; represents the Gaussian white noise at the receiving end of the eavesdropper; Obtain the equivalent channel of the eavesdropper, and minimize the equivalent gain of the equivalent channel of the eavesdropper by optimizing the reflection coefficient of the controllable reflecting element; acquiring a signal-to-interference ratio corresponding to the eavesdropper is expressed as: ; Quantify the difference in capability between the legitimate user and the eavesdropper through the secure transmission rate, denoted as: ; wherein, denotes the user in the group; denotes the achievable rate of a legitimate user; denotes the achievable rate of an eavesdropper; denotes the taking of the non-negative operation.

4. The method of claim 3, wherein a leakage control weight coefficient λ is introduced to guarantee the perception function assisted secure communication, and an optimization objective is denoted as: In S3, the perception gain constraint of the direction of leakage is introduced where, and denote the polar and azimuth angles of the potential target direction of leakage, respectively, denotes the leakage beam gain of the system in the direction of a certain potential eavesdropper, and is expressed as: ; wherein, ISAC sensing leakage threshold; 5. The method of claim 4, wherein S4 is specifically: ; wherein, denotes the sum of all legitimate user precoding vectors; denotes the eavesdropping channel conjugate transpose in the corresponding direction; denotes the leakage gain threshold. A joint optimization framework MAD2RA is constructed, and the reflection coefficient matrix of the controllable reflecting element, the beamforming vector and the power allocation coefficient are optimized through multi-agent cooperation to maximize the total secure rate of the system. The optimization task is divided into optimization objective functions corresponding to the base station side and the RIS side.

6. The method of claim 5, wherein the optimization process on the base station side is denoted as: S401: An environment is constructed, which includes a base station, a user and a controllable reflecting element in communication with each other; S402: A multi-agent architecture MADDPG is constructed based on the environment, and the channel of each user in the environment is set as a user channel agent; the multi-agent architecture includes a policy network and a target network; S404: The policy network of the multi-agent architecture generates the power allocation coefficient of the user channel agent based on the global state; S405: The total secure transmission rate is calculated based on the power allocation coefficient, and the total secure transmission rate is input as a reward to the target network of the multi-agent architecture; the target network of the multi-agent architecture calculates the timing difference error, and reversely updates the weight parameters of the policy network; and the step S3 is executed again.

7. The method of claim 6, wherein the optimization objective function on the base station side is specifically: the power allocation coefficient and the beamforming vector are optimized given the reflection coefficient of the controllable reflecting element; and is formalized as: S403: The policy network of the multi-agent architecture obtains the global state of the user channel agent at the time slot , which includes the equivalent channel from the base station to the user and the intra-group and inter-group interference in the time slot ; 8. The method of claim 7, wherein the optimization objective function on the RIS side is specifically: the reflection coefficient of the controllable reflecting element is optimized given the power allocation coefficient and the beamforming vector; and is formalized as: It includes: A scene construction module for constructing a downlink scene; A signal transmission module for constructing a signal transmission mechanism, assuming that channel state information (CSI) is available, adjusting the phase shift of the controllable reflecting element (RIS) based on an intelligent controller, and making the base station generate a corresponding beamforming vector; ; in, Indicates the total secure transmission rate; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The transmission rate of each legitimate user; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The transmission rate of an eavesdropper; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The interference ratio of a legitimate user; This indicates that, given the reflection coefficient of a controllable reflective element, the first... Group 1 The interference ratio of a single eavesdropper's signal; This indicates the maximum allocated power. A data acquisition module for acquiring the signal-to-interference-and-noise ratio (SINR) of users, including eavesdroppers and legitimate users; and acquiring the secure transmission rate of each user based on the difference in SINR between the legitimate user and the eavesdropper; ​ ; wherein, denotes the transmission rate of the kth legitimate user of the mth group given the power allocation factor and the beamforming vector; denotes the transmission rate of the kth eavesdropper of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth legitimate user of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth eavesdropper of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth legitimate user of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth eavesdropper of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth legitimate user of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth eavesdropper of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth legitimate user of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth eavesdropper of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth legitimate user of the mth group given the power allocation factor and the beamforming vector; denotes the signal-to-interference ratio of the kth eavesdropper of the mth group given the power allocation factor and the beamforming vector; denotes the phase shift matrix of the controllable reflecting element.

9. A multi-agent reinforcement learning based security resource optimization system for Internet of Things, for implementing the multi-agent reinforcement learning based security resource optimization method according to any one of claims 1-8. ​ ​ ​ ​ A joint optimization module, denoted as MAD2RA, is used to maximize the total safe transmission rate of the system by jointly optimizing the reflection coefficients of the controllable reflecting elements, the base station beamforming vectors and the user power allocation coefficients through multi-agent collaboration optimization.

Citation Information

Patent Citations

  • DRL-based dual-RIS-position-assisted millimeter wave communication system optimal configuration method

    CN117320043A

  • Industrial operating system resource instance scheduling optimization method

    CN119149200A

  • RIS auxiliary vehicle edge calculation method based on improved DRL

    CN119363779A

  • DRL-based active RIS-assisted MISO communication system joint optimization method

    CN120151897A

  • Multi-RIS communication network rate improving method based on MADDPG

    CN114727318A