A method and system for double reconfigurable intelligent surface assisted one-time pad wireless secure communication

By combining dual reconfigurable smart surfaces with multi-agent reinforcement learning, the phase shift and key generation processes of the RIS reflector unit are optimized, solving the problems of low key generation rate and insufficient data transmission in wireless secure communication, and realizing efficient and secure wireless data transmission.

CN122458023APending Publication Date: 2026-07-24SOUTHEAST UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHEAST UNIV
Filing Date
2026-03-19
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing wireless physical layer security technologies suffer from low key generation rates and insufficient optimization of data transmission and security. Traditional optimization methods struggle to handle the dynamic coupling between RIS phase shift and the number of key generation rounds, resulting in limited performance of secure wireless communication.

Method used

A dual-reconfigurable intelligent surface-assisted approach is adopted, combined with a multi-agent reinforcement learning (MARL) optimization architecture. By controlling the phase shift and key generation process of the RIS reflection unit, a dynamic balance between channel randomness and determinism is achieved, optimizing key generation and data transmission. XOR encryption is used to achieve one-time pad data transmission.

Benefits of technology

It significantly improves key generation rate, enhances data transmission efficiency, strengthens anti-eavesdropping capabilities, reduces training and deployment complexity, adapts to various wireless communication protocols, supports multi-user and heterogeneous network scenarios, and improves communication stability and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122458023A_ABST
    Figure CN122458023A_ABST
Patent Text Reader

Abstract

This invention proposes a dual reconfigurable smart surface (RIS)-assisted "one-time pad" secure wireless communication method to address the problems of low physical layer key generation rate and insufficient key rate to meet the requirements of one-time pad data transmission in wireless communication. This method constructs a controllable-random channel environment by deploying two N-unit RISs. It utilizes a multi-agent reinforcement learning (MARL) algorithm to jointly optimize the RIS phase shift and the number of key negotiation rounds, achieving a dynamic balance between data transmission rate and key generation rate. Based on this, it achieves absolutely secure communication for one-time pad data transmission. First, the legitimate communicating parties (Alice and Bob) construct a multipath random wireless channel by controlling the dual RISs based on channel reciprocity in time-division duplex (TDD) mode. Second, the MARL algorithm (including a pre-training mechanism) is used to control and optimize the phase shift of the dual RISs and determine the optimal number of key negotiation rounds for one-time pad. Finally, the physical layer key is generated through channel probing, channel estimation, quantization negotiation, and privacy amplification. Before data transmission, the generated key is used to XOR the transmitted data, achieving lightweight encryption of the transmitted data. This method significantly improves the physical layer key generation rate, provides a feasible solution for realizing "one-time key" data transmission, and is applicable to various wireless security communication scenarios such as the Internet of Things, the Internet of Vehicles, and satellite communication. It can take into account both the complexity of edge device deployment and real-time requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of physical layer security in wireless communication, specifically involving a wireless secure communication method that combines dual reconfigurable smart surfaces (RIS) and multi-agent reinforcement learning (MARL) to simultaneously improve the key generation rate and data transmission efficiency of wireless communication. Through an intrinsic design that integrates communication security, it adapts to the data security transmission requirements of "one-time pad". Background Technology

[0002] With the evolution of 5G / 6G wireless networks, the demand for wireless communication transmission in scenarios such as the Internet of Things (IoT) and the Internet of Vehicles (IoV) is surging. However, the openness of wireless networks leads to serious eavesdropping threats to data transmission. Traditional secure wireless communication relies on upper-layer cryptographic mechanisms, and the key generation and distribution process is highly complex, which is not suitable for resource-constrained IoT and IoV scenarios. Complex cryptographic algorithms are difficult to deploy on edge IoT devices. In addition, the development of quantum computing exposes traditional cryptographic systems to the risk of being cracked.

[0003] Physical layer key generation technology generates secure keys based on the inherent characteristics of wireless channels, such as reciprocity and randomness, without relying on upper-layer protocols, offering advantages such as low latency and lightweight design. In Time Division Duplex (TDD) mode, legitimate communicating parties (Alice, Bob) can utilize channel reciprocity to generate a shared key for encrypted wireless data transmission. However, limited by channel coherence time, the physical layer key generation rate (KGR) is relatively low; simultaneously, a passive eavesdropper (Eve) can obtain relevant wireless channel information by approaching the legitimate communicating parties, thereby inferring the key sequence and degrading the security of the physical layer key.

[0004] Reconfigurable smart surfaces (RIS), as a novel information metamaterial, can artificially construct programmable wireless channels by adjusting the phase shift of units, providing a new approach to improving physical layer security. In existing technologies, RIS are mostly used alone in the field of communication for channel enhancement or interference suppression, but lack consideration for the integrated design of key generation and data transmission. Regarding the optimization of this integrated design, traditional optimization methods struggle to handle the dynamic coupling relationship between RIS phase shift and the number of key negotiation rounds, resulting in limited performance of the key generation and data transmission system. Therefore, there is an urgent need for a one-time pad wireless secure communication method that can jointly optimize RIS configuration and the key generation process, improving key generation rate while maintaining wireless data transmission efficiency. Summary of the Invention

[0005] This invention aims to address the problems of low key generation rate and insufficient co-optimization of data transmission and security in existing wireless physical layer security technologies. It provides a dual RIS-assisted one-time pad wireless secure communication method that uses the MARL-optimized architecture to regulate key generation and data transmission to achieve a dynamic balance between the two, thereby realizing secure one-time pad data transmission.

[0006] To achieve the above objectives, the present invention provides the following solution: a dual reconfigurable smart surface-assisted one-time-one-key wireless secure communication method, the method comprising the following steps:

[0007] Step 1: Construct a wireless communication system assisted by dual reconfigurable smart surfaces RIS1 and RIS2. In time-division duplex (TDD) mode, acquire channel state information between the legitimate communicating parties Alice and Bob, as well as information related to the eavesdropper Eve. Set the value ranges for the total number of communication rounds L, the number of key negotiation rounds Q, and the phase shift value range for the RIS reflection unit. Step 2: In the key generation phase, a random phase shift is applied to RIS1 to construct a multipath random wireless channel, and a controllable phase shift is applied to RIS2 to adjust the equivalent channel, thereby performing channel detection and channel estimation within a predetermined Q rounds to obtain a reciprocal channel observation sequence for physical layer key generation; Step 3: The phase shift control strategy of RIS2 and the key negotiation round number Q are trained and optimized using the multi-agent reinforcement learning (MARL) algorithm with a pre-training mechanism. A reward function is constructed using the one-time key-data matching objective and the communication rate objective, outputting the optimal RIS2 phase shift configuration and the optimal key negotiation round number Q*; Step 4: Based on the optimal key negotiation round number Q and the optimal RIS2 phase shift configuration, channel quantization, information negotiation / error correction consistency processing, and privacy amplification processing are performed in the key generation phase to generate a shared key of length Q·Rk; Step 5: In the data transmission phase, the RIS1 phase shift is kept constant and the RIS2 phase shift is fixed to the optimal RIS2 phase shift configuration. The shared key is used to pair keys of length Q*. The data is encrypted and decrypted once to achieve secure wireless data transmission; Step 6: When the channel environment changes or the key-data matching constraint does not meet the preset threshold, return to steps 2 to 5 to update the RIS phase shift and key negotiation rounds.

[0008] Before data transmission, the transmitted data is XORed using the generated key to achieve lightweight encryption of the transmitted data.

[0009] The core technical solution of this invention includes three parts: dual RIS random wireless channel construction, MARL optimization framework design, and key generation and data transmission coordination mechanism, as detailed below:

[0010] Constructing a dual-RIS random wireless channel requires deploying two N-unit RIS (RIS1 and RIS2) units between Alice and Bob. Randomness is introduced through random phase shift variations in RIS1, while determinism is introduced through controllable phase shifts in RIS2. Taking the Rayleigh fading model as an example, the channel gain h follows a Gaussian distribution with mean 0 and variance σ². The equivalent channel can be expressed as:

[0011]

[0012] Where Φ1 is the RIS1 random phase matrix, Φ2 is the RIS2 controllable phase matrix, and z is the signal estimation error. .

[0013] The basic algorithm of the MARL optimization framework adopts the MADDPG algorithm, which includes two Actor networks and one Critic network:

[0014] Actor1: Input local CSI, }

[0015] Output the phase shift matrix of RIS2 A pre-training mechanism is adopted, which uses random channel data and optimal phase configuration samples for training to improve the initial performance of the model.

[0016] Actor2: Input global CSI,

[0017] }

[0018] Output key negotiation rounds (1≤Q≤L / 2, where L is the total number of negotiation rounds); using the cross-entropy loss function for training, the target number of rounds to approximate the theoretical optimal number can be expressed as:

[0019]

[0020] The Critic network takes the global CSI and the actions output by the two Actors as input, calculates the state-action value (Q value) using the MSE loss function, and guides the update of the Actor network parameters. The training process employs an experience replay and target network soft update mechanism to ensure the stability of convergence.

[0021] Key generation and data transmission coordination mechanism

[0022] Key generation process: Alice and Bob exchange probe signals, perform CSI estimation, quantize channel characteristics, negotiate keys, apply privacy amplification to the keys, and generate physical layer keys. The key generation rate is expressed as:

[0023]

[0024] Data transmission process: Based on the generated security key, the sending end Alice encrypts the transmitted data using an XOR method, and then sends the data to the receiving end Bob via a wireless channel. The data transmission rate is... Represented as:

[0025]

[0026] in, , The number of antennas for Alice and Bob, respectively. For legitimate channel correlation coefficients, Here, γ is the correlation coefficient of the eavesdropping channel, and γ is the reference signal-to-noise ratio. Let h be the phase shift diagonal matrix of the i-th RIS, and h be the channel. ,

[0027] Key generation and data collaboration can be represented as optimization objectives:

[0028]

[0029] By balancing the key generation rate and data transmission rate, and optimizing the number of key negotiation rounds, it approaches Shannon's "one-time pad" wireless secure communication.

[0030] This method supports multi-antenna configuration and is compatible with various wireless communication protocols such as LoRa, WiFi, and 5G NR. It can be applied to scenarios such as secure communication of IoT devices, encrypted data transmission in vehicle networks, secure authentication of satellite communication, and privacy protection of smart homes.

[0031] The MARL module employs, but is not limited to, deep reinforcement learning algorithms such as MADDPG. The training process includes experience replay and soft updates of the target network. The optimizer uses Adam to ensure training stability and convergence speed.

[0032] Beneficial effects:

[0033] Significantly improved key generation rate and enhanced adaptability to low-entropy environments: Through the collaborative mechanism of "random perturbation + controllable optimization" of dual RIS, the key generation stage is significantly improved by random RIS ( By randomly varying the phase shift to enhance channel time variation and multipath randomness, additional entropy sources can be continuously injected even in relatively static or low-motion scenarios, thus significantly improving the key generation rate. Compared to schemes without RIS assistance, the key generation rate can be improved by more than 60%, with particularly significant gains in long coherence time scenarios.

[0034] Key-data matching more closely approximates the one-time pad theoretical constraints, improving effective secure throughput: Controllable RIS is jointly optimized through the MARL module. The phase shift and the number of key negotiation rounds Q make the ratio of the generated key length to the data length closer to 1, thereby avoiding "insufficient keys leading to OTP failure" or "excessive keys causing resource waste" and improving the effective security throughput under one-time pad conditions. When the matching degree is close to the theoretical optimum, the required number of key negotiation rounds is reduced by about 10%–15% compared to the algorithm without RIS assistance.

[0035] Reduced training and deployment complexity, adaptable to large-scale RIS and real-time control: A "CSI→" model is established before entering the online reinforcement learning phase through a pre-training mechanism (and optional transfer learning freeze mechanism). The efficient initialization mapping of "phase shift" reduces the exploration cost of continuous high-dimensional action spaces, decreases the number of training iterations, and improves convergence stability. In the online deployment phase, only forward inference is required to output the results. Phase shift and This reduces the computational overhead of real-time control and makes it suitable for scalable deployments as the number of RIS units increases.

[0036] Enhanced resistance to eavesdropping and convergence of leakage surfaces: stochastic RIS ( Introducing independent perturbations during the key generation stage reduces the correlation between the eavesdropper's observation channel and the legitimate channel, thereby lowering the probability that the eavesdropper will infer the key from the relevant channel. At the same time, multiple rounds of key negotiation and privacy amplification further compress the information that may be leaked, improve the statistical security and robustness of the key, and reduce the risk of eavesdropping.

[0037] A two-stage RIS control strategy reduces data phase fluctuations and improves communication stability: This invention explicitly divides the communication session into a key generation phase and a data transmission phase: the key generation phase is activated... Randomization to improve KGR, data transmission phase stops Random variation and fixed Phase-shift configuration avoids rate fluctuations caused by random disturbances in data transmission, thereby improving the stability and reliability of OTP encrypted data transmission.

[0038] It exhibits better robustness and recoverability to CSI errors and environmental changes: The MARL module updates its policy based on interactive feedback, enabling continuous adaptation in the presence of channel estimation errors, measurement noise, or environmental changes; and it can be configured to trigger a re-optimization mechanism when the key-data ratio deviates from the threshold, automatically updating the RIS configuration and negotiation rounds when the system performance deviates from the one-time one-pad matching target, thereby improving the stability and recoverability of the system in long-term operation.

[0039] The hardware feasibility and protocol compatibility are stronger: RIS phase shift can be implemented using finite bit quantization (such as B-bit discrete phase), and the pre-training / discrete action set mechanism is naturally adapted to this hardware constraint; at the same time, the system architecture of this invention does not depend on a specific modulation method or a specific upper-layer protocol, and can be integrated with various communication protocol stacks such as multi-antenna MIMO, IoT / V2X, etc., and has good engineering feasibility.

[0040] Easily scalable to multi-user / multi-RIS / heterogeneous network scenarios: Due to its modular design (channel estimation module, MARL module, key generation module, OTP module), this invention can be extended to scenarios with multi-user access, multiple RIS collaboration, different topologies, and different security requirements. Expansion can be achieved by adding agents or adjusting state / action definitions, demonstrating good versatility. Attached Figure Description

[0041] Figure 1. Schematic diagram of transfer learning.

[0042] Figure 2 MARL module training framework diagram.

[0043] Figure 3 shows the training and inference process of the policy network and value network in the system model. Detailed Implementation

[0044] Example 1: See Figures 1-3 A one-time, one-key secure wireless communication method assisted by a dual reconfigurable smart surface, the method comprising the following steps:

[0045] Step 1: Constructing a dual reconfigurable smart surface and The auxiliary wireless communication system, in time-division duplex (TDD) mode, acquires channel state information between the legitimate communicating parties Alice and Bob, as well as information related to the eavesdropper Eve, and sets the range of values ​​for the total number of communication rounds L, the number of key negotiation rounds Q, and the phase shift range of the RIS reflection unit; Step 2: In the key generation stage, for Apply random phase shifts to construct multipath random wireless channels, and for A controllable phase shift is applied to adjust the equivalent channel, thereby performing channel detection and channel estimation within a predetermined Q rounds to obtain a reciprocal channel observation sequence for physical layer key generation; Step 3: The multi-agent reinforcement learning (MARL) algorithm with a pre-training mechanism is used to... The phase shift control strategy and the key negotiation round number Q are trained and optimized. A reward function is constructed using the one-time key-data matching objective and the communication rate objective, and the optimal output is obtained. Phase shift configuration and optimal key negotiation rounds Step 4: Based on the optimal key negotiation round number Q and the optimal... Phase-shift configuration performs channel quantization, information negotiation / error correction consistency processing, and privacy amplification processing during the key generation phase, generating a key with a length of [length missing]. Step 5: Maintain the shared key during data transmission. The phase shift no longer changes randomly and is fixed. The phase shift is the optimal Phase shift configuration, utilizing the shared key pair with a length of The data is encrypted and decrypted once to achieve secure wireless data transmission; Step 6: When the channel environment changes or the key-data matching constraint does not meet the preset threshold, return to steps 2 to 5 to update the RIS phase shift and key negotiation rounds.

[0046] Example 2: A dual reconfigurable smart surface-assisted one-time pad wireless secure communication system, comprising legitimate communication nodes (Alice, Bob), an eavesdropping node (Eve), a dual RIS module, a wireless channel estimation module, a MARL module, a key generation module, and communication nodes; the dual RIS module includes a random phase RIS (RIS1) and a controllable phase RIS (RIS2), deployed in the wireless propagation path between Alice and Bob; the channel estimation module is located between the legitimate communication nodes and the dual RIS module, used to collect channel state information (CSI); the MARL module is located behind the channel estimation module, used to output the phase shift matrix of RIS2 and the optimal key negotiation round number Q; the key generation module is located behind the MARL module, used to generate the key for data encryption; the communication nodes perform an XOR operation between the generated physical layer key and the transmitted data to achieve one-time pad data secure transmission;

[0047] In the dual RIS module: the phase shift matrix Φ1 of RIS1 changes randomly to introduce channel randomness and prevents it from changing randomly during data transmission; the phase shift matrix Φ2 of RIS2 is a controllable variable, dynamically optimized through a reinforcement learning algorithm. Both Φ1 and Φ2 are diagonal matrices, satisfying... , where θ n The phase shift of the nth RIS unit (0≤θ≤2π);

[0048] The MARL module includes two Actor networks (Actor1 and Actor2) and one Critic network, with Actor1 employing a pre-training mechanism; wherein:

[0049] (1) Actor1 receives local CSI related to RIS2:

[0050]

[0051] Output the phase shift matrix of RIS2: ;

[0052] (2) Actor2 receives global CSI:

[0053]

[0054] Output key negotiation round number: ;

[0055] (3) The Critic network receives the actions from the global CSI and the dual Actor outputs, evaluates the state-action value through the mean squared error (MSE) loss function, and guides the Actor network parameter update.

[0056] The pre-training process of Actor1 includes: randomly generating multiple sets of Channel State Information (CSI); randomly sampling multiple sets of RIS2 phase shift configurations for each CSI and calculating the corresponding reward value; selecting the phase shift configuration with the largest reward value as the supervision label; and training Actor1 with the mean square error (MSE) as the loss function so that its output approximates the supervision label.

[0057] The implementation steps of the key generation module include:

[0058] Step 1: Alice and Bob exchange probe signals, and the CSI estimation is completed through the channel estimation module to obtain the equivalent channel. ;

[0059] Step 2: Based on the results of multiple rounds of detection and channel estimation, the channel characteristics are quantized;

[0060] Step 3: Modify the inconsistent key through a key negotiation mechanism;

[0061] Step 4: Perform privacy amplification to enhance key security.

[0062] The optimization goal of secure communication methods is to maximize the sum of data transmission rate and key generation rate on the basis of one-time pad, while minimizing the deviation between the number of key negotiation rounds and the theoretical optimal value, and finally achieve one-time pad;

[0063] The key generation and data transmission coordination mechanism is as follows.

[0064] Key generation process: Alice and Bob exchange probe signals, perform CSI estimation, quantize channel characteristics, negotiate keys, apply privacy amplification to the keys, and generate physical layer keys. The key generation rate is expressed as:

[0065]

[0066] Data transmission process: Based on the generated security key, the sending end Alice encrypts the transmitted data using an XOR method, and then sends the data to the receiving end Bob via a wireless channel. The data transmission rate is... Represented as:

[0067] in, , The number of antennas for Alice and Bob, respectively. For legitimate channel correlation coefficients, Here, γ is the correlation coefficient of the eavesdropping channel, and γ is the reference signal-to-noise ratio. Let h be the phase shift diagonal matrix of the i-th RIS, and h be the channel. ,

[0068] Key generation and data collaboration can be represented as optimization objectives:

[0069]

[0070] By balancing the key generation rate and data transmission rate, and optimizing the number of key negotiation rounds, it approximates Shannon's "one-time pad" wireless security communication. For data transmission rate, For key generation rate, Let the angle be the nth element of the RIS. L represents the number of rounds for unilateral key negotiation, and L represents the total number of rounds in the communication session.

[0071] The system satisfies a dual RIS-assisted equivalent channel model, and the channel estimation error is complex Gaussian noise. The equivalent channels of both the legitimate link and the eavesdropping link contain... and Multipath components, direct components, and estimation error noise components of the reflection.

[0072] The characteristic is that the channel satisfies:

[0073]

[0074] in, For inter-node channel gain, For channel estimation error, The noise error generated for Bob The noise error generated by Alice Noise error generated for Eve.

[0075] The pre-training phase will Phase shift control actions are discretized into a set of actions containing K representative phase shift configurations. And based on the action set, value learning or supervised pre-training is performed to reduce the complexity of action space exploration.

[0076] After completing the pre-training, The hidden layer parameters are frozen, and the output layer is replaced with a continuous phase-shift regression output layer. Only the output layer parameters are fine-tuned during training. Output continuous phase shift control vector To reduce gradient propagation complexity and improve training stability;

[0077] The system sets a condition for determining whether the ratio of key length to data length deviates from a threshold ε. This triggers retraining or re-optimization. Phase shift and key negotiation rounds.

[0078] Example 3: To make the technical solution of the present invention clearer, the following detailed description is provided in conjunction with specific embodiments, including the deployment and configuration of the dual RIS-assisted wireless secure communication system, the system training and optimization process, the method deployment, and performance evaluation:

[0079] 1: System Deployment and Configuration

[0080] Legitimate communication partners: Alice and Bob each have 4 antennas. Eve the eavesdropper is equipped with two antennas. ;

[0081] Dual RIS configuration: Both RIS1 and RIS2 contain 20 reflection units (N=20). The phase of RIS1 is updated randomly (update interval 2.5ms), and the phase of RIS2 is dynamically output by Actor1.

[0082] Communication parameters: Total number of communication rounds L=400, reference signal-to-noise ratio Channel estimation noise variance Reinforcement learning parameters: hidden layer size 256, learning rate Batch size 64, training rounds 500.

[0083] 2: System Training and Optimization Process

[0084] Pre-training phase:

[0085] Since this experiment uses the continuous change of RIS phase shift as the action space, and the key generation rate is less sensitive to the change of RIS phase shift than the data transmission rate, a pre-trained network is used to assist the stable start-up of multi-agent reinforcement learning.

[0086] 10,000 sets of channel state information are randomly generated, and 20 sets of RIS2 phase configurations are sampled for each set of CSI.

[0087] Calculate the reward function for each configuration group. Solving for the optimal configuration ;

[0088] Actor1 is trained using MSE as the loss function, with CSI as the input and phase shift as the output, to make the network approximate... The pre-training process involved 5000 iterations.

[0089] like Figure 1 As shown, the hidden layer of the neural network is first trained discretely using DQN to optimize its feature parsing ability for CSI. Then, the hidden layer is frozen and transferred to DDPG to train continuous actions.

[0090] MARL module training phase:

[0091] like Figure 2 As shown, the pre-trained Actor1 network (responsible for controlling RIS2) and Actor2 network (responsible for outputting the number of key negotiation rounds) are put into the MADDPG training framework for training.

[0092] Initialize MADDPG algorithm parameters: Actor1, Actor2 and Critic network weights, target network synchronization initial weights, and experience replay buffer D initialization;

[0093] Each episode resets the environment state s, iterating through 400 time steps: a. Actor1 is based on Output Actor2 is based on Output b. Add exploration noise ε; c. Perform the action to obtain the next state s' and reward. ,Will c. Store the data in D; d. Randomly sample 64 samples from D and calculate the Critic target value.

[0094] d. Update the weights of the Critic network; e. Update the weights of Actor1 and Actor2 via policy gradient; f. Softly update the weights of the target network. .

[0095] Deployment and execution phase:

[0096] The edge device loads the trained model and collects CSI data in real time;

[0097] Actor1 outputs the optimal phase shift for RIS2, and Actor2 outputs the optimal number of key negotiation rounds Q.

[0098] The key generation process (Q-round negotiation) and data transmission process (L-2Q-round transmission) are executed, and data encryption wireless communication is achieved based on a simple XOR operation.

[0099] 3: Performance evaluation, key generation rate: 1.2 Mb / s in RIS-assisted scenarios, an improvement of more than 70% compared to scenarios without RIS;

[0100] Key negotiation rounds: 140-145 rounds after optimization, close to the theoretical optimal value of 133-134 rounds, while the number of data transmission rounds increases;

[0101] Inference latency: Low inference latency, meeting the real-time requirements of edge devices.

[0102] It should be noted that the above embodiments are not intended to limit the scope of protection of the present invention. Equivalent transformations or substitutions made based on the above technical solutions all fall within the scope of protection of the claims of the present invention.

Claims

1. A one-time, one-key wireless secure communication method assisted by a dual reconfigurable smart surface, characterized in that, The method includes the following steps: Step 1: Constructing a dual reconfigurable smart surface and The auxiliary wireless communication system, in time-division duplex (TDD) mode, acquires channel state information between the legitimate communicating parties Alice and Bob, as well as information related to the eavesdropper Eve, and sets the range of values ​​for the total number of communication rounds L, the number of key negotiation rounds Q, and the phase shift range of the RIS reflection unit; Step 2: In the key generation stage, for Apply random phase shifts to construct multipath random wireless channels, and for A controllable phase shift is applied to adjust the equivalent channel, thereby performing channel detection and channel estimation within a predetermined Q rounds to obtain a reciprocal channel observation sequence for physical layer key generation; Step 3: The multi-agent reinforcement learning (MARL) algorithm with a pre-training mechanism is used to... The phase shift control strategy and the key negotiation round number Q are trained and optimized. A reward function is constructed using the one-time key-data matching objective and the communication rate objective, and the optimal output is obtained. Phase shift configuration and optimal key negotiation rounds Step 4: Based on the optimal key negotiation round number Q and the optimal... Phase-shift configuration performs channel quantization, information negotiation / error correction consistency processing, and privacy amplification processing during the key generation phase, generating a key with a length of [length missing]. Step 5: Maintain the shared key during data transmission. The phase shift no longer changes randomly and is fixed. The phase shift is the optimal Phase shift configuration, utilizing the shared key pair with a length of The data is encrypted and decrypted once to achieve secure wireless data transmission; Step 6: When the channel environment changes or the key-data matching constraint does not meet the preset threshold, return to steps 2 to 5 to update the RIS phase shift and key negotiation rounds.

2. A dual reconfigurable smart surface-assisted one-time-key wireless secure communication system, characterized in that, Used to implement the secure communication method of claim 1, The system includes legitimate communication nodes (Alice, Bob), an eavesdropping node (Eve), a dual RIS module, a wireless channel estimation module, a MARL module, a key generation module, and communication nodes. The dual RIS module includes a random phase RIS (RIS1) and a controllable phase RIS (RIS2), deployed in the wireless propagation path between Alice and Bob. The channel estimation module is located between the legitimate communication nodes and the dual RIS module and is used to collect channel state information (CSI). The MARL module is located behind the channel estimation module and is used to output the phase shift matrix of RIS2 and the optimal key negotiation round number Q. The key generation module is located behind the MARL module and is used to generate the key for data encryption. The communication nodes perform an XOR operation between the generated physical layer key and the transmitted data to achieve "one-time pad" secure data transmission.

3. The dual reconfigurable smart surface-assisted one-time-key wireless secure communication system according to claim 2, characterized in that, In the dual RIS module: the phase shift matrix Φ1 of RIS1 changes randomly to introduce channel randomness and prevents it from changing randomly during data transmission; the phase shift matrix Φ2 of RIS2 is a controllable variable, dynamically optimized through a reinforcement learning algorithm. Both Φ1 and Φ2 are diagonal matrices, satisfying... , where θ n The phase shift of the nth RIS unit .

4. The dual reconfigurable smart surface-assisted one-time-key wireless secure communication system according to claim 3, characterized in that, The MARL module includes two Actor networks (Actor1 and Actor2) and one Critic network, with Actor1 employing a pre-training mechanism; wherein: (1) Actor1 receives local CSI related to RIS2: Output the phase shift matrix of RIS2: ; (2) Actor2 receives global CSI: Output key negotiation round number: ; (3) The Critic network receives the actions from the global CSI and the dual Actor outputs, evaluates the state-action value through the mean squared error (MSE) loss function, and guides the Actor network parameter update.

5. The dual reconfigurable smart surface-assisted one-time-key wireless secure communication system according to claim 4, characterized in that: The pre-training process of Actor1 includes: randomly generating multiple sets of Channel State Information (CSI); randomly sampling multiple sets of RIS2 phase shift configurations for each CSI and calculating the corresponding reward value; selecting the phase shift configuration with the largest reward value as the supervision label; and training Actor1 with mean square error (MSE) as the loss function so that its output approximates the supervision label.

6. The dual reconfigurable smart surface-assisted one-time-key wireless secure communication system according to claim 5, characterized in that, The implementation steps of the key generation module include: Step 1: Alice and Bob exchange probe signals, and the CSI estimation is completed through the channel estimation module to obtain the equivalent channel. ; Step 2: Based on the results of multiple rounds of detection and channel estimation, the channel characteristics are quantized; Step 3: Modify the inconsistent key through a key negotiation mechanism; Step 4: Perform privacy amplification to enhance key security.

7. The dual reconfigurable smart surface-assisted one-time-key wireless secure communication system according to claim 6, characterized in that, The optimization goal of secure communication methods is to maximize the sum of data transmission rate and key generation rate on the basis of one-time pad, while minimizing the deviation between the number of key negotiation rounds and the theoretical optimal value, and finally achieve one-time pad; The key generation and data transmission coordination mechanism is as follows. Key generation process: Alice and Bob exchange probe signals, perform CSI estimation, quantize channel characteristics, negotiate keys, apply privacy amplification to the keys, and generate physical layer keys. The key generation rate is expressed as: Data transmission process: Based on the generated security key, the sending end Alice encrypts the transmitted data using an XOR method, and then sends the data to the receiving end Bob via a wireless channel. The data transmission rate is... Represented as: in, , The number of antennas for Alice and Bob, respectively. For legitimate channel correlation coefficients, Here, γ is the correlation coefficient of the eavesdropping channel, and γ is the reference signal-to-noise ratio. Let h be the phase shift diagonal matrix of the i-th RIS, and h be the channel. , Key generation and data collaboration can be represented as optimization objectives: By balancing the key generation rate and data transmission rate, and optimizing the number of key negotiation rounds, it approximates Shannon's "one-time pad" wireless security communication. For data transmission rate, For key generation rate, Let the angle be the nth element of the RIS. L represents the number of rounds for unilateral key negotiation, and L represents the total number of rounds in the communication session.

8. The dual reconfigurable smart surface-assisted one-time-one-key wireless secure communication system according to claim 7, wherein the system satisfies the equivalent channel model assisted by dual RIS, and the channel estimation error is complex Gaussian noise, and the equivalent channels of the legitimate link and the eavesdropping link both contain... and Multipath components, direct components, and estimation error noise components of the reflection. Its features are, The channel satisfies: in, For inter-node channel gain, For channel estimation error, The noise error generated for Bob The noise error generated by Alice Noise error generated for Eve.

9. The dual reconfigurable smart surface-assisted one-time-key wireless secure communication system according to claim 5, characterized in that: The pre-training phase will Phase shift control actions are discretized into a set of actions containing K representative phase shift configurations. And based on the action set, value learning or supervised pre-training is performed to reduce the complexity of action space exploration.

10. The dual reconfigurable smart surface-assisted one-time-key wireless secure communication system according to claim 9, characterized in that: After completing the pre-training, The hidden layer parameters are frozen, and the output layer is replaced with a continuous phase-shift regression output layer. Only the output layer parameters are fine-tuned during training. Output continuous phase shift control vector To reduce gradient propagation complexity and improve training stability; The system sets the key length to data length ratio to deviate from a threshold. The judgment condition is when Triggering retraining or re-optimization Phase shift and key negotiation rounds.