A railway internet of things covert communication method and system based on deep reinforcement learning

By introducing intelligent metasurface RIS into the railway IoT communication system, constructing a RIS-MIMO model and using deep reinforcement learning to optimize beamforming and RIS phase shift, the reliability and cost issues of covert transmission in 5G-R and 6G-R railway communication systems are solved, and secure and efficient data transmission in high-speed mobile scenarios is achieved.

CN119697671BActive Publication Date: 2025-10-10LANZHOU JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510133079.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-10-10
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

In 5G-R and 6G-R railway communication systems, existing technologies make it difficult to achieve reliable and covert transmission while ensuring the quality of received signals and effectively reducing communication network operating costs. Especially in high-speed mobile scenarios, where channels change rapidly and timeliness requirements are high, existing methods find it difficult to accurately obtain channel status information, resulting in system performance loss.

Method used

A railway Internet of Things covert communication method based on deep reinforcement learning is adopted. By adding an intelligent metasurface RIS to the railway Internet of Things communication system, a RIS-MIMO model is constructed. Combined with the Doppler frequency shift compensation caused by high-speed train movement and the channel uncertainty introduced by artificial noise, a channel model is constructed. The non-convexity and coupling are handled through a deep reinforcement learning framework, and the beamforming at the transmitting and receiving ends and the phase shift of the intelligent metasurface RIS are optimized to ensure reliable and low-cost covert communication under fast time-varying channels.

Benefits of technology

Secure transmission of railway IoT data is achieved under fast time-varying channels, ensuring the security and rationality of privacy data, reducing system resource consumption, ensuring driving safety and passenger privacy, and effectively reducing the security risks faced by the new generation of railway IoT. It is low-cost and easy to deploy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119697671B_ABST
    Figure CN119697671B_ABST
Patent Text Reader

Abstract

The application discloses a railway Internet of Things covert communication method and system based on deep reinforcement learning, and belongs to the technical field of communication security, which comprises the following steps: adding an intelligent metasurface in a railway Internet of Things communication system, and constructing a high-speed railway Internet of Things covert communication model based on RIS-MIMO; constructing a channel model of the railway Internet of Things covert communication system based on the time-varying nature of the channel; designing a concealment constraint condition according to the channel model and a received signal model based on hypothesis testing, and considering the priority of the railway Internet of Things communication service; and taking the maximization of the concealment throughput as the target, and forming a joint optimization problem with the designed concealment condition, the total power of the base station and the RIS phase shift as the constraints. After implementing the joint optimization through the proposed low-complexity deep reinforcement learning framework, the railway Internet of Things communication system can be more reliable in covert transmission, power resources can be saved, and the cost is low and the deployment is easy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication security technology, and in particular to a railway Internet of Things covert communication method and system based on deep reinforcement learning. Background Art

[0002] In intelligent high-speed rail communication systems powered by fifth-generation mobile communications technology (5G), the scale of IoT devices and the amount of data transmitted are expected to explode. However, to meet the needs of IoT devices for unfettered mobility, the wireless networking methods and broadcast nature of wireless channels make both 5G-Rail (5G-R) and the upcoming 6G-R vulnerable to interference and attacks from eavesdropping devices. The data generated by railway IoT devices includes confidential business data and user information. Leakage of this data could significantly harm the privacy of both businesses and users.

[0003] While cryptography-based upper-layer encryption and physical-layer security technologies can protect data transmission to a certain extent, this fundamental approach to ensuring information is not deciphered by unauthorized eavesdroppers is becoming increasingly vulnerable with the advent of quantum computing and other supercomputing technologies. Secure transmission of private data involves more than just the content of communications; concealing the communication itself is also crucial. Only by remaining undetected can communication be protected from interference or deciphering. Therefore, wireless covert communication technology, with its unique advantage of concealing communication behavior, offers a solution to the need for confidentiality in data communication within the railway Internet of Things. However, in the field of high-speed rail communication technology, eavesdroppers typically use high-precision energy meters for detection. Existing technologies have yet to find a secure communication method that can guarantee the quality of received signals on trains in 5G-R networks, while simultaneously achieving reliable and covert transmission and effectively reducing communication network operating costs.

[0004] Reconfigurable Intelligent Surfaces (RIS) have recently attracted widespread attention due to their ability to independently reflect incident electromagnetic waves with adjustable phase shifts, thereby enabling intelligent wireless environments that improve the quality of legitimate communications and protect receivers from detection by unauthorized devices. RIS-assisted wireless communication networks also feature low cost, low power consumption, and a simple structure, making them widely used in the field of information security.

[0005] In existing research, researchers have focused more on improving the overall covert throughput of RIS-assisted high-speed rail communication systems, while neglecting the priority of services within the railway Internet of Things. Allocating the same power resources and concealment levels to railway services of varying priorities is inappropriate and wastes resources. For example, high-priority services (railway emergency calls, emergency voice communications, etc.) should be assigned a higher covert transmission level to ensure train safety. Furthermore, in high-speed mobile scenarios, channels change rapidly and timeliness is crucial. Existing technologies and methods struggle to accurately acquire channel state information, resulting in significant losses in system performance. Summary of the Invention

[0006] The purpose of the present invention is to provide a railway Internet of Things covert communication method and system based on deep reinforcement learning, which solves the problem of secure transmission of railway Internet of Things data under fast time-varying channels.

[0007] To achieve the above object, the present invention provides the following solutions:

[0008] A railway Internet of Things covert communication method based on deep reinforcement learning, comprising:

[0009] Incorporating intelligent metasurface RIS into the railway IoT communication system, and constructing a high-speed railway IoT covert communication model based on RIS-MIMO.

[0010] Based on the RIS-MIMO-based high-speed rail IoT covert communication model, a channel model for the railway IoT covert communication system is constructed by considering the Doppler frequency shift compensation caused by the high-speed movement of the train and the channel uncertainty introduced by artificial noise.

[0011] According to the communication model of the railway Internet of Things covert communication system, a received signal model based on hypothesis testing is constructed;

[0012] Obtaining concealment constraints according to the received signal model and railway communication service priorities;

[0013] Combined with the aforementioned concealment constraints, with the goal of maximizing the concealment throughput, a joint optimization problem of the transmit-receive beamforming and the smart metasurface RIS phase shift is constructed.

[0014] According to the time-varying nature of the high-speed rail channel, the joint optimization problem is transformed into a Markov decision process problem;

[0015] Based on the transformed Markov decision process problem, a deep reinforcement learning framework is used to deal with the non-convexity and coupling of the problem, obtain the optimal solution to the problem, and ensure reliable and low-cost covert communication in fast time-varying channel scenarios.

[0016] Optionally, the RIS-MIMO-based high-speed rail IoT covert communication model includes:

[0017] A train moves from the center of the base station BS coverage area to the edge of the coverage area, a trackside base station equipped with multiple antennas and a large uniform planar array UPA, the smart metasurface RIS deployed between the base station BS and the train, the eavesdropper Willie and the trackside buildings. It is assumed that a mobile relay MR is installed on the top of the train for communicating with the base station BS.

[0018] Optionally, considering the Doppler shift compensation caused by high-speed train movement and the channel uncertainty introduced by artificial noise, the channel model of the railway IoT covert communication system is constructed, including:

[0019] According to the RIS-MIMO-based high-speed rail IoT covert communication model, the positions of the base station BS, the intelligent metasurface RIS, and the eavesdropper Willie are fixed, and the mobile relay MR moves with the train. The channel matrix H of the BS-RIS link is BR , the channel matrix W of the BS-Willie link BW And the channel matrix W of the RIS-Willie link RW It is constructed as a quasi-static channel without Doppler shift, and the channel matrix G of the BS-MR link is BM , the channel matrix H of the RIS-MR link RM And the channel matrix W of the MR-Willie link MW It is constructed as a time-varying channel with Doppler shift;

[0020] The construction of the time-varying channel is to represent the Doppler frequency shift by the Clarke-Jakes spectrum. For any time-varying channel H[t] at any time, let and H[t+Ts] represent the outdated channel state information CSI and the real-time channel state information CSI, respectively. The relationship between H[t+Ts] is:

[0021]

[0022] Where, κ=J0(2πf D T S ) represents the time correlation coefficient, f D is the Doppler frequency shift, J0(·) is the first kind of zero-order Bessel function, ΔH is the error term, t is any time during the train operation, and Ts is the transmission delay;

[0023] The mobile relay MR generates interference signals of different powers, and the self-interference channel of the mobile relay MR is f MCN(0,φ), φ∈[0,1] is the self-interference cancellation coefficient, CN is the symbolic representation of the complex Gaussian distribution, and it is assumed that the interference signal transmission power P of the mobile relay MR Mi Obey the interval [0, P Mi,max ], the probability density function is:

[0024]

[0025] Where, P Mi,max is the maximum self-interference transmission power, x is [0, P Mi,max ] on different values.

[0026] Optionally, constructing a received signal model based on hypothesis testing includes: constructing a received signal model at the eavesdropper Willie;

[0027] The received signal model at the eavesdropper Willie is:

[0028]

[0029] Among them, Y W is the receiving signal model of the eavesdropper Willie, W RW The conjugate transpose of , Θ is the reflection coefficient matrix of RIS, s M is the self-interference signal from the mobile relay MR, Z W is the additive white Gaussian noise at the eavesdropper Willie, H0 indicates that no covert communication occurs, M is the receiving beamforming matrix at the mobile relay MR, W BW The conjugate transpose of F H is the conjugate transpose of F, F is the transmit beamforming matrix at the base station BS, s B is a covert signal from the BS, and H1 indicates that covert communication occurs.

[0030] Optionally, obtaining the concealment constraint condition according to the received signal model and the railway communication service priority includes:

[0031] Obtaining a total detection error probability at the eavesdropper Willie according to a received signal model at the eavesdropper Willie;

[0032] According to the total detection error probability of the eavesdropper Willie, KL divergence and total variation distance are introduced, and Pinsker inequality is used to express the hidden constraint as follows:

[0033]

[0034] Where, T V (Lx , L y ) represents the variable L x With L y The total variation distance between x ||L y ) represents the value of the variable L x to L y KL divergence of (x,y) = (0,1) or (1,0);

[0035] Let 0≤ε≤1, and D(L x ||L y )≤2ε 2 Replace ξ≥1-ε as the concealment constraint, where ε is the prior value of the concealment constraint and ξ is the total detection error probability at the eavesdropper Willie;

[0036] Define the priority of high-speed rail communication services and use digital quantification to identify them. x ||L y )≤2ε 2 The values ​​of (x, y) and ε are related to the service priority.

[0037] Optionally, in combination with the concealment constraint condition, with the goal of maximizing the concealment throughput, constructing a joint optimization problem of the transmit-receiver beamforming and the smart metasurface RIS phase shift includes:

[0038] Obtaining a concealed throughput at the mobile relay MR based on short packet communication with a finite block length;

[0039] Taking the maximization of concealment throughput as the goal and combining the concealment constraint condition to construct the joint optimization problem;

[0040] The joint optimization problem is:

[0041]

[0042]

[0043] Where F is the transmit beamforming matrix at the base station BS, M is the receive beamforming matrix at the mobile relay MR, Θ is the reflection coefficient matrix of RIS, η is the concealed throughput at the mobile relay MR, D(L x ‖L y ) represents the value of the variable L x to L y KL divergence, ε is the prior value of the hidden constraint, P B is the maximum transmission power of the base station, N Rv is the total number of rows of RIS units, N Rhis the total number of columns of RIS units, m is the number of rows of RIS units, n is the number of columns of RIS units, C1, C2, and C3 are the covert constraints that satisfy the covert communication conditions, the maximum transmission power constraint of the base station BS, and the unit modulus constraint of the RIS phase shift, respectively. δ is the decoding error probability at MR, K is the packet length, and R is the information transmission rate from the base station BS to the mobile relay MR.

[0044] Optionally, converting the joint optimization problem into a Markov decision process problem includes:

[0045] The Markov decision process problem consists of a state set S, an action space A, and a state transition probability P. sa , reward function Re and γ discount factor;

[0046] The channel information of all nodes in the railway Internet of Things covert communication system, the transmission rate of the last time slot and the prior value of the covert constraint are used as the state set S;

[0047] Taking the transmit beamforming matrix at the base station BS, the receive beamforming matrix at the mobile relay MR and the RIS phase shift as an action space A;

[0048] The reward function Re is:

[0049]

[0050] Where Re is the reward function, M is the receive beamforming matrix at the mobile relay MR, H RM The conjugate transpose of , Θ is the reflection coefficient matrix of RIS, H BR is the channel matrix of the BS-RIS link, F H is the conjugate transpose of F, F is the transmit beamforming matrix at the base station BS, ω1 and ω2 are the weights of part 2 and part 3, respectively, ρ1 and ρ2 represent the satisfaction of the concealment constraint and the BS maximum transmit power constraint in each time slot t', respectively. Part 1 represents the direct utility, that is, the effective received signal, and parts 2 and 3 represent the cost function of each time slot t', corresponding to the cases where the concealment constraint requirements and the base station BS maximum transmit power requirements are not met, respectively.

[0051] Optionally, a deep reinforcement learning framework is used to address the non-convexity and coupling of the problem. Obtaining the optimal solution to the problem includes:

[0052] A deep reinforcement learning framework based on the dual actor-critic method DACM jointly optimizes the transmit-receive beamforming and the intelligent metasurface RIS phase shift, wherein the Actor network and the Critic network of the deep reinforcement learning framework are composed of two DNN networks.

[0053] Optional, dual-actor-critic method (DACM)-based deep reinforcement learning framework to jointly optimize the transmit-receive beamforming and smart metasurface RIS phase shifting, including:

[0054] S1. The central controller of the base station BS obtains status information s t’ , the state information s t’ Input the Actor network to perform action selection and obtain action a t’ 、Reward t’ and the next state s t’+1 ;

[0055] S2. Sequence (s t’ , a t’ ,Re t’ , s t’+1 ) is stored in the experience replay pool as a data set for training the network. According to the Banach fixed point theorem and the Bellman equation, the Critic network performs behavior evaluation based on the state value function and the action value function;

[0056] S3. Randomly sample some data from the experience replay pool to calculate the loss function and train the network parameters of the DNN network, return to S1, and repeatedly train until the network converges;

[0057] Wherein, the loss function is:

[0058]

[0059] in, is the loss function, τ is the size of the Mini-batch of the sampled dataset, V * is the state value, s v is the state of the vth data set, is the network parameter of the DNN network, Q * is the action value, a v is the vth data set action.

[0060] The present invention also provides a railway Internet of Things covert communication system based on deep reinforcement learning, comprising: a communication model construction module, a channel model construction module, a received signal model construction module, a constraint condition design module, an optimization problem construction module, a deep reinforcement learning algorithm design module, and an optimal solution calculation module;

[0061] The communication model construction module is used to add the intelligent metasurface RIS to the railway Internet of Things communication system to build a high-speed railway Internet of Things covert communication model based on RIS-MIMO;

[0062] The channel model construction module is used to construct a channel model for the railway Internet of Things covert communication system based on the RIS-MIMO-based high-speed railway Internet of Things covert communication model, taking into account the Doppler frequency shift compensation caused by the high-speed movement of the train and the channel uncertainty introduced by artificial noise;

[0063] The received signal model construction module is used to construct a received signal model based on hypothesis testing according to the communication model of the railway Internet of Things covert communication system;

[0064] The constraint condition design module is used to obtain hidden constraint conditions according to the received signal model and the railway communication service priority;

[0065] The optimization problem construction module is used to construct a joint optimization problem of the transceiver beamforming and the smart metasurface RIS phase shift in combination with the concealment constraint conditions with the goal of maximizing the concealment throughput;

[0066] The deep reinforcement learning algorithm design module is used to handle the non-convexity and coupling of the problem using a deep reinforcement learning framework based on the transformed Markov decision process problem;

[0067] The optimal solution calculation module is used to adjust the model parameters of the deep reinforcement learning framework and train the network model until the model converges to obtain the optimal solution to the problem.

[0068] The beneficial effects of the present invention are as follows: the present invention solves the problem of secure data transmission of the railway Internet of Things under fast time-varying channels, ensures the security and rationality of privacy data transmission in the railway Internet of Things, ensures driving safety and passenger privacy security while saving system resources, effectively reduces the security risks faced by the new generation of railway Internet of Things, and is low-cost and easy to deploy. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0070] Figure 1 This is a high-speed rail Internet of Things covert communication model based on RIS-MIMO in an embodiment of the present invention;

[0071] Figure 2 A deep reinforcement learning framework based on a dual-actor-critic method according to an embodiment of the present invention;

[0072] Figure 3This is a flow chart of a railway Internet of Things covert communication method based on deep reinforcement learning according to an embodiment of the present invention;

[0073] Figure 4 Schematic diagram of the impact of different hidden constraints on system performance according to an embodiment of the present invention;

[0074] Figure 5 is the network parameter learning rate β of the deep reinforcement learning framework of the embodiment of the present invention π Schematic diagram of the impact on the performance of the DRL algorithm;

[0075] Figure 6 Schematic diagram of the training loss of a neural network in a deep reinforcement learning framework according to an embodiment of the present invention. DETAILED DESCRIPTION

[0076] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0077] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0078] Example 1:

[0079] This embodiment provides a railway Internet of Things covert communication method based on deep reinforcement learning, including:

[0080] Incorporating intelligent metasurface RIS into the railway IoT communication system, and constructing a high-speed railway IoT covert communication model based on RIS-MIMO.

[0081] Based on the RIS-MIMO-based high-speed rail IoT covert communication model, the channel model of the railway IoT covert communication system is constructed by considering the Doppler frequency shift compensation caused by high-speed train movement and the channel uncertainty introduced by artificial noise.

[0082] According to the communication model of the railway Internet of Things covert communication system, a received signal model based on hypothesis testing is constructed;

[0083] Obtain hidden constraints based on the received signal model and railway communication service priorities;

[0084] Combined with the concealment constraints, with the goal of maximizing the concealment throughput, a joint optimization problem of the transmit-receive beamforming and the smart metasurface RIS phase shift is constructed.

[0085] According to the time-varying nature of the high-speed rail channel, the joint optimization problem is transformed into a Markov decision process problem;

[0086] Based on the transformed Markov decision process problem, a deep reinforcement learning framework is used to deal with the non-convexity and coupling of the problem, obtain the optimal solution to the problem, and ensure reliable and low-cost covert communication in fast time-varying channel scenarios.

[0087] Furthermore, the RIS-MIMO-based high-speed rail IoT covert communication model includes:

[0088] A train moves from the center of the base station BS coverage area to the edge of the coverage area, a trackside base station equipped with multiple antennas and a large uniform planar array (UPA), an intelligent metasurface (RIS) deployed between the base station BS and the train, an eavesdropper (Willie), and trackside buildings. It is assumed that a mobile relay (MR) is installed on the top of the train to communicate with the base station BS.

[0089] Furthermore, considering the Doppler shift compensation caused by high-speed train movement and the channel uncertainty introduced by artificial noise, the channel model of the railway IoT covert communication system is constructed, including:

[0090] According to the RIS-MIMO-based high-speed rail IoT covert communication model, the positions of the base station BS, the intelligent metasurface RIS, and the eavesdropper Willie are fixed, and the mobile relay MR moves with the train. The channel matrix H of the BS-RIS link is BR , the channel matrix W of the BS-Willie link BW And the channel matrix W of the RIS-Willie link RW It is constructed as a quasi-static channel without Doppler shift, and the channel matrix G of the BS-MR link is BM 、H RM And the channel matrix W of the MR-Willie link MW It is constructed as a time-varying channel with Doppler shift.

[0091] Furthermore, constructing a received signal model based on hypothesis testing includes: a received signal model at the eavesdropper Willie.

[0092] Furthermore, based on the received signal model and railway communication service priority, the hidden constraints are obtained, including:

[0093] According to the received signal model, the total detection error probability of the eavesdropper Willie is obtained; according to the total detection error probability of the eavesdropper Willie, the KL divergence and total variation distance are introduced, and the hidden constraint is expressed using the Pinsker inequality; the priority of the high-speed rail communication service is defined and quantified with a digital symbol, and D(L x ||Ly )≤2ε 2 The values ​​of (x, y) and ε are related to the service priority.

[0094] Furthermore, combined with the concealment constraints and aiming to maximize the concealment throughput, the joint optimization problem of the transmit-receive beamforming and the smart metasurface RIS phase shift is constructed, including:

[0095] Obtain the concealed throughput at the mobile relay based on short packet communication with finite block length;

[0096] The goal is to maximize the concealment throughput and combine it with concealment constraints to construct a joint optimization problem.

[0097] Furthermore, converting the joint optimization problem into a Markov decision process problem includes:

[0098] The Markov decision process problem consists of the state set S, action space A, and state transition probability P. sa , reward function Re and γ discount factor;

[0099] The channel information of all nodes in the railway IoT covert communication system, the transmission rate of the last time slot, and the prior value of the covert communication constraint are taken as the state set S;

[0100] The transmit beamforming matrix at the base station BS, the receive beamforming matrix at the mobile relay MR, and the RIS phase shift are used as the action space A;

[0101] The reward function Re is:

[0102]

[0103] Where M is the receive beamforming matrix at the mobile relay MR, H RM The conjugate transpose of , Θ is the reflection coefficient matrix of RIS, H BR is the channel matrix of the BS-RIS link, F H is the conjugate transpose of F, F is the transmit beamforming matrix at the base station BS, ω1 and ω2 are the weights of part 2 and part 3, respectively, ρ1 and ρ2 represent the satisfaction of the concealment constraint and the BS maximum transmit power constraint in each time slot t', respectively. Part 1 represents the direct utility, that is, the effective received signal, and parts 2 and 3 represent the cost function of each time slot t', corresponding to the cases where the concealment constraint requirements and the base station BS maximum transmit power requirements are not met, respectively.

[0104] Furthermore, the Actor network and Critic network of the deep reinforcement learning framework are composed of two DNN networks.

[0105] Furthermore, the DACM-based deep reinforcement learning framework jointly optimizes the transmit-receive beamforming and the smart metasurface RIS phase shift, including:

[0106] S1. The central controller of the base station obtains status information s t’ , the status information s t’ Input the Actor network to perform action selection and obtain action a t’ 、Reward t’ and the next state s t’+1 ;

[0107] S2. Sequence (s t’ , a t’ ,Re t’ , s t’+1 ) is stored in the experience replay pool as a data set for training the network. According to the Banach fixed point theorem and Bellman equation, the Critic network performs behavior evaluation based on the state value function and action value function;

[0108] S3. Randomly sample some data from the experience replay pool to calculate the loss function and train the network parameters of the DNN network. Return to S1 and repeat the training until the network converges.

[0109] Among them, the loss function is:

[0110]

[0111] in, is the loss function, τ is the size of the Mini-batch of the sampled dataset, V * is the status value, is the network parameter of the DNN network, Q * is the action value, a v is the vth data set action, s v is the state of the vth data set.

[0112] The following combination Figures 1-6 The method of this embodiment is further described:

[0113] like Figure 3 As shown, the present invention provides a railway Internet of Things covert communication method based on deep reinforcement learning, comprising the following steps:

[0114] Incorporating intelligent metasurfaces into the railway IoT communication system to build a high-speed railway IoT covert communication model based on RIS-MIMO;

[0115] Based on the RIS-MIMO-based high-speed rail IoT covert communication model, the channel model of the railway IoT covert communication system is constructed by considering the Doppler frequency shift compensation caused by high-speed train movement and the channel uncertainty introduced by artificial noise.

[0116] Based on the constructed railway IoT covert communication system channel model, a received signal model based on hypothesis testing is constructed;

[0117] Based on the received signal model based on hypothesis testing, the concealment constraints are derived;

[0118] Based on the security requirements and real-time performance of each high-speed rail sub-service, the priorities of high-speed rail communication services are defined and quantified with corresponding numbers. The design of hidden constraints will be linked to the service priority.

[0119] Based on the concealment constraints and service priorities, considering the covert communication with limited block length in high-speed rail, a joint optimization problem of beamforming and RIS phase shift at the transceiver and transmitter is constructed with the goal of maximizing the covert throughput.

[0120] Based on the joint optimization problem, considering the time-varying nature of the high-speed rail channel, the optimization problem is transformed into a Markov Decision Process (MDP) problem.

[0121] Based on the transformed MDP problem, a low-complexity and intelligent deep reinforcement learning framework is proposed to deal with the non-convexity and coupling of the problem and obtain the optimal solution to the problem.

[0122] In this embodiment, if Figure 1 As shown in the figure, adding intelligent metasurface to the railway IoT communication system and constructing a high-speed railway IoT covert communication model based on RIS-MIMO include:

[0123] Deploy the passive intelligent metasurface RIS between adjacent base stations. Assume that there are multiple illegal eavesdroppers with different locations and numbers, as well as trackside buildings with different numbers and heights. B A large uniform planar array (UPA) with N antennas W In the presence of an illegal eavesdropper named Willie, a root antenna secretly sends communication signals to a mobile relay (MR), which is equipped with N R The RIS is deployed between the BS and the MR, and is composed of N RA reflection unit is composed, which is to reduce the network cost, assist the BS and the train to carry out covert communication, evade the obstruction of obstacles to legal communication, and relieve the signal deviation caused by Doppler effect. According to the above analysis, if the interference signal existing in the environment is not considered, the received signal at the MR is:

[0124]

[0125] The reflection coefficient matrix of the RIS is defined as And Wherein, βm , n, θ m,n Correspond to the reflection coefficient and phase shift of the mth row and nth column unit on the surface of the RIS, e j is a complex exponential function, The phase shift of the N R th RIS unit is represented by q i The phase shift of the i th RIS unit is represented by q m,n ∈[0,2π), βm , n∈[0,1], Z M is the additive white Gaussian noise at the mobile relay MR.

[0126] In this embodiment, according to the high-speed rail Internet of Things covert communication model based on RIS-MIMO, the Doppler frequency shift compensation caused by high-speed train movement and the channel uncertainty introduced by artificial noise are considered, and the channel model of the railway Internet of Things covert communication system is constructed, including:

[0127] In the high-speed rail Internet of Things covert communication model based on RIS-MIMO, the positions of the BS, RIS and eavesdropper Willie are fixed, and the MR moves at high speed with the train. Therefore, H BR , the channel matrix W BW of the BS-Willie link, and the channel matrix W RW of the RIS-Willie link can be modeled as quasi-static channels without Doppler frequency shift, and G BM , H RM and the channel matrix W MW of the MR-Willie link can be modeled as time-varying channels with Doppler frequency shift. The quasi-static channels H BR , W BW , W RW can be represented as:

[0128]

[0129] In the formula, N B , N R and N WThey represent the number of BS antennas, the number of RIS units and the number of Willie antennas respectively, and L BR , L BW and L RW Represents channel H BR 、W BW and W RW The total number of signal paths, ρ is the channel H BR The ρth signal path, ω is the channel W BW The ωth signal path, χ is the channel W RW The χth signal path, α ρ For channel H BR The complex channel gain, α ω is the channel W BW The complex channel gain, α χ is the channel W RW The complex channel gain, For channel H BR The normalized receive array response vector for the ρth signal path, For channel H BR The normalized transmit array response vector of the ρth signal path, is the channel W BW The normalized receive array response vector for the ωth signal path, is the channel W BW The normalized transmit array response vector of the ωth signal path, is the channel W RW The normalized receiver array response vector of the xth signal path, is the channel W RW The normalized transmit array response vector of the xth signal path. c H Represents vector a c The conjugate transpose of . and Represents channel H BR The signal arrival azimuth and elevation angle of the ρth path is: and Represents channel H BR The signal of the ρth path leaves the azimuth and elevation angles, and Represents channel W BW The signal arrival azimuth and elevation angles of the ωth path are: and Represents channel W BW The signal of the ωth path leaves the azimuth and elevation angles, and Represents channel W RWThe signal arrival azimuth and elevation angle of the xth path, and Represents channel W RW The signal of the xth path leaves the azimuth and elevation angles.

[0130] In high-speed mobile communication networks, the high time-varying channel characteristics caused by the high mobility of trains make it difficult for base stations to obtain perfect CSI. If outdated CSI is used to design the transmit and receive beamforming and RIS phase shift, it will lead to very significant performance loss. Therefore, this embodiment uses the Clarke-Jakes spectrum, which is more suitable for the time-varying characteristics of the channel, to represent the Doppler shift. Specifically, for any time-varying channel H[t] at any time, let and H[t+Ts] represent outdated CSI and real-time CSI respectively, and the relationship between them is:

[0131]

[0132] Where, κ=J0(2πf D T S ) represents the time correlation coefficient, f D is the Doppler shift, J0( · ) is the first kind zero-order Bessel function, ΔH is the error term, t is any time during the train operation, and Ts is the transmission delay.

[0133] Therefore, the time-varying channel G BM 、H RM and W MW This method can be used to model the system and avoid the influence of channel time variation caused by Doppler frequency shift on the system. BM 、H RM and W MW They are modeled as:

[0134]

[0135] Where, and G BM [t+Ts] represents the outdated CSI and real-time CSI between BS and MR, respectively. and H RM [t+Ts] represents the outdated CSI and real-time CSI between RIS-MR, and W MW [t+Ts] represents the outdated CSI and real-time CSI between MR and Willie, t is any time during train operation, and Ts is the transmission delay.

[0136] Assume that MR generates interference signals of different powers to confuse Willie’s detection of covert communication. MR’s self-interference channel fM CN(0,φ), where φ∈[0,1] is the self-interference cancellation coefficient, which is determined by the efficiency of the self-interference cancellation (SIC) performed at the MR. In this embodiment, it is assumed that the interference signal transmission power P of the MR Mi Obey the interval [0, P Mi,max ], the probability density function is:

[0137]

[0138] Where, P Mi,max is the maximum self-interference transmission power, x is [0, P Mi,max ] on different values.

[0139] In this embodiment, based on the constructed railway IoT covert communication system channel model, constructing a received signal model based on hypothesis testing includes:

[0140] This embodiment considers Willie's binary detection model. In which, the null hypothesis H0 indicates that no covert communication occurs, the BS sends an artificial noise signal in the form of chaotic noise, and the MR sends an artificial noise signal in the form of [0, P Mi,max ]; Assume that H1 indicates that there is a covert communication behavior, BS sends a covert signal in the form of chaotic noise, and MR sends a self-interference signal that obeys [0, P Mi,max ] is a self-interference signal uniformly distributed on the surface. Based on the above analysis, the MR received signal under the H0-H1 assumption can be obtained:

[0141]

[0142] The Willie received signal under the H0-H1 assumption can be expressed as:

[0143]

[0144] Where Y M is the received signal of the mobile relay MR, G BM is the channel matrix of the BS-MR link, H RM The conjugate transpose of Z M is the additive white Gaussian noise at the mobile relay MR, Y W is the receiving signal model of the eavesdropper Willie, P Mi is the interference signal transmission power of the mobile relay MR, W MW is the channel matrix of the MR-Willie link, W RW The conjugate transpose of , Θ is the reflection coefficient matrix of RIS, H RMis the channel matrix of the RIS-MR link, f M is the self-interference channel of the mobile relay MR, s M is the self-interference signal from the mobile relay MR, Z W is the additive white Gaussian noise at the eavesdropper Willie, H0 indicates that no covert communication occurs, M is the receiving beamforming matrix at the mobile relay MR, W BW The conjugate transpose of H BR is the channel matrix of the BS-RIS link, F H is the conjugate transpose of F, F is the transmit beamforming matrix at the base station BS, s B is a covert signal from the BS, and H1 indicates that covert communication occurs.

[0145] In this embodiment, the concealment constraint is derived based on a received signal model based on hypothesis testing;

[0146] In the covert communication process of this embodiment, the total detection error probability at the eavesdropper Willie is expressed as:

[0147] ξ=π0Pr{D1H0}+π1Pr{D0H1};

[0148] In the formula, D1 and D0 are binary decisions to infer whether the BS transmission has occurred. D0 indicates that the base station BS has not communicated, and D1 indicates that the base station BS is communicating. π0 and π1 = 1-π0 represent the prior probabilities of the two hypotheses H0 and H1, respectively. H0 indicates that Willie judges that the BS has not communicated, and H1 indicates that Willie judges that the BS is communicating. Pr is the probability solution symbol. It should be emphasized that pre-setting the prior probability is beneficial to improving the detection performance of the eavesdropper Willie. This embodiment assumes π0 = π1 = 0.5 (i.e., equal prior probability). From the perspective of the base station, the larger ξ, the better. In this way, even if Willie uses the best detector, he cannot detect the BS transmission behavior. Generally, the concealment constraint of the legitimate communication party can be modeled as ξ ≥ 1-ε, where 0 ≤ ε ≤ 1 is the specified concealed communication constraint prior value to ensure that the transmitted information is sufficiently concealed.

[0149] This embodiment introduces KL divergence and total variation distance to provide a stricter hidden constraint. Specifically, let:

[0150] ξ=1-T V (L x ,L y );

[0151] Where, T V (L x , L y ) represents the variable Lx With L y The total variation distance between (x,y) = (0,1) or (1,0). Usually, T V (L x , L y ) is not easy to solve. To solve this problem, we can use Pinsker inequality to get:

[0152]

[0153] Where, D(L x , L y ) represents the value of the variable L x to L y KL divergence of .

[0154] From the analysis of the above embodiments, it can be seen that D(L x ||L y )≤2ε 2 is a stricter constraint than ξ≥1-ε. Therefore, in order to better achieve reliable covert communication under a given ε, this embodiment adopts D(L x ||L y )≤2ε 2 Represents an implicit constraint.

[0155] In this embodiment, based on the security requirements and real-time performance of each high-speed rail sub-service, the priority of high-speed rail communication services is defined and quantified with corresponding numbers. The design of hidden constraints will be associated with the level of service priority and include:

[0156] The priority of high-speed rail communication services is defined as follows: services with high security requirements and immediate response (generally in milliseconds or even shorter) are defined as high priority, with a quantitative identification of 3; services with a certain time tolerance from service generation to transmission (the response time is significantly longer than high-priority services) and no strong security requirements are defined as medium priority, with a quantitative identification of 2; customers' ordinary short data communications / file transfers and low-security requirements services that do not require response measures are defined as low priority, with a quantitative identification of 1. Among them, high-priority services generally refer to railway emergency calls, emergency voice / video communications, driving-related dispatching communications and other services. Medium-priority services generally refer to non-driving-related dispatching communications, operation and maintenance voice / video communications and other services. Low-priority services generally refer to ordinary data communications / file transfers and other services for users on board. In this definition, the priority of medium-priority and low-priority services can be increased as appropriate.

[0157] Based on the above priority definition, when constructing the optimization problem and designing the hidden constraint value, D(L x ||L y )≤2ε 2The values ​​of (x, y) and ε are related to the service priority. Specifically, higher priority services should use (x, y) = (1, 0) and a smaller ε value to ensure the reliability of covert transmission.

[0158] In this embodiment, based on the concealment constraints and service priorities, and considering the limited block length of high-speed rail covert communication, with the goal of maximizing covert throughput, a joint optimization problem of beamforming and RIS phase shifting at the transceiver end is constructed, including:

[0159] Considering short packet communication with finite block length, the concealed throughput at MR can be expressed as:

[0160] η=KR(1-δ);

[0161] Where δ is the decoding error probability at the MR, K is the packet length, and R represents the information transmission rate from the BS to the MR:

[0162]

[0163] Where Q -1 (·) represents the inverse Q function, γ M represents the signal to interference plus noise ratio (SINR) at MR and can be written as:

[0164]

[0165] Where M is the receiving beamforming matrix at the mobile relay MR, G BM is the channel matrix of the BS-MR link, H RM The conjugate transpose of H RM is the channel matrix of the RIS-MR link, Θ is the reflection coefficient matrix of RIS, H BR is the channel matrix of the BS-RIS link, F H is the conjugate transpose of F, F is the transmit beamforming matrix at the base station BS, P Mi is the interference signal transmission power of MR, φ is the self-interference cancellation coefficient, σ M is the additive white Gaussian noise power at the mobile relay.

[0166] When the packet length K and the decoding error probability δ at MR are determined, the throughput maximization problem of high-speed rail IoT covert communication based on RIS-MIMO can be modeled as:

[0167]

[0168] Where N Rv ×N Rh=N R is the number of RIS units, N Rv is the total number of rows of RIS units, N Rh is the total number of columns of RIS units, m is the number of rows of RIS units, and n is the number of columns of RIS units. C1 is the covert constraint that satisfies the covert communication conditions, C2 is the BS maximum transmit power constraint, C3 is the unit modulus constraint of RIS phase shift, ε is the prior value of the covert constraint, and P B is the maximum transmit power of the base station.

[0169] In this embodiment, based on the joint optimization problem, the optimization problem is transformed into a Markov Decision Process (MDP) problem taking into account the time-varying nature of the high-speed rail channel.

[0170] For the optimization problem constructed, this embodiment first converts it into a Markov Decision Process (MDP) problem. The MDP problem contains a large and complex state space and action space. Usually, an MDP consists of a five-tuple (S, A, P sa , Re, γ), where S represents the state set, A is the action space, P sa is the state transition probability, Re is the reward function, and γ is the discount factor. In each cycle, the agent observes the system state s t’ ∈S, select action a t’ ∈A, and then transfer to the new state s t’+1 , determine the benefit type by calculating the discount factor. Finally, obtain the timely reward Re(s) from the environmental feedback t’ ,a t’ The key elements of MDP are defined as follows:

[0171] State space: The state of a time slot t' includes the channel information of all nodes in the system, the transmission rate of the last time slot, and the prior value of the covert communication constraint, that is:

[0172]

[0173] Action space: Based on the observed system state s t’ , the agent takes the transmit beamforming matrix at BS, the receive beamforming matrix at MR and the RIS phase shift as the output of the Actor network, namely:

[0174] a t' =[M (t') ,F (t') ,Θ (t') ];

[0175] State transition probability: P sa (s t’+1 |s t’ ,a t’ ) means to execute action a in the current state t’ Then transfer to the new state s t’+1 The probability, P sa Usually determined by the object model.

[0176] Reward function: When the agent performs an action in the current state, the reward serves as a signal to evaluate the quality of the covert beamforming strategy. System performance can only be improved when the reward function of each learning step is related to the desired goal. Therefore, the goal of deep reinforcement learning is to maximize the signal-to-interference-to-noise ratio at the MR while satisfying the concealment constraint, the BS maximum transmit power constraint, and the unit modulus constraint of the RIS phase shift. The reward function is expressed as:

[0177]

[0178] in,

[0179]

[0180] Where part 1 represents the direct utility, or effective signal reception. Parts 2 and 3 represent the cost function at a given time t, corresponding to failure to meet the concealment constraint and the BS maximum transmit power requirement, respectively. The coefficients ω1 and ω2 are the weights of parts 2 and 3, respectively, balancing utility and cost. ρ1 and ρ2 represent the satisfaction level imposed on the concealment constraint and the BS maximum transmit power constraint, respectively, at each time slot t'. If the relevant constraints are met at the current time slot, ρ1 = 0 and ρ2 = 0, indicating no penalty in the reward function.

[0181] In this embodiment, based on the transformed MDP problem, a low-complexity and intelligent deep reinforcement learning framework is proposed to handle the non-convexity and coupling of the problem and obtain the optimal solution to the problem, including:

[0182] After converting the optimization problem into an MDP problem, the state space of the MDP will be composed of a large amount of RIS-assisted high-speed rail environment information, and the variables in the action space will be continuous. Therefore, a deep reinforcement learning framework based on the Double Actor-Critic Methods (DACM) is proposed to jointly optimize the transmit and receive end beamforming and RIS phase shift. The specific framework is shown in Figure 2The Actor network and the Critic network are both composed of two DNNs, which are the training network and the target network, respectively. The experience replay pool is responsible for recycling reward information and state information and extracting samples for the learning and training of the reinforcement network.

[0183] The DACM-based deep reinforcement learning framework for solving the maximum hidden throughput mainly includes the following four steps:

[0184] Step 1: The central controller of the BS is responsible for sensing and collecting state information s in the environment t’ and inputting it into the Actor network to perform action selection. The Actor network outputs the corresponding action a t’ , the environment obtains the immediate reward Re t’ , and then moves to the next state s t’+1 .

[0185] Step 2: Store the sequence (s t’ , a t’ , Re t’ , s t’+1 ) in the experience replay pool as the dataset for the training network. The Critic network evaluates the behavior through the state value function and the action value function. According to the Banach fixed point theorem and the Bellman equation, the optimal state value function V*(s t’ ) and the optimal action value function Q*(s t’ , a t’ ) under the optimal policy π* satisfy the following equations:

[0186] V * (s t' ) = E π [Re(s t'+1 |s, a) + γV * (s t'+1 )|s0 = s];

[0187] Q * (s t' , a t' ) = E π [Re(s t'+1 |s, a) + γQ * (s t'+1 , a t'+1 )|s0 = s, a0 = a];

[0188] where s0 is the initial state, a0 is the initial action, and E π represents the expectation of the variable under the policy π.

[0189] Step 3: In order to avoid the correlation of state information, some data is randomly sampled from the experience replay pool to calculate the loss function and train the network parameters of the DNN network. The loss function is given by the following formula:

[0190]

[0191] In the above formula, is the loss function, τ is the size of the Mini-batch of the sampled dataset, V * is the state value, s v is the state of the vth data set, Q * is the action value, a v is the vth data set action, is the network parameter of the DNN network, which is given by the following formula:

[0192]

[0193] Where, β π is the learning rate, is the first-order partial derivative, is the loss function calculated at a certain time slot t'.

[0194] Strategy parameter π t’ The update rule is given by:

[0195] π t'+1 =argmaxQ(s t' ,a t' );

[0196] In the formula, argmax represents the maximum value of the action value function Q (s t’ ,a t’ ).

[0197] Step 4: The agent continuously obtains state information from the environment and repeatedly trains until the network converges, completing the training process and obtaining the optimal solution for the transmit and receive beamforming matrix and RIS phase shift.

[0198] In this embodiment, the execution steps for solving the concealed throughput maximization are shown in Table 1.

[0199] Table 1

[0200]

[0201] The specific conditions of this example are: system center frequency is 28 GHz; bandwidth is 100 MHz; noise power is -134 dBm; path loss model: PL (d x )=PL(d0)+10ulog(d x / d0)+X σ, where PL(d0) represents the intercept, u ​​is the path loss index, d0 is the reference distance, and d x Indicates the distance between the transmitter and receiver related to the train position, X σ is shadow fading and obeys a Gaussian random distribution with a mean of 0.

[0202] Due to the asymmetry of relative entropy, the results obtained by using D(L0||L1) and D(L1||L0) as hidden constraints are generally different. Figure 4 The figure shows a comparison of a set of experimental results conducted by the present invention: under two KL divergences with the same initial conditions, the variation of the system's concealed throughput with the number of BS antennas is observed. Figure 4 As can be seen from the figure: As the number of BS antennas N B With the increase of , the throughput under both concealment constraints increases, but D(L0||L1)≤2ε 2 The concealed throughput under the constraint is always greater than D(L1||L0)≤2ε 2 The hidden throughput under the constraint, which means D(L1||L0)≤2ε 2 is the ratio D(L0||L1)≤2ε 2 More stringent concealment constraints. Therefore, when designing concealment constraints, the present invention considers adopting D(L1||L0)≤2ε 2 As the hidden constraint term of the concealed throughput maximization problem, the hidden constraint prior value ε is designed according to the priority of the service to achieve better system performance.

[0203] Figure 5 The figure shows the changing trend of system rewards with the number of training sets under different learning rates. It can be seen that different learning rates have different effects on the performance of deep reinforcement learning algorithms. Specifically, when β π = 0.1, the learning rate is too large, the behavior is oscillatory, and its reward performance is much lower than β π = 0.001. In addition, if the learning rate is set too small, the convergence time will be longer. As can be seen from the figure, the model built in this embodiment has a learning rate of β π = 0.001, it can learn the problem well and converge quickly. In addition, when setting β π =0.001, Figure 5 It can also be seen that when the number of training sets is about 100, the algorithm has converged quickly and remains stable as the number of training sets increases.

[0204] Figure 6 Shown is the relationship between the loss value of the neural network during training / validation and the number of training sets. Figure 6As can be seen from the graph, the training / loss value drops rapidly before the number of training sets reaches 100, and stabilizes after the number of training sets reaches 200. In addition, the validation loss is slightly higher than the training loss, which indicates that the DNN network parameters designed in the present invention have a good fitting ability in mapping between input samples and output samples.

[0205] When solving the concealed throughput maximization problem proposed by the DACM framework in this paper, the time complexity of the training phase is where N D , T is the number of cycles shown in Table 1, τ is the size of Mini-batch, La, Lc are the number of training layers of Actor network and Critic network respectively, p e Indicates the number of neurons in the corresponding DNN layer e, but the algorithm can be run offline during the training phase with high complexity. In addition, once the network performance finally converges, the complexity of each step can be reduced to Comparison shows that the complexity of the algorithm proposed in this invention is much lower than that of the existing technology. Furthermore, the algorithm has the ability to continuously obtain information from the environment and optimize the network training model, which is very suitable for fast time-varying channel scenarios.

[0206] Example 2:

[0207] A railway Internet of Things covert communication system based on deep reinforcement learning, including:

[0208] A communication model building module is used to incorporate intelligent metasurface RIS into the railway IoT communication system and build a high-speed railway IoT covert communication model based on RIS-MIMO;

[0209] The channel model construction module is used to build a channel model for the railway IoT covert communication system based on the RIS-MIMO high-speed railway IoT covert communication model, taking into account the Doppler frequency shift compensation caused by high-speed train movement and the channel uncertainty introduced by artificial noise;

[0210] A receiving signal model building module is used to build a receiving signal model based on hypothesis testing according to the communication model of the railway Internet of Things covert communication system;

[0211] The constraint condition design module is used to obtain hidden constraint conditions based on the received signal model and railway communication service priorities;

[0212] An optimization problem construction module is used to construct a joint optimization problem of the transmit-receive beamforming and the smart metasurface RIS phase shift, combining concealment constraints with the goal of maximizing the concealment throughput.

[0213] The deep reinforcement learning algorithm design module is used for processing the non-convexity and coupling of the problem by using a deep reinforcement learning framework based on the converted Markov decision process problem.

[0214] The optimal solution calculation module is used for adjusting model parameters of the deep reinforcement learning framework, training a network model, and obtaining an optimal solution of the problem until the model converges.

[0215] The above-described embodiments are only used to describe the preferred modes of the present application, and are not used to limit the scope of the present application. Without departing from the design spirit of the present application, various modifications and improvements of the technical solutions of the present application made by those skilled in the art shall fall within the protection scope of the present application defined by the claims.

Claims

1. A railway Internet of Things covert communication method based on deep reinforcement learning, characterized in that: include: Incorporating intelligent metasurface RIS into the railway IoT communication system, and constructing a high-speed railway IoT covert communication model based on RIS-MIMO. Based on the RIS-MIMO-based high-speed rail IoT covert communication model, a channel model for the railway IoT covert communication system is constructed by considering the Doppler frequency shift compensation caused by the high-speed movement of the train and the channel uncertainty introduced by artificial noise. According to the communication model of the railway Internet of Things covert communication system, a received signal model based on hypothesis testing is constructed; Obtaining a concealment constraint condition according to the received signal model and the railway communication service priority; Combined with the aforementioned concealment constraints, a joint optimization problem of the transmit-receive beamforming and the smart metasurface RIS phase shift is constructed with the goal of maximizing the concealment throughput. According to the time-varying nature of the high-speed rail channel, the joint optimization problem is transformed into a Markov decision process problem; Based on the transformed Markov decision process problem, a deep reinforcement learning framework is used to deal with the non-convexity and coupling of the problem, obtain the optimal solution to the problem, and ensure reliable and low-cost covert communication in fast time-varying channel scenarios.

2. The railway Internet of Things covert communication method based on deep reinforcement learning according to claim 1 is characterized in that: The RIS-MIMO-based high-speed rail IoT covert communication model includes: A train moves from the center of the base station BS coverage area to the edge of the coverage area, a trackside base station equipped with multiple antennas and a large uniform planar array UPA, the smart metasurface RIS deployed between the base station BS and the train, the eavesdropper Willie and the trackside buildings. It is assumed that a mobile relay MR is installed on the top of the train for communicating with the base station BS.

3. The railway Internet of Things covert communication method based on deep reinforcement learning according to claim 2 is characterized in that: Considering the Doppler shift compensation caused by high-speed train movement and the channel uncertainty introduced by artificial noise, the channel model for the railway IoT covert communication system includes: According to the RIS-MIMO-based high-speed rail IoT covert communication model, the positions of the base station BS, the intelligent metasurface RIS, and the eavesdropper Willie are fixed, and the mobile relay MR moves with the train. The channel matrix H of the BS-RIS link is BR , the channel matrix W of the BS-Willie link BW And the channel matrix W of the RIS-Willie link RW It is constructed as a quasi-static channel without Doppler shift, and the channel matrix G of the BS-MR link is BM , the channel matrix H of the RIS-MR link RM And the channel matrix W of the MR-Willie link MW It is constructed as a time-varying channel with Doppler shift; The construction of the time-varying channel is to represent the Doppler frequency shift by the Clarke-Jakes spectrum. For any time-varying channel H[t] at any time, let and H[t+Ts] represent the outdated channel state information CSI and the real-time channel state information CSI, respectively. The relationship between H[t+Ts] is: Where, κ=J0(2πf D T S ) represents the time correlation coefficient, f D is the Doppler frequency shift, J0(·) is the first kind of zero-order Bessel function, ΔH is the error term, t is any time during the train operation, and Ts is the transmission delay; The mobile relay MR generates interference signals of different powers, and the self-interference channel of the mobile relay MR is f M ~CN(0,φ), φ∈[0,1] is the self-interference cancellation coefficient, CN is the symbolic representation of the complex Gaussian distribution, assuming that the interference signal transmission power P of the mobile relay MR is Mi Obey the interval [0, P Mi,max ], the probability density function is: Where, P Mi,max is the maximum self-interference transmission power, x is [0, P Mi,max ] on different values.

4. The railway Internet of Things covert communication method based on deep reinforcement learning according to claim 3 is characterized in that: Constructing a received signal model based on hypothesis testing includes: constructing a received signal model at the eavesdropper Willie; The received signal model at the eavesdropper Willie is: Among them, Y W is the receiving signal model of the eavesdropper Willie, W RW The conjugate transpose of , Θ is the reflection coefficient matrix of RIS, s M is the self-interference signal from the mobile relay MR, Z W is the additive white Gaussian noise at the eavesdropper Willie, H0 indicates that no covert communication occurs, M is the receiving beamforming matrix at the mobile relay MR, W BW The conjugate transpose of F H is the conjugate transpose of F, F is the transmit beamforming matrix at the base station BS, s B is a covert signal from the BS, and H1 indicates that covert communication occurs.

5. The railway Internet of Things covert communication method based on deep reinforcement learning according to claim 4 is characterized in that: Obtaining concealment constraints based on the received signal model and railway communication service priority includes: Obtaining a total detection error probability at the eavesdropper Willie according to a received signal model at the eavesdropper Willie; According to the total detection error probability of the eavesdropper Willie, KL divergence and total variation distance are introduced, and Pinsker inequality is used to express the hidden constraint as follows: Where, T V (L x , L y ) represents the variable L x With L y The total variation distance between x ||L y ) represents the value of the variable L x to L y KL divergence of (x,y) = (0,1) or (1,0); Let 0≤ε≤1, and D(L x ||L y )≤2ε 2 Replace ξ≥1-ε as the concealment constraint, where ε is the prior value of the concealment constraint and ξ is the total detection error probability at the eavesdropper Willie; Define the priority of high-speed rail communication services and use digital quantification to identify them. x ||L y )≤2ε 2 The values ​​of (x, y) and ε are related to the service priority.

6. The railway Internet of Things covert communication method based on deep reinforcement learning according to claim 3 is characterized in that: In combination with the aforementioned concealment constraints, with the goal of maximizing the concealment throughput, the joint optimization problem of the transmit-receive beamforming and the smart metasurface RIS phase shift is constructed, which includes: Obtaining a concealed throughput at the mobile relay MR based on short packet communication with a finite block length; Taking the maximization of concealment throughput as the goal and combining the concealment constraint condition to construct the joint optimization problem; The joint optimization problem is: Where F is the transmit beamforming matrix at the base station BS, M is the receive beamforming matrix at the mobile relay MR, Θ is the reflection coefficient matrix of RIS, η is the concealed throughput at the mobile relay MR, D(L x ‖L y ) represents the value of the variable L x to L y KL divergence, ε is the prior value of the hidden constraint, P B is the maximum transmission power of the base station, N Rv is the total number of rows of RIS units, N Rh is the total number of columns of RIS units, m is the number of rows of RIS units, n is the number of columns of RIS units, C1, C2, and C3 are the covert constraints that satisfy the covert communication conditions, the maximum transmission power constraint of the base station BS, and the unit modulus constraint of the RIS phase shift, respectively. δ is the decoding error probability at MR, K is the packet length, and R is the information transmission rate from the base station BS to the mobile relay MR.

7. The railway Internet of Things covert communication method based on deep reinforcement learning according to claim 6 is characterized in that: Converting the joint optimization problem into a Markov decision process problem involves: The Markov decision process problem consists of a state set S, an action space A, and a state transition probability P. sa , reward function Re and γ discount factor; The channel information of all nodes in the railway Internet of Things covert communication system, the transmission rate of the last time slot and the prior value of the covert constraint are used as the state set S; Taking the transmit beamforming matrix at the base station BS, the receive beamforming matrix at the mobile relay MR and the RIS phase shift as an action space A; The reward function Re is: Where Re is the reward function, M is the receive beamforming matrix at the mobile relay MR, H RM The conjugate transpose of , Θ is the reflection coefficient matrix of RIS, H BR is the channel matrix of the BS-RIS link, F H is the conjugate transpose of F, F is the transmit beamforming matrix at the base station BS, ω1 and ω2 are the weights of part2 and part3, respectively, ρ1 and ρ2 represent the satisfaction of the concealment constraint and the BS maximum transmit power constraint in each time slot t', respectively. Part 1 represents the direct utility, that is, the effective received signal, and parts 2 and 3 represent the cost function of each time slot t', corresponding to the cases where the concealment constraint requirements and the base station BS maximum transmit power requirements are not met, respectively.

8. The railway Internet of Things covert communication method based on deep reinforcement learning according to claim 7 is characterized in that: The deep reinforcement learning framework is used to deal with the non-convexity and coupling of the problem. The optimal solution to the problem is obtained by: A deep reinforcement learning framework based on the dual actor-critic method DACM jointly optimizes the transmit-receive beamforming and the intelligent metasurface RIS phase shift, wherein the Actor network and the Critic network of the deep reinforcement learning framework are composed of two DNN networks.

9. The railway Internet of Things covert communication method based on deep reinforcement learning according to claim 8 is characterized in that: A deep reinforcement learning framework based on the dual-actor-critic method (DACM) to jointly optimize the transmit-receive beamforming and the smart metasurface RIS phase shift includes: S1. The central controller of the base station BS obtains status information s t’ , the state information s t’ Input the Actor network to perform action selection and obtain action a t’ 、Reward t’ and the next state s t’+1 ; S2. Sequence (s t’ , a t’ ,Re t’ , s t’+1 ) is stored in the experience replay pool as a data set for training the network. According to the Banach fixed point theorem and the Bellman equation, the Critic network performs behavior evaluation based on the state value function and the action value function; S3. Randomly sample some data from the experience replay pool to calculate the loss function and train the network parameters of the DNN network, return to S1, and repeatedly train until the network converges; Wherein, the loss function is: in, is the loss function, τ is the size of the Mini-batch of the sampled dataset, V * is the state value, s v is the state of the vth data set, is the network parameter of the DNN network, Q * is the action value, a v is the vth data set action.

10. A railway Internet of Things covert communication system based on deep reinforcement learning, used to implement the railway Internet of Things covert communication method based on deep reinforcement learning according to any one of claims 1 to 9, characterized in that: include: Communication model construction module, channel model construction module, received signal model construction module, constraint condition design module, optimization problem construction module, deep reinforcement learning algorithm design module and optimal solution calculation module; The communication model construction module is used to add the intelligent metasurface RIS to the railway Internet of Things communication system to build a high-speed railway Internet of Things covert communication model based on RIS-MIMO; The channel model construction module is used to construct a channel model for the railway Internet of Things covert communication system based on the RIS-MIMO-based high-speed railway Internet of Things covert communication model, taking into account the Doppler frequency shift compensation caused by the high-speed movement of the train and the channel uncertainty introduced by artificial noise; The received signal model construction module is used to construct a received signal model based on hypothesis testing according to the communication model of the railway Internet of Things covert communication system; The constraint condition design module is used to obtain hidden constraint conditions according to the received signal model and the railway communication service priority; The optimization problem construction module is used to construct a joint optimization problem of the transmit-receive end beamforming and the smart metasurface RIS phase shift in combination with the concealment constraint conditions with the goal of maximizing the concealment throughput; According to the time-varying nature of the high-speed rail channel, the joint optimization problem is transformed into a Markov decision process problem; The deep reinforcement learning algorithm design module is used to handle the non-convexity and coupling of the problem using a deep reinforcement learning framework based on the transformed Markov decision process problem; The optimal solution calculation module is used to adjust the model parameters of the deep reinforcement learning framework and train the network model until the model converges to obtain the optimal solution to the problem.