Security control method for networked switching system based on encryption and decryption switching q-learning replay attack detection
By constructing a replay attack detection mechanism based on Q-learning during encryption/decryption switching, the security control problem of networked switching systems under replay attacks is solved, enabling accurate detection and analysis of replay attacks and ensuring the security and stability of the system under replay attacks.
Patent Information
- Application Number
- CN202410860014.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2044-06-28
AI Technical Summary
Existing networked switching systems struggle to accurately determine whether they are under attack when facing replay attacks, leading to inaccurate system status analysis and impacting security control.
A replay attack detection mechanism based on Q-learning during encryption/decryption switching is constructed. By building a switching system model and a controller model, designing an event triggering mechanism and a Q-learning algorithm, encryption and decryption processing is performed, the estimated values and residual norms of state and modal signals are calculated, and the security controller equations are established simultaneously to construct switching rules to judge and defend against replay attacks.
It enables accurate detection and analysis of replay attacks, reduces asynchronous duration, ensures the security and stability of the system under replay attacks, and improves the reliability of security control.
Smart Images

Figure CN118759931B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of security control technology for handover systems, and in particular to a security control method for networked handover systems based on encryption / decryption handover Q-learning replay attack detection. Background Technology
[0002] As modern control systems increasingly move towards networking, the complexity of industrial system models is also constantly increasing. To meet the demands of changing working environments and improve operational flexibility, many systems in practical applications are beginning to exhibit multimodal characteristics. Simultaneously, the need for remote spatial distribution and low-cost control is growing, requiring remote monitoring and control via networks. Faced with complex industrial production environments, networked switching systems and their safety control have become current research hotspots.
[0003] As a typical network attack, replay attacks can record legitimate data packets in a system and resend them at a later time. Because the replayed data is highly similar to the system data, replay attacks are more likely to successfully bypass attack detection mechanisms. Common designs for replay attack detection mechanisms include adding watermarking mechanisms, designing encryption / decryption functions, and reinforcement learning detection methods. However, these methods mainly focus on designing detection mechanisms to determine whether state signals have been attacked. In fact, the switching signals of a switching system can also be captured and replayed by replay attacks during network transmission. That is, the switching system is vulnerable to attacks where both system state signals and switching signals are replayed, causing complex and multiple consecutive asynchronous switching events. This switching behavior is not only difficult to analyze but also greatly increases the difficulty of system stability. Under such replay attacks, because the state data of each subsystem in the switching system is shared, the energy flow between inactive subsystems can also affect the security control of the switching system. Summary of the Invention
[0004] This invention provides a security control method for a networked handover system based on encryption / decryption switching Q-learning replay attack detection, to overcome the technical problems of existing network handover systems that cannot accurately determine whether they are under replay attack and cannot accurately analyze the system status, thus failing to perform security control.
[0005] To achieve the above objectives, the technical solution of the present invention is as follows:
[0006] A security control method for a networked handover system based on encryption / decryption handover Q-learning replay attack detection includes:
[0007] S1: Construct the switching system model and controller model to obtain state signals and modal signals;
[0008] S2: Construct an event triggering mechanism based on replay attack detection, the event triggering mechanism including error detection conditions, modality matching conditions and attack detection conditions; use the event triggering mechanism to process the state signal and modality signal to obtain the subsystem modality signal and state signal at the trigger time;
[0009] S3: Construct a replay attack detection mechanism based on switching Q-learning according to the Q-learning algorithm. The replay attack detection mechanism is used to encrypt and decrypt the subsystem state signal, and calculate the estimated values of the subsystem modal signal and state signal after encryption and decryption, as well as the state estimation residual norm. The event triggering mechanism and the replay attack detection mechanism are used to process the subsystem modal signal and state signal at the triggering time to obtain the state estimate, modal estimate and state estimation residual norm, and determine the detection result of the replay attack.
[0010] S4: Substitute the state estimate and the modality estimate into the controller model to obtain the security controller equation. Combine the security controller equation, the switching system model, and the equations of the state estimator and modality estimator of the replay attack detection mechanism to obtain the closed-loop system equation.
[0011] S5: Construct switching rules based on replay attack detection. The switching rules are used to analyze the switching behavior of the closed-loop system under replay attack based on the replay attack detection results, obtain the judgment conditions for system security and stability, and perform security control on the closed-loop system.
[0012] Furthermore, S1 constructs a switching system model and a controller model to obtain state signals and modal signals, including:
[0013] S11. Construct the switching system model, as shown in formula (1).
[0014]
[0015] in, To switch the system's state vector, control input, system output, and external disturbances; It is a switching signal, A σ B σ C σ D σ It is a given constant matrix;
[0016] S12. Construct the controller model, as shown in formula (2).
[0017] u(t) = K φ x(t) (2)
[0018] Where φ = φ(t), φ(t) ∈ L are the modal signals of the controller, and K φThis is the controller gain.
[0019] Furthermore, S2 constructs an event-triggered mechanism based on replay attack detection, which includes error detection conditions, modality matching conditions, and attack detection conditions. The event-triggered mechanism is used to process the state signal and modal signal to obtain the subsystem modal signal and state signal at the trigger time, including:
[0020] S21. Construct an event triggering mechanism based on replay attack detection, as shown in formula (3).
[0021]
[0022] Wherein, the initial trigger time t1h=l0, α(g r,m h) represents the time interval at which the data arrives at the controller at the (m-1)th sampling time after the r-th trigger. The attack detection results obtained above, These represent error detection conditions, modality matching conditions, and attack detection conditions, respectively.
[0023] S22. Construct error detection conditions, as shown in formula (4).
[0024]
[0025] in, Indicates sampling time g r,m The parameter matrix at h, where and Let ε be the weight matrix to be designed. r (t)=x(g r,m h)-x(t r h), g r,m h = t r h+mh represents the r-th trigger time t. r At the m-th sampling time after h, δ σ >0 is the trigger threshold; x(t) r h) represents the subsystem state signal at the trigger moment; ε r (t) represents the trigger time t. r The error between the state of h and its state at the m-th sampling time;
[0026] S23. Construct modal matching conditions, as shown in formula (5).
[0027]
[0028] Wherein, σ(g) r,m h) represents the r-th trigger time t r The modal signal at the m-th sampling time after h, then tr The modal signal at time h is σ(t) r h);
[0029] S24. Construct attack detection conditions, as shown in formula (6).
[0030]
[0031] Furthermore, the replay attack detection mechanism includes encryption measures, decryption measures, state estimation measures, modality estimation measures, and a Q-learning-based switching Q-learning method;
[0032] The encryption measure is used to encrypt the subsystem state signal at the trigger time to obtain the encrypted state signal, and to obtain the encrypted state signal and modal signal under replay attack based on the encrypted state signal, and to obtain the actual modal signal based on the encrypted modal signal.
[0033] The encryption measures are used to encrypt the subsystem state signal at the triggering time to obtain the encrypted state signal, and based on the encrypted state signal, to obtain the encrypted state signal and modal signal under replay attack, and based on the encrypted modal signal and the event triggering mechanism, to obtain the actual switching signal;
[0034] The decryption measures are used to decrypt the encrypted state signal under the replay attack, and to obtain the actual system state signal based on the decrypted system signal and the event triggering mechanism.
[0035] The state estimation measure is implemented by a state estimator, which is used to calculate the state estimate of the system state signal in the switching system under the actual switching signal.
[0036] The modal estimation measure is implemented by a modal estimator, which is used to calculate the modal estimate of the modal signal in the switching system under the actual switching signal.
[0037] The Q-learning-based switching Q-learning method is used to determine the detection result of a replay attack based on the difference between the actual system state signal and the state estimate, i.e., the state estimate residual norm, and the modality estimate.
[0038] Furthermore, the subsystem modal signals and state signals at the triggering moment are processed using the event triggering mechanism and the replay attack detection mechanism to obtain state estimates and modal estimates, including:
[0039] S31. The encryption measures are as shown in formula (7).
[0040] x c (t r h)=x(t rh)+c(t r h) (7)
[0041] Wherein, c(t) r h) is a time-varying encryption function generated from a random seed, x c (t r h) represents the encrypted state signal, x(t) r h) represents the trigger time t before encryption. r The system state at point h;
[0042] The encrypted state signal and switching signal are i.e., the modal signal σ(t) r h) under the influence of replay attacks, it can be rewritten as formula (8).
[0043]
[0044] Where α(t) represents the detection result of the replay attack;
[0045] S32. The decryption measures are as shown in formula (9).
[0046]
[0047] in, d(t) represents the decrypted system state signal. r h) is the decryption key generated from the random seed, and the energy limit of the time-varying encryption function and the decryption key satisfies the condition. Given a constant matrix;
[0048] Under the action of the event triggering mechanism based on replay attack detection and the encryption / decryption measures, the actual switching signal and actual state signal of the switching system are as shown in formulas (10) and (11).
[0049]
[0050] Where, η r (t)=tg r,m h represents the piecewise time delay function; ε r (t) represents the trigger time t. r The error between the state of h and its state at the m-th sampling time; ε p (t) represents the error between the state at time p and the state at the m-th sampling time thereafter;
[0051] S33, The state estimator is as shown in formula (12).
[0052]
[0053] in, L is the state estimate of the switching system under the actual switching signal of x(t). φ(t) The observer gain, φ = φ(t), φ(t) ∈ L, represents the modal estimate of the switching system under the actual switching signal; A φ(t) B φ(t) D φ(t) It is the system parameter matrix of the φ(t)th state estimator;
[0054] S34. The modal estimator is as shown in formula (13).
[0055]
[0056] in, This represents the estimated state partition corresponding to the i-th state estimator. yes The boundary, and
[0057] Furthermore, the Q-learning-based switching Q-learning method includes switching between the Q-table training phase and the attack detection phase:
[0058] S35, Switch to Q-table training phase:
[0059] Construct a quintuple (A, S, γ) f ,π f For a Markov decision process of f(k), train the Q-table and update the Q-value, where f∈F={0,1,…,F} is the number of training iterations;
[0060] Let φ be a set of environment variables. f (g r,m h) is the r,m-th learning interval Φ in the f-th training round. r,m The modal estimator provides estimates of the subsystem modes. It is the residual norm of the state estimation at step r,m in the f-th round of training. It is in the fth round of training
[0061] Action Set Indicates that the detector is based on environment variables. The detection alarm values are defined as follows: when a1 = 0, it indicates that no replay attack has been detected; when a2 = 1, it indicates that a replay attack has been detected.
[0062] reward function γ f (r,m) represents the environment variables at step r,m in the f-th training round. In performing the action The rewards obtained later, among which
[0063] Strategy This represents the environment variables at step r,m in the f-th training round. Execute action The probability of the discount parameter k∈[0,1] represents the degree of attention to long-term rewards;
[0064] S351. Initialize the actions for the f-th round of training. Environment variables at steps r and m According to strategy Select Action As shown in formula (14),
[0065]
[0066] The environmental variables at step r,m By performing actions Transform to environment variables in step r,m+1 According to the reward function γ f (r,m) Calculate the action to be executed The reward obtained later is shown in formula (15).
[0067]
[0068] S352. Update the Q value according to the reward function and the value of the (f-1)th training round, as shown in formula (16).
[0069]
[0070] in, For the environment variable-action pair during the f-th training round The Q-value after the (f-1)th training round. Environment variables for the f-th round of training The Q-table value α corresponding to the optimal action after the (f-1)th round of training. f =0.1·(2+arctan(f-α) * )) is the learning step size, α * Let be the learning constant, where the initial value of the Q table is . r∈N + m∈N;
[0071] S353, Regarding the environment variables at step r,m+1 Repeat the process of S351-S352 to obtain the corresponding Q value until the system stops running, and the training of the fth round ends.
[0072] S354. Repeat steps S351-S353 to perform the next round of training until f = F. Training ends, and the system returns a list containing all environment variable-action pairs. Complete switching of the Q table;
[0073] S36. During the attack detection phase, the corresponding optimal action is selected from the complete switching Q table.
[0074] S361. Determine environmental variables based on the modal estimates and state estimation residual norms in the complete switching Q table:
[0075] Let r = 1, m = 0, when t ∈ Φ r,m At that time, observe the current environmental variables and the current modal estimates;
[0076] S362. Select the optimal action corresponding to the current environment variable from the complete switching Q table.
[0077] S363. Output the attack detection results, as shown in formula (17).
[0078] α(t)=a r,m ,t∈Φ r,m (17).
[0079] Furthermore, S4 substitutes the state estimate and the modal estimate into the controller model to obtain the security controller equation. Then, by simultaneously solving the security controller equation, the switching system model, and the equations of the state estimator and modal estimator of the replay attack detection mechanism, a closed-loop system equation is obtained, including:
[0080] Substituting the state estimate and the modal estimate into the controller model, a safety controller is obtained, as shown in formula (18).
[0081]
[0082] φ = φ(t), φ(t) ∈ L, representing the modal estimate obtained from the modal estimator;
[0083] By simultaneously solving the safety controller equation, the modal estimator equation, the state estimator equation, and the switching system equation, the closed-loop system equation is obtained, as shown in equation (19).
[0084]
[0085] in,
[0086] Extended state vector of closed-loop system Extended System Matrix
[0087] A φ(t) B φ(t) D φ(t) This is the system parameter matrix of the φ-th state estimator, where φ(t) represents the mode estimates obtained through the mode estimator; the system parameter matrix System parameter matrix Extended perturbation matrix Extended output matrix System parameter matrix System parameter matrix Initial state vector Initial time Subsystem and controller modes
[0088] Furthermore, S5 constructs switching rules based on replay attack detection. These switching rules are used to analyze the switching behavior of the closed-loop system under a replay attack based on the replay attack detection results, obtain the judgment conditions for system security and stability, and perform security control on the closed-loop system, including:
[0089] S51. Construct a switching rule based on replay attack detection, as shown in formula (20).
[0090]
[0091] The initial value of the switching rule is... State partitioning of the i-th subsystem D i boundary and
[0092] S52. Analyze the replay attack results based on the switching rules and the event triggering mechanism, determine the operating status of the subsystem mode and the controller mode under the replay attack, obtain the judgment conditions for system security and stability, and use the judgment conditions to perform security control on the switching system.
[0093] This invention establishes a switching Q-learning attack detection mechanism incorporating time-varying encryption / decryption functions. This dynamically changes the encryption method of the system state signals, amplifies the estimation residual value generated by the state estimator under replay attacks, and accurately analyzes whether a replay attack has occurred. By designing an event-triggered mechanism based on replay attack detection, the attack detection results are incorporated into various triggering conditions, shortening the asynchronous duration caused by replay attacks. Switching rules based on replay attack detection are designed to maintain the currently active subsystem mode unchanged or switch to a subsystem mode that meets the state switching conditions, thereby mitigating the continuous asynchronous operation between subsystems and subcontrollers caused by replay attacks. Finally, based on the event-triggered mechanism and switching rules, a security control scheme for the switching system under replay attacks is constructed. Considering the state data sharing among the subsystems of the switching system, security control judgment conditions are given to ensure the switching system's security under replay attacks, thus achieving secure control of the switching system. Attached Figure Description
[0094] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0095] Figure 1 This is a flowchart of the security control method for a networked handover system based on encryption / decryption switching Q-learning replay attack detection, as described in this invention.
[0096] Figure 2 This is a block diagram of the switching system structure of the present invention;
[0097] Figure 3 This diagram illustrates four scenarios for the switching signal in the closed-loop system of this invention.
[0098] Figure 4 This is a topology diagram of a switching system according to an embodiment of the present invention;
[0099] Figure 5 This is a trajectory curve diagram on the X and Y axes of an embodiment of the present invention;
[0100] Figure 6 This is a graph showing the change in state estimation residuals with and without an encryption / decryption module in one embodiment of the present invention.
[0101] Figure 7 This is an event triggering interval diagram for replay attack detection in one embodiment of the present invention;
[0102] Figure 8 One embodiment of the invention includes a switching signal and an asynchronous duration graph for replay attack detection;
[0103] Figure 9 This is a diagram showing the switching signal and asynchronous duration for replay attack detection in one embodiment of the present invention;
[0104] Figure 10 This is a diagram showing the X-axis tracking and positioning error of multiple unmanned vessels in one embodiment of the present invention.
[0105] Figure 11 This is a Y-axis tracking and positioning error diagram of multiple unmanned vessels in one embodiment of the present invention;
[0106] Figure 12 This is a diagram showing the heading angle tracking and positioning error of multiple unmanned vessels in one embodiment of the present invention. Detailed Implementation
[0107] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0108] This embodiment provides a security control method for a networked handover system based on encryption / decryption handover Q-learning replay attack detection, such as... Figure 1 As shown, it includes:
[0109] S1: Construct the switching system model and controller model to obtain state signals and modal signals;
[0110] S2: Construct an event triggering mechanism based on replay attack detection, the event triggering mechanism including error detection conditions, modality matching conditions and attack detection conditions; use the event triggering mechanism to process the state signal and modality signal to obtain the subsystem modality signal and state signal at the trigger time;
[0111] S3: Construct a replay attack detection mechanism based on switching Q-learning according to the Q-learning algorithm. The replay attack detection mechanism is used to encrypt and decrypt the subsystem state signal, and calculate the estimated values of the subsystem modal signal and state signal after encryption and decryption, as well as the state estimation residual norm. The event triggering mechanism and the replay attack detection mechanism are used to process the subsystem modal signal and state signal at the triggering time to obtain the state estimate, modal estimate and state estimation residual norm, and determine the detection result of the replay attack.
[0112] S4: Substitute the state estimate and the modality estimate into the controller model to obtain the security controller equation. Combine the security controller equation, the switching system model, and the equations of the state estimator and modality estimator of the replay attack detection mechanism to obtain the closed-loop system equation.
[0113] S5: Construct switching rules based on replay attack detection. The switching rules are used to analyze the switching behavior of the closed-loop system under replay attack based on the replay attack detection results, obtain the judgment conditions for system security and stability, and perform security control on the closed-loop system.
[0114] Specifically, firstly, a switching system model and a controller model are constructed to obtain state signals and modal signals, thus building an initial switching system and controller model to prepare for subsequent steps. Secondly, an event-triggered mechanism based on replay attack detection is constructed, which includes error detection conditions, modal matching conditions, and attack detection conditions. The event-triggered mechanism is used to process the state signals and modal signals to obtain the subsystem modal signals and state signals at the trigger time. Adding attack detection results to each trigger condition optimizes the data update time, reducing asynchronous duration. A replay attack detection mechanism based on switching Q-learning is constructed according to the Q-learning algorithm. This mechanism is used to encrypt and decrypt the subsystem state signals and calculate the estimated values of the encrypted and decrypted subsystem modal signals and state signals, as well as the state estimation residual norm. The event-triggered mechanism and the replay attack detection mechanism are used to process the subsystem modal signals and state signals at the trigger time. The system obtains state estimates, modal estimates, and the state estimate residual norm to determine the detection result of replay attacks. The attack detection mechanism based on switching Q-learning can obtain the optimal evaluation threshold, thereby more intelligently judging whether a replay attack has occurred, reducing false alarms, and improving the reliability of security control. Substituting the state estimates and modal estimates into the controller model, the security controller equation is obtained. Combining the security controller equation, the switching system model, and the equations of the state estimator and modal estimator of the replay attack detection mechanism, a closed-loop system equation is obtained. This equation, in conjunction with the attack detection algorithm, can effectively defend against replay attacks and reduce false alarms caused by normal system switching. Finally, a switching rule based on replay attack detection is constructed. This switching rule can determine whether the current state signal and the switching signal are historical signals replayed by the replay attack based on the replay attack detection result, and determine whether a subsystem switch is needed to ensure system stability.
[0115] A switching system block diagram of this embodiment is as follows: Figure 2 As shown,
[0116] In a specific embodiment, the scheme for constructing the switching system model and controller model to obtain state signals and modal signals is as follows:
[0117] S11. Construct the switching system model, as shown in formula (21).
[0118]
[0119] in, To switch the system's state vector, control input, system output, and external disturbances; It is a switching signal, A σ B σ C σ D σ It is a given constant matrix;
[0120] S12. Construct the controller model, as shown in formula (22).
[0121] u(t) = K φ x(t) (22)
[0122] Where φ = φ(t), φ(t) ∈ L are the modal signals of the controller, and K φ This is the controller gain.
[0123] This approach establishes an initial switching system and controller model to prepare for subsequent steps.
[0124] In a specific embodiment, an event-triggered mechanism based on replay attack detection is constructed. This mechanism includes error detection conditions, modality matching conditions, and attack detection conditions. The scheme for processing the state signal and modality signal using this mechanism to obtain the subsystem modality signal and state signal at the trigger moment is as follows:
[0125] S21. Construct an event triggering mechanism based on replay attack detection, as shown in formula (23).
[0126]
[0127] Wherein, the initial trigger time t1h=l0, α(g r,m h) represents the time interval at which the data reaches the controller at the (m-1)th sampling moment after the rth trigger by the attack detection mechanism. The attack detection results obtained above, These represent error detection conditions, modality matching conditions, and attack detection conditions, respectively.
[0128] S22. Construct error detection conditions, as shown in formula (24).
[0129]
[0130] in, Indicates sampling time g r,m The parameter matrix at h, where and Let ε be the weight matrix to be designed. r (t)=x(g r,m h)-x(t r h), g r,m h = t r h+mh represents the r-th trigger time t. r At the m-th sampling time after h, δ σ >0 is the trigger threshold; x(t) r h) represents the subsystem state signal at the trigger moment; ε r (t) represents the trigger time t. r The error between the state of h and its state at the m-th sampling time;
[0131] S23. Construct modal matching conditions, as shown in formula (25).
[0132]
[0133] Wherein, σ(g) r,m h) represents the r-th trigger time t r The modal signal at the m-th sampling time after h, then t r The modal signal at time h is σ(t) r h);
[0134] S24. Construct attack detection conditions, as shown in formula (26).
[0135]
[0136] This scheme constructs an event-triggered mechanism based on replay attack detection. By incorporating attack detection results into each triggering condition to optimize data update timing, it reduces asynchronous duration. Specifically, in the error detection condition, the update timing of correct data is optimized based on whether x(t) and σ(t) are historical data packets replayed in a replay attack. In the modality matching condition, triggering is promptly arranged after subsystem switching based on whether x(t) has been replayed, shortening the asynchronous duration caused by subsystem switching. In the attack detection condition, triggering is promptly arranged after the replay attack ends, mitigating the asynchronous duration caused by the replay attack. Furthermore, the trigger weight matrix in the error detection condition... and Add dissipative parameter matrix Y σ This enables the adjustment of external data transmission based on changes in the system's internal energy.
[0137] In a specific embodiment, a replay attack detection mechanism based on switching Q-learning is constructed according to the Q-learning algorithm. The subsystem modal signals and state signals at the triggering time are processed using the event triggering mechanism and the replay attack detection mechanism to obtain state estimates, modal estimates, and state estimate residual norms. The scheme for determining the detection result of the replay attack is as follows:
[0138] S31. Construct a replay attack detection mechanism based on switching Q-learning, wherein the replay attack detection mechanism includes encryption measures, decryption measures, state estimation measures, mode estimation measures, and a switching Q-learning method based on Q-learning:
[0139] S32. The encryption measure is used to encrypt the subsystem state signal at the trigger time to obtain the encrypted state signal, as shown in formula (27).
[0140] x c (t r h)=x(t r h)+c(t r h) (27)
[0141] Wherein, c(t) r h) is a time-varying encryption function generated from a random seed, x c (t r h) represents the encrypted state signal, x(t) r h) represents the trigger time t before encryption. r The system state at point h;
[0142] The encrypted state signal and the switching signal, i.e., the modal signal σ(t) r h) under the influence of replay attacks, it can be rewritten as formula (28).
[0143]
[0144] Where α(t) represents the detection result of the replay attack detection to be solved;
[0145] S33. The decryption measure is used to decrypt the encrypted signal to obtain the decrypted transmission signal, as shown in formula (29).
[0146]
[0147] in, d(t) represents the decrypted system state signal. r h) is the decryption key generated from the random seed, and the energy limit of the time-varying encryption function and the decryption key satisfies the condition. Given a constant matrix;
[0148] Due to transmission delays in the communication network, the data x transmitted at the trigger time... c (t r h) and σ(t) r h) After passing through the network, The time arrives at the sub-controller, in This represents the transmission delay of the data triggered for the rth time, and therefore the signal's hold interval. It can be divided into Where Φ r,m The subinterval is represented as shown in formula (30).
[0149]
[0150] Represents the maximum number of subintervals, and the piecewise time delay function η r (t)=tg r,m h satisfies the condition
[0151] Under the action of the event triggering mechanism based on replay attack detection and the encryption / decryption measures, the actual switching signal and actual state signal of the switching system are as shown in formulas (31) and (32).
[0152]
[0153] Where, η r (t)=tg r,m h represents the piecewise time delay function; ε r (t) represents the trigger time t. r The error between the state of h and its state at the m-th sampling time; ε p (t) represents the error between the state at time p and the state at the m-th sampling time thereafter;
[0154] S34. The state estimation measure is implemented by a state estimator, which is used to calculate the state estimate of the system state signal under the actual switching signal, as shown in formula (33).
[0155]
[0156] in, L is the state estimate of the switching system under the actual switching signal of x(t). φ(t) The observer gain, φ = φ(t), φ(t) ∈ L, represents the modal estimate of the switching system under the actual switching signal; A φ(t) B φ(t) Dφ(t) It is the system parameter matrix of the φ(t)th state estimator;
[0157] S35. The modal estimation measure is implemented by a modal estimator, which is used to calculate the modal estimate value of the modal signal in the switching system under the actual switching signal, as shown in formula (34).
[0158]
[0159] in, This represents the estimated state partition corresponding to the i-th state estimator. yes The boundary, and
[0160] This scheme constructs a replay attack detection mechanism based on handover Q-learning. By dynamically changing the encryption method of the system state signals through encryption / decryption functions, it amplifies the estimation residual value generated by the state estimator under replay attacks, thereby distinguishing replay attacks from normal handover behavior. In the absence of replay attacks, the original transmission signal is recovered from the encrypted information for normal system use.
[0161] In a specific embodiment, the Q-learning-based switching Q-learning method is used to calculate the difference between the state estimate and the original transmitted signal, i.e., the state estimate residual norm, and to determine the detection result of the replay attack based on the state estimate residual norm and the modality estimate, including a Q-table training phase and an attack detection phase:
[0162] S36. Switch to Q-table training phase:
[0163] Construct a quintuple (A, S, γ) f ,π f For a Markov decision process of f(k), train the Q-table and update the Q-value, where f∈F={0,1,…,F} is the number of training iterations;
[0164] Let φ be a set of environment variables. f (g r,m h) is the r,m-th learning interval Φ in the f-th training round. r,m The modal estimator provides estimates of the subsystem modes. It is the residual norm of the state estimation at step r,m in the f-th round of training. It is in the fth round of training To reduce computational complexity, Divide into K mutually exclusive intervals [θ] based on their numerical values. k ,[θ k+1), k∈K, where [θ k Given a sequence of real numbers θ0 = 0 < θ1 < ... < θ k <∞, therefore, the environment variable at step r,m is expressed as
[0165] Action Set Indicates that the detector is based on environment variables. The detection alarm values are defined as follows: when a1 = 0, it indicates that no replay attack has been detected; when a2 = 1, it indicates that a replay attack has been detected.
[0166] reward function γ f (r,m) represents the environment variables at step r,m in the f-th training round. In performing the action The rewards obtained later, among which
[0167] Strategy This represents the environment variables at step r,m in the f-th training round. Execute action The probability of the discount parameter k∈[0,1] represents the degree of attention to long-term rewards;
[0168] S361, Training in the f-th round: Initializing the action Environment variables at steps r and m According to strategy Select Action As shown in formula (35),
[0169]
[0170] Specifically, generate a random number c∈[0,1], if So
[0171] The environmental variables at step r,m By performing actions Transform to environment variables in step r,m+1 According to the reward function γ f (r,m) Calculate the action to be executed The reward obtained later is shown in formula (36).
[0172]
[0173] Where c1 and c2 are positive constants, and α f These are known attack scenarios during training;
[0174] S362. Update the Q value according to the reward function and the value of the (f-1)th training round, as shown in formula (37).
[0175]
[0176] in, For the environment variable-action pair during the f-th training round The Q-value after the (f-1)th training round. Environment variables for the f-th round of training The Q-table value α corresponding to the optimal action after the (f-1)th round of training. f =0.1·(2+arctan(f-α) * )) is the learning step size, α * Let be the learning constant, where the initial value of the Q table is . r∈N + m∈N;
[0177] S363, Regarding the environment variables at step r,m+1 Repeat the process of S361-S362 to obtain the corresponding Q value until the system stops running, and the training of the fth round ends.
[0178] S364. Repeat steps S361-S363 to perform the next round of training until f = F. Training ends, and the system returns a list containing all environment variable-action pairs. Complete switching of the Q table;
[0179] S37. During the attack detection phase, the corresponding optimal action is selected from the complete switching Q table.
[0180] S371. Determine environmental variables based on the modal estimates and state estimation residual norms in the complete switching Q table:
[0181] Let r = 1, m = 0, when t ∈ Φ r,m At that time, observe the current environmental variables and the current modal estimates;
[0182] S372. Select the optimal action corresponding to the current environment variable from the complete switching Q table.
[0183] S373. Output the attack detection results, as shown in formula (38).
[0184] α(t)=a r,m ,t∈Φ r,m (38).
[0185] This embodiment constructs a handover Q-learning algorithm for detecting replay attacks. Compared with the common method of setting a fixed evaluation threshold to detect attacks in handover systems, the attack detection mechanism based on handover Q-learning can obtain the optimal evaluation threshold, thereby more intelligently determining whether a replay attack has occurred or whether the system is undergoing a normal handover, thus reducing false alarms and improving the reliability of security control.
[0186] In a specific embodiment, the scheme of substituting the state estimate and the modal estimate into the controller model to obtain the security controller equation, and simultaneously solving the security controller equation, the switching system model, and the equations of the state estimator and modal estimator of the replay attack detection mechanism to obtain the closed-loop system equation is as follows:
[0187] Substituting the state estimate and the modal estimate into the controller model, the safety controller equation is obtained, as shown in formula (39).
[0188]
[0189] φ = φ(t), φ(t) ∈ L, representing the modal estimate obtained from the modal estimator;
[0190] By simultaneously solving the safety controller equation, the modal estimator equation, the state estimator equation, and the switching system equation, the closed-loop system equation is obtained, as shown in formula (40).
[0191]
[0192] Among them, the extended state vector of the closed-loop system Extended System Matrix
[0193] A φ(t) B φ(t) D φ(t) This is the system parameter matrix of the φ-th state estimator, where φ(t) represents the mode estimates obtained through the mode estimator; the system parameter matrix System parameter matrix Extended perturbation matrix Extended output matrix System parameter matrix System parameter matrix Initial state vector Initial time Subsystem and controller modes
[0194] This solution constructs a closed-loop system, which includes a state estimator and a modality estimator. It can work with attack detection algorithms to effectively defend against replay attacks and reduce false alarms caused by normal system switching.
[0195] In a specific embodiment, a switching rule based on replay attack detection is constructed. This switching rule is used to analyze the switching behavior of the closed-loop system under a replay attack based on the replay attack detection results, obtain the judgment conditions for system security and stability, and the scheme for security control of the closed-loop system is as follows:
[0196] S51. Construct a switching rule based on replay attack detection, as shown in formula (41).
[0197]
[0198] The initial value of the switching rule is... State partitioning of the i-th subsystem D i boundary and
[0199] S52. Analyze the replay attack results based on the switching rules and the event triggering mechanism, determine the operating status of the subsystem mode and controller mode under the replay attack, obtain the judgment conditions for system security and stability, and use the judgment conditions to perform security control on the switching system:
[0200] To facilitate the analysis of handover behavior, it is assumed that the j-th and i-th subsystems are in adjacent handover intervals [l q-1 ,l q ) and [l q ,l q+1 The system is activated within the subsystem. Due to the attack, normal state signals and modal signals are altered into historical data M. This depends on whether the attack covers the subsystem. or sub-controller switching time And replay attacks either did not log or logged at least one switching action, any [l] q ,l q+1 The changes in the modal signal (σ,φ) over the interval can be divided into the following categories: Figure 3 The four cases shown:
[0201] In this embodiment, the network switching is affected by a random replay attack. The probability of the attack occurring follows a Bernoulli distribution α(t)∈{0,1}. When α(t)=1, historical data is recorded before the attack replays; when α(t)=0, data is transmitted normally. Assume the attacker can record an upper limit of data packets. If there are 1, then the stored data is described as follows: The p-th replay attack also needs to satisfy the following assumptions about the number of attacks;
[0202] Assumption 1 (Number of attacks): There exists w p ≥0 and τ p ≥h, such that in In the interval The number of times an internal replay attack occurred;
[0203] Definitions of supply ratio and cross-supply ratio:
[0204] If in t∈[l] q ,l q+1 When σ(t) = i, there exists a positive definite storage energy function. Supply rate Π ii (y,w) and cross-supply rate Π ij (χ,y,w,t), This makes formulas (42) and (43) true.
[0205]
[0206] for exist This makes formula (44) true.
[0207]
[0208] So when the control input u(t) = 0, the sequence is switched. The switching system is dissipative, in which
[0209] Π ii (t)=w T (t)Q i w(t)+2w T (t)U i y(t)+y T (t)Y i y(t),
[0210] Π ij (t)=a ij [w T (t)Q i w(t)+2w T (t)U i y(t)+y T (t)Y i y(t)]+w T (t)Q ij w(t),
[0211] εij (t)=χ T (t)F ij χ(t),
[0212] a ij >0 is a given scalar, Q ij >0, F ij >0, Q i U i ,Y i A matrix of appropriate dimension;
[0213] Theorem 1: Under Assumption 1, for a given positive constant h, η , and a ij >0, parameter And i≠j, The closed-loop system is asymptotically stable under the encryption / decryption switching Q-learning attack detection mechanism if the following conditions are met.
[0214] (1) There exists a matrix S 0 >0, R 0 >0, Q ij >0, F ik >0, pair matrix M 1 M 2 and a matrix of appropriate dimension Satisfying formulas (45), (46) and (47),
[0215]
[0216] in It consists of the following block matrices:
[0217]
[0218]
[0219] In the formula
[0220]
[0221] (2) There exists an indeterminate matrix of appropriate dimension. Satisfying formulas (48), (49) and (50),
[0222]
[0223] in It consists of the following block matrices:
[0224]
[0225] Proof: To achieve the control objective, it is proven that the closed-loop system is asymptotically stable. First, candidate Lyapunov functions are selected, as shown in equation (51).
[0226]
[0227] in,
[0228]
[0229] Calculate the time derivative of formula (51) along the closed-loop system and take the expected value to obtain...
[0230]
[0231] right Applying Jensen's integral inequality, Park's lemma, and formula (46) to the integral terms in the equations, we obtain formulas (52) and (53).
[0232]
[0233] in,
[0234] Will and Substituting into formula (41), we obtain the inequality.
[0235]
[0236] Substituting the inequalities and formulas (52) and (53) into Based on formulas (45) and (46), we obtain formula (54).
[0237]
[0238] Based on the following conditions, the changes in subsystem mode σ and controller mode φ in AD analysis are discussed, and the equation (51) is applied in any interval [l]. q ,l q+1 Changes within )
[0239] Case A: In the interval Above, (σ,φ)=(i,j), in the interval Above, (σ,φ)=(i,i), in the interval Above, (σ,φ)=(i,φ) p ), in the interval Above, (σ,φ)=(i,i), in the interval Above, (σ,φ)=(i,i);
[0240] By combining the modal estimator formula (34), the switching rule formulas (41) and (48), it can be derived that under case A, formula (51) is in the switching interval [l q ,l q+1 The evolution within ) is shown in equation (55).
[0241]
[0242] Case B: In the interval Above, (σ,φ)=(i,j), in the interval Above, (σ,φ)=(i,φ) p ), in the interval Above, (σ,φ)=(i,i);
[0243] By combining the modal estimator formula (34), the switching rule formulas (41) and (48), it can be derived that under case B, formula (51) is in the switching interval. The evolution within is shown in formula (56).
[0244]
[0245] Case C: In the interval Above, (σ,φ)=(i,j), in the interval Above, (σ,φ)=(i,i), in the interval superior, In the interval Above, (σ,φ)=(i,i);
[0246] By combining the modal estimator formula (34), the switching rule formulas (41) and (48), it can be derived that under case C, formula (51) is in the switching interval [l q ,l q+1 The evolution within ) is shown in formula (57).
[0247]
[0248] Case D: In the interval Above, (σ,φ)=(i,j), in the interval superior, In the interval Upper(σ,φ)=(i,i);
[0249] By combining the modal estimator formula (34), the switching rule formulas (41) and (48), it can be derived that under case D, formula (51) is in the switching interval [lq ,l q+1 The evolution within ) is shown in formula (58).
[0250]
[0251] Therefore, we can obtain the result through formula (48). and For case AD, formula (51) is valid in any interval [l q ,l q+1 The changes in ) can be summarized by formula (59).
[0252]
[0253] Discussion of subsystem switching points Changes in formula (51) before and after:
[0254] Under the attack detection-based switching rule, formula (60) is obtained.
[0255]
[0256] Discuss V(t) and its operation on the interval [0,t). Relational derivation:
[0257] Based on formula (49) and formulas (59) and (60), we can obtain formula (61).
[0258]
[0259] In formula (61) Move to the left, due to F ik ≥0, we can obtain formula (62),
[0260]
[0261] That is, formulas (42) and (43) hold true. On the other hand, according to formula (47) and Q ij If the value is greater than 0, then formula (63) can be obtained.
[0262] Π ik (t)<0 (63)
[0263] That is, formula (44) holds, therefore, the closed-loop system satisfies the dissipation property of Theorem 1;
[0264] Prove the asymptotic stability of the closed-loop system:
[0265] When w(t) = 0, substituting formula (62) into formula (58) yields formula (64).
[0266]
[0267] The control objective was achieved;
[0268] Formula (45) in Theorem 1 and formula (48) Linearization is performed on the nonlinear terms:
[0269] Theorem 2: Under Assumption 1, for a given positive constant h, η , w and a ij >0, parameter And i≠j, If formulas (46), (47), (49), and (50) in Theorem 1 hold, and there exists a matrix S 0 >0, R 0 >0, Q ij >0, F ik >0, k∈{i,j} Opposing Matrix M 1 M 2 With a matrix of appropriate dimensions Satisfying formulas (65) and (66),
[0270]
[0271] The closed-loop system then becomes asymptotically stable, where equation (65) states... It consists of the following matrices:
[0272]
[0273] In formula (66) It consists of the following block matrices:
[0274]
[0275] The other block matrices in the matrix are zero matrices of appropriate dimension;
[0276] Proof: First, multiply both sides of formula (45) by a matrix. and its transpose, in which It can be obtained that consists of 10×10 matrix blocks
[0277]
[0278] Multiply the left and right sides of formula (48) by a matrix and its transpose, in which It can be obtained that consists of 7x7 matrix blocks. Applying Schur's complement lemma, we obtain -[4R]. -1 <4w 2 R-2wI,-[4(1-α)R] -1 <4w 2 (1-α)R-2wI,-[4αR] -1 <4w 2 αR-2wI, thus having and Equivalent to formulas (45) and (48); with the help of Schur's supplementary lemma, formulas (45) and (48) can be derived from formulas (65) and (66).
[0279] This solution constructs a switching rule based on replay attack detection. Based on the replay attack detection results, and considering whether the current state signal and the switching signal are historical signals replayed in a replay attack, it maintains the currently active subsystem mode or switches to a subsystem mode that meets the state switching conditions. This mitigates the continuous asynchronous operation between subsystems and subcontrollers caused by replay attacks, determines whether a subsystem switch is necessary, and ensures system stability. (Supply rate Π) ii (y,w) and cross-supply rate Π ij (χ,y,w,t) not only describes the energy exchange between each subsystem and external inputs, but also reflects the impact of activated and inactive subsystems on the overall energy exchange of the switching system. Treating the energy of an attack as a special type of external input, and analyzing the situation constructed under AD... The changes in the system's internal energy and external input under replay attacks are used to reveal the relationship between the system's internal energy and external input. The impact of replay attacks on the overall stability of the system is described, and security conditions are constructed to ensure the system operates under these conditions, thus achieving security control under replay attacks.
[0280] Simulation is performed using an unmanned surface vessel as an example:
[0281] The effectiveness of the method is verified using a case study of dynamic positioning control for multi-surface unmanned surface vessels. The dynamics of the unmanned vessel can be expressed as formula (67).
[0282]
[0283] in and Indicates the X-axis position and Y-axis position of the nth ship. Let n represent the ship's heading angle, where n ∈ {1, 2, 3}. Indicates oscillation speed, Indicates sway speed, This represents the yaw rate. u n To control the input, M n , Π n and for
[0284]
[0285] Where m n x n,g and I n,z Let represent the mass of the nth ship, its center of gravity on the X-axis, and its moment of inertia on the heading, respectively. It is the added mass in the longitudinal direction. and Indicates the additional mass in the sway direction. The additional moment of inertia in the heading direction, and The damping coefficient is linear.
[0286] Unmanned surface vessel No. 1 acts as the leader and is responsible for receiving positioning information. Unmanned surface vessels No. 2 and No. 3 act as follower vessels and maintain cooperative tracking and positioning by receiving relative position information from No. 1.
[0287] Specifically, let the state of the nth unmanned vessel be... In formula (67), the dynamic positioning objective of the unmanned vessel is to adjust the vessel's position and course so that... Among them ι n Represents the expected formation vector. For a single desired location point, ρ n For preset scalar;
[0288] The relevant parameters for the unmanned surface vessel are:
[0289]
[0290] The initial state of the three unmanned boats is as follows: Other parameters are shown in Table 1.
[0291] Table 1 Parameters of the three unmanned vessels
[0292]
[0293] Figure 4The simulation verification of the safety control method uses a switching system topology. Due to the limited communication channels of unmanned surface vessel 1, it can only communicate with one accompanying vessel at a time, and accompanying vessels 2 and 3 cannot communicate directly with each other. Specifically, when unmanned surface vessel 1 sends information only to unmanned surface vessel 2, it is topology 1, as shown below. Figure 4 As shown on the left. On the other hand, when unmanned surface vessel 1 only sends information to unmanned surface vessel 3, it is topology 2, as shown... Figure 4 As shown on the right, since only one follower ship can receive information at a time, the other follower ship that is not being communicated with may deviate from the queue, making each topology unstable when operating alone. To ensure that the unmanned ships can achieve cooperative tracking and positioning, it is necessary to switch between the two topologies;
[0294] based on Figure 4 Due to changes in the topology, the cooperative localization problem of the three ships can be modeled as a switching system containing two unstable subsystems, as shown in equation (68).
[0295]
[0296] in In the formula K n For the gain of the controller to be designed, The communication topology matrix is ρ∈{1,2}, and the other coefficient matrices are as follows.
[0297]
[0298] The control objective of formula (67) is then transformed into the stabilization problem of formula (68);
[0299] Select a replay attack that satisfies Hypothesis 1: The input parameters of the algorithm are: discount parameter κ = 0.42, training iterations F = 50, action set A = {0, 1}, and exploration constant β. * =3, learning constant α * =5, in the reward function c1=10, c2=5; state estimation residual norm partitioning parameters: θ0=0, θ1=0.02, θ2=0.04, θ3=0.06, θ4=0.08, θ5=0.1, θ6=0.2, θ7=0.3, θ8=0.4, θ9=1; encryption / decryption function is designed as c(t)=0.05arctan(t+μ c ), d(t)=0.05arctan(t+μ) d ), μ c ,μ d Random numbers generated using the same random seed; training data consists of subsystem modes generated during system runtime. and the residual norm of state estimation Choose h = 0.01, η =0.003, δ1=δ2=0.1, a 12 =a 21 =0.2, n 12 =-1, w=0.9, solving Theorem 2 yields the controller gains for the three ships.
[0300]
[0301] In this embodiment, the position and state of the unmanned vessel are as follows: Figures 5-12 As shown,
[0302] Figure 5 The image shows that the three unmanned vessels have reached the desired location.
[0303] Figure 6 The changes in the norm of the state estimation residuals were compared with and without encryption / decryption modules. Specifically, the switching Q-learning method for detecting replay attacks, which incorporates encryption and decryption modules into the switching Q-learning algorithm, divides the state residuals into {θ0,…,θ9} and performs 50 rounds of training; the replay attack detection process without encryption / decryption modules is consistent with the switching Q-learning algorithm. The comparison results are as follows: Figure 6 As shown, the encryption / decryption process can effectively amplify the residual norm value of the state estimation under the attack shadow area, making replay attacks easier to detect, thus verifying the effectiveness of the encryption / decryption module.
[0304] Figure 7 Compare the event trigger counts, switching points, and asynchronous durations after the attack ends with / without replay attack detection with Table 2. Specifically, the attack detection-based event triggering mechanism incorporates the attack detection result α(t) into each triggering condition; the remaining triggering conditions of the multi-condition event triggering mechanism without incorporating the attack detection result remain consistent with the event triggering mechanism.
[0305] Table 2. Number of event triggers and asynchronous duration (T=60s) for replay attack detection (with / without replay attack)
[0306] Event triggering mechanism Trigger count Switching point and asynchronous duration after attack ends Event-triggered mechanism based on attack detection 103 2.89s Event triggering mechanism without attack detection 139 4.17s
[0307] The results are shown in Table 2 and Figure 7 As shown, from the second column of Table 2 and Figure 7 It can be seen that the event triggering mechanism of introducing α(t) can avoid the triggering of erroneous data and increase the triggering interval; as can be seen from the third column of Table 2, the event triggering mechanism based on attack detection can effectively reduce the asynchronous duration between the switching point and the end of the attack.
[0308] Figure 8 and Figure 9The asynchronous duration and the number of consecutive asynchronous switches were compared under switching rules with and without replay attack detection. Specifically, the switching rule based on replay attack detection introduced the attack detection result α(t); the switching rule without attack detection, except for not introducing α(t), maintained the same switching conditions as the fantasy rule. It can be seen that, under the same replay attack, the switching rule introducing α(t) effectively shortened the asynchronous time caused by the replay attack and effectively avoided the six consecutive asynchronous switches that occurred in the switching rule without attack detection.
[0309] Figure 10 , Figure 11 and Figure 12 Comparing the tracking and positioning errors of multiple unmanned surface vessels (USVs) on the X, Y axes, and heading angles, considering the cross-feed rate of the inactive subsystem and without introducing dissipation, it can be seen that considering the cross-feed rate of the inactive subsystem results in better convergence speed and stability.
[0310] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A security control method for a networked handover system based on encryption / decryption handover Q-learning replay attack detection, characterized in that, include: S1: Construct the switching system model and controller model to obtain state signals and modal signals; the specific steps are as follows: S11. Construct the switching system model, as shown in formula (1). (1) in, , , , This is to switch the system's state vector, control input, system output, and external disturbances; It is a switching signal. , , , It is a given constant matrix; S12. Construct the controller model, as shown in formula (2). (2) in, These are the controller's modal signals. For controller gain; S2: Construct an event-triggered mechanism based on replay attack detection. The event-triggered mechanism includes error detection conditions, modality matching conditions, and attack detection conditions. Use the event-triggered mechanism to process the state signal and modality signal to obtain the subsystem modality signal and state signal at the trigger time. Specific steps are as follows: S21. Construct an event triggering mechanism based on replay attack detection, as shown in formula (3). (3) Among them, the initial trigger time , This represents the time interval at which the data arrives at the controller at the (m-1)th sampling time after the r-th trigger. The attack detection results obtained above , , These represent error detection conditions, modality matching conditions, and attack detection conditions, respectively. S22. Construct error detection conditions, as shown in formula (4). (4) in, , indicating the sampling time The parameter matrix at the point, where and The weight matrix to be designed, , for The next m Next sampling time This is the trigger threshold; The subsystem state signal indicating the trigger moment; Indicates the trigger time The state and its subsequent m The error between the states at each sampling time; S23. Construct modal matching conditions, as shown in formula (5). (5) in, express The next m The modal signal at the next sampling time, then The modal signal at time t is ; S24. Construct attack detection conditions, as shown in formula (6). (6); S3: Construct a replay attack detection mechanism based on switching Q-learning according to the Q-learning algorithm. The replay attack detection mechanism is used to encrypt and decrypt the subsystem state signal, and calculate the estimated values of the subsystem modal signal and state signal after encryption and decryption, as well as the state estimation residual norm. The event triggering mechanism and the replay attack detection mechanism are used to process the subsystem modal signal and state signal at the triggering time to obtain the state estimate, modal estimate and state estimation residual norm, and determine the detection result of the replay attack. S4: Substitute the state estimate and the modality estimate into the controller model to obtain the security controller equation. Combine the security controller equation, the switching system model, and the equations of the state estimator and modality estimator of the replay attack detection mechanism to obtain the closed-loop system equation. S5: Construct switching rules based on replay attack detection. The switching rules are used to analyze the switching behavior of the closed-loop system under replay attack based on the replay attack detection results, obtain the judgment conditions for system security and stability, and perform security control on the closed-loop system.
2. The security control method for a networked handover system based on Q-learning replay attack detection during encryption / decryption switching, as described in claim 1, is characterized in that... The replay attack detection mechanism includes encryption measures, decryption measures, state estimation measures, modality estimation measures, and a Q-learning-based switching Q-learning method; The encryption measures are used to encrypt the subsystem state signal at the triggering time to obtain the encrypted state signal, and based on the encrypted state signal, to obtain the encrypted state signal and modal signal under replay attack, and based on the encrypted modal signal and the event triggering mechanism, to obtain the actual switching signal; The decryption measures are used to decrypt the encrypted state signal under the replay attack, and to obtain the actual system state signal based on the decrypted system signal and the event triggering mechanism. The state estimation measure is implemented by a state estimator, which is used to calculate the state estimate of the system state signal in the switching system under the actual switching signal. The modal estimation measure is implemented by a modal estimator, which is used to calculate the modal estimate of the modal signal in the switching system under the actual switching signal. The Q-learning-based switching Q-learning method is used to determine the detection result of replay attack based on the difference between the actual system state signal and the state estimate, i.e., the state estimate residual norm, and the modality estimate.
3. The security control method for a networked handover system based on Q-learning replay attack detection during encryption / decryption switching, as described in claim 2, is characterized in that... The subsystem modal signals and state signals at the triggering moment are processed using the event triggering mechanism and the replay attack detection mechanism to obtain state estimates and modal estimates, including: S31. The encryption measures are as shown in formula (7). (7) in, It is a time-varying encryption function generated from a random seed. This indicates the encrypted state signal. Indicates the trigger time before encryption. The system status at that location; The encrypted state signal and switching signal are i.e., mode signals. Under the influence of replay attacks, it is rewritten as formula (8). (8) in This represents the detection result of the replay attack detection problem to be solved; S32. The decryption measures are as shown in formula (9). (9) in, This indicates the system status signal after decryption. The decryption key is generated from the random seed, and the energy limit of the time-varying encryption function and the decryption key satisfies the following condition. , , Given a constant matrix; Under the action of the event triggering mechanism based on replay attack detection and the encryption / decryption measures, the actual switching signal and the actual system state signal of the switching system are as shown in formulas (10) and (11). (10) (11) in, , representing a piecewise time delay function; Indicates the trigger time The state and its subsequent m The error between the states at each sampling time; Represents the state at time p and its subsequent state. m The error between the states at each sampling time; S33, The state estimator is as shown in formula (12). (12) in, yes The state estimate of the switching system under the actual switching signal. It is the observer gain. , representing the modal estimate of the switching system under the actual switching signal; , , It is the first The system parameter matrix of each state estimator; S34. The modality estimator is as shown in formula (13). (13) in, Indicates the first i The estimated state partitioning corresponding to each state estimator , , yes The boundary, and .
4. The security control method for a networked handover system based on Q-learning replay attack detection during encryption / decryption switching, as described in claim 2, is characterized in that... The Q-learning-based switching Q-learning method includes switching between the Q-table training phase and the attack detection phase: S35, Switch to Q-table training phase: Construct a quintuple The Markov decision process trains the Q-table and updates the Q-value, where For the number of training iterations; For the set of environment variables, where It is the first f In the first round of training Learning intervals The modal estimator provides estimates of the subsystem modes. It is the first f In the first round of training On-step state estimation residual norm , It is in the fth round of training ; Action Set Indicates that the detector is based on environment variables. The detection alarm value, when This indicates that no replay attack was detected. This indicates that a replay attack has been detected. reward function Indicates the first f In the first round of training, Step environment variables In performing the action The rewards obtained later, among which , ; Strategy Indicates the first f In the first round of training, Step environment variables Execute action The probability of discount parameters This indicates the level of attention paid to long-term rewards; S351. Initialize the actions for the f-th round of training. ;No. Step environment variables According to strategy Select Action As shown in formula (14), (14) The first Step environment variables By performing actions Switch to the Step environment variables According to the reward function Calculate the execution action The reward obtained later is shown in formula (15). (15) S352, Based on the reward function and the... f-1 The Q-value is updated based on the values from each training round, as shown in formula (16). (16) in, For the first f Environmental variables-action pairs during round training In the f-1 The Q-value after one round of training. For the first f Environment variables for round training In the f-1 The Q-table value corresponding to the optimal action after one round of training. To learn step length, Let be the learning constant, where the initial value of the Q table is . , ; S353, regarding the first Step environment variables Repeat steps S351-S352 to obtain the corresponding Q value until the system stops running. f Round of training has ended; S354. Repeat steps S351-S353 to proceed to the next round of training until... Training is complete, and the returned value contains all environment variables and action pairs. Complete switching of the Q table; S36. During the attack detection phase, the corresponding optimal action is selected from the complete switching Q table. : S361. Determine environmental variables based on the modal estimates and state estimation residual norms in the complete switching Q table: make , ,when At that time, observe the current environmental variables and the current modal estimates; S362. Select the optimal action corresponding to the current environment variable from the complete switching Q table. ; S363. Output the attack detection results, as shown in formula (17). , (17) 。 5. The security control method for a networked handover system based on Q-learning replay attack detection during encryption / decryption switching according to claim 1, characterized in that, S4 substitutes the state estimate and the modal estimate into the controller model to obtain the security controller equation. Simultaneously, it establishes the security controller equation, the switching system model, and the equations of the state estimator and modal estimator of the replay attack detection mechanism to obtain the closed-loop system equation, including: Substituting the state estimate and the modal estimate into the controller model, the safety controller equation is obtained, as shown in formula (18). , (18) , representing the mode estimate obtained from the mode estimator; By simultaneously solving the safety controller equation, the modal estimator equation, the state estimator equation, and the switching system equation, the closed-loop system equation is obtained, as shown in formula (19). (19) in, Extended state vector of closed-loop system Extended system matrix , , , It is the first The system parameter matrix of each state estimator This represents the modal estimates obtained through the modal estimator; the system parameter matrix. System parameter matrix , extended perturbation matrix Expand the output matrix System parameter matrix System parameter matrix Initial state vector Initial time Subsystem and controller modes .
6. The security control method for a networked handover system based on Q-learning replay attack detection during encryption / decryption switching according to claim 1, characterized in that, S6 constructs a handover rule based on replay attack detection, inputs the replay attack detection result into the event triggering mechanism and the handover rule, analyzes the handover behavior under replay attack, obtains the judgment conditions for system security and stability, and performs security control on the handover system, including: S51. Construct a switching rule based on replay attack detection, as shown in formula (20). (20) The initial value of the switching rule is... State partitioning of the i-th subsystem , boundary ,and ; S52. Analyze the replay attack results based on the switching rules and the event triggering mechanism, determine the operating status of the subsystem mode and the controller mode under the replay attack, obtain the judgment conditions for system security and stability, and use the judgment conditions to perform security control on the switching system.
Citation Information
Patent Citations
Switching type adaptive event trigger control method and system and storage medium
CN116800463A
Power system load frequency safety control method and system under network attack
CN117459283A