Security control method for networked switching system based on switching q-learning deception attack detection

By constructing a spoofing attack detection mechanism based on handover Q-learning, the security control problem of networked handover systems under spoofing attacks is solved, achieving accurate detection and secure control of state signals and modal signals, and ensuring system stability.

CN118759932BActive Publication Date: 2026-01-06DALIAN MARITIME UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410860051.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2026-01-06
Estimated Expiration
2044-06-28

AI Technical Summary

Technical Problem

Existing networked switching systems cannot accurately determine whether they are under attack when faced with spoofing attacks, resulting in inaccurate system status analysis and difficulty in security control.

Method used

A deception attack detection mechanism based on switching Q-learning is constructed, including an event triggering mechanism, state and modal signal processing, Q-learning algorithm and switching rules. By detecting errors, modal matching and attack conditions, it is determined whether the system signals have been tampered with, and a security controller equation is constructed to achieve security control.

Benefits of technology

It can accurately detect spoofing attacks, reduce false alarms, shorten the system's asynchronous duration, alleviate continuous asynchronous operations, ensure the security and stability of the switching system under spoofing attacks, and provide security control judgment conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118759932B_ABST
    Figure CN118759932B_ABST
Patent Text Reader

Abstract

The application discloses a kind of security control method of networked switching system based on switching Q learning deception attack detection, comprising: constructing switching system model and controller model, obtain state signal and modal signal;Event-triggering mechanism based on deception attack detection is constructed, including error detection condition, modal matching condition and attack detection condition;According to Q learning algorithm, construct deception attack detection mechanism based on switching Q learning, use event-triggering mechanism and deception attack detection mechanism to process modal signal and state signal, obtain state estimation value, modal estimation value, state estimation residual norm and modal estimation residual norm, determine the detection result of attack;State estimation value and controller modal are substituted into controller, obtain security controller, simultaneous security controller, state estimator, modal estimator and switching system, obtain closed-loop system;Switching rule based on deception attack detection is constructed, according to attack detection result analysis deception attack under closed-loop system switching behavior, obtain the judgment condition of system security stability, the present application can determine optimal evaluation threshold, to judge whether the state signal and modal value of system transmission to controller are tampered with by deception attack respectively, give the security control discrimination condition of ensuring switching system under deception attack, realize the security control of switching system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of handover system security control technology, and in particular to a security control method for networked handover systems based on handover Q-learning spoofing attack detection. Background Technology

[0002] As modern control systems increasingly move towards networking, the complexity of industrial system models is also constantly increasing. To meet the demands of changing working environments and improve operational flexibility, many systems in practical applications are beginning to exhibit multimodal characteristics. Simultaneously, the need for remote spatial distribution and low-cost control is growing, requiring remote monitoring and control via networks. Faced with complex industrial production environments, networked switching systems and their safety control have become current research hotspots.

[0003] As a typical network attack, spoofing attacks severely disrupt normal system operation and pose a significant threat to system stability by tampering with network communication data packets. Particularly in switched systems, subsystem modalities and state information can be tampered with during network transmission to the controller, leading to multiple consecutive asynchronous operations between the subsystem and the controller. This switching behavior is not only difficult to analyze but also greatly increases the difficulty of system stability. For spoofing attack detection in non-spoofing systems, common methods include state estimation-based detection, cumulative sum detection, and reinforcement learning-based attack detection, which primarily focus on designing observers or filters to detect whether state signals have been attacked. Therefore, when attackers use tampering with switching signals to cause asynchronous phenomena in the system, the design of switching signal detection is particularly important. For spoofing attack detection in switched systems, existing coding detection and game theory detection methods do not consider distinguishing between the detection of state signals and switching signals. Therefore, designing a spoofing attack detection mechanism to separately detect whether state signals and switching signals have been tampered with, and then providing corresponding control strategies based on the detection results, presents a considerable challenge. Summary of the Invention

[0004] This invention provides a security control method for a networked handover system based on Q-learning spoofing attack detection, to overcome the technical problems of existing network handover systems that cannot accurately determine whether they are under spoofing attacks and cannot accurately analyze the system status, thus hindering security control.

[0005] To achieve the above objectives, the technical solution of the present invention is as follows:

[0006] A security control method for a networked handover system based on handover Q-learning spoofing attack detection includes:

[0007] S1: Construct the switching system model and controller model to obtain state signals and modal signals;

[0008] S2: Construct an event triggering mechanism based on deception attack detection, the event triggering mechanism including error detection conditions, modality matching conditions and attack detection conditions; use the event triggering mechanism to process the state signal and modality signal to obtain the subsystem modality signal and state signal at the trigger time;

[0009] S3: Construct a deception attack detection mechanism based on switching Q-learning according to the Q-learning algorithm, use the event triggering mechanism and the deception attack detection mechanism to process the subsystem modal signal and state signal at the triggering time, obtain the state estimate, modal estimate, state estimate residual norm and modal estimate residual norm, and determine the deception attack detection result;

[0010] S4: Obtain the controller mode based on the deception attack detection result, and substitute the state estimate and controller mode into the controller model to obtain the security controller equation. Combine the security controller equation, the state estimator of the deception attack detection mechanism, the mode estimator, and the switching system equation to obtain the closed-loop system equation.

[0011] S5: Construct switching rules based on deception attack detection. The switching rules are used to analyze the switching behavior of the closed-loop system under deception attack based on the deception attack detection results, obtain the judgment conditions for system security and stability, and perform security control on the closed-loop system.

[0012] Furthermore, S1 constructs a switching system model and a controller model to obtain state signals and modal signals, including:

[0013] S11. Construct the switching system model, as shown in formula (1).

[0014]

[0015] in, and To switch the system's state vector and control input; It is a modal signal, A σ B ω It is a given constant matrix;

[0016] S12. Construct the controller model, as shown in formula (2).

[0017] u(t) = K φ x(t) (2)

[0018] Where φ = φ(t) ∈ L is the modal information of the controller, K φ This is the controller gain.

[0019] Furthermore, S2 constructs an event-triggered mechanism based on deception attack detection, which includes error detection conditions, modality matching conditions, and attack detection conditions; the event-triggered mechanism is used to process the state signal and modality signal to obtain the subsystem modality signal and state signal at the trigger time, including:

[0020] S21. Construct an event triggering mechanism based on deception attack detection, as shown in formula (3).

[0021]

[0022] Wherein, the initial trigger time t1h = l0, These represent error detection conditions, modality matching conditions, and attack detection conditions, respectively.

[0023] S22. Construct error detection conditions, as shown in formula (4).

[0024]

[0025] in,

[0026]

[0027] Indicates sampling time g r,m The parameter matrix at point h, ε(t) = x(g) r,m h)-x(t r h), g r,m h = t r h+mh represents the r-th trigger time t. r At the m-th sampling time after h, δ σ >0 is the trigger threshold; x(t) r h) represents the subsystem state signal at the trigger time; ε(t) represents the error between the state at the trigger time and the state at the m-th sampling time thereafter; (α) x (g r,m h),α σ (g r,m h)) represents the deception attack detection result to be calculated at the trigger time;

[0028] S23. Construct modal matching conditions, as shown in formula (5).

[0029]

[0030] Wherein, σ(g) r,m h) represents the r-th trigger time t r The modal signal at the m-th sampling time after h, then t r The modal signal at time h is σ(t) r h);

[0031] S24. Construct attack detection conditions, as shown in formula (6).

[0032]

[0033] Furthermore, the deception attack detection mechanism based on switching Q-learning includes state estimation measures, modality estimation measures, and a switching Q-learning method based on Q-learning.

[0034] The state estimation measure is implemented by a state estimator, which is used to calculate the state estimate value of the system state signal in the switching system;

[0035] The modal estimation measure is implemented by a modal estimator, which is used to calculate the modal estimate of the modal signal in the switching system;

[0036] The Q-learning-based switching Q-learning method is used to calculate the difference between the state estimate and the state signal under the spoofing attack, i.e., the state estimate residual norm, and the difference between the modality estimate and the modality signal under the spoofing attack, i.e., the modality estimate residual norm. The spoofing attack detection result is determined based on the state estimate residual norm, the modality estimate residual norm, and the modality estimate.

[0037] Furthermore, the subsystem modal signals and state signals at the triggering time are processed using the event triggering mechanism and the deception attack detection mechanism to obtain state estimates and modal estimates, including:

[0038] S31. Under a deception attack, the modal signals and state signals of the switching system are as shown in formulas (7) and (8).

[0039]

[0040] ε(t) represents the error between the state at the trigger time and the state at the m-th sampling time thereafter, α σ (t) and α x (t) represents the deception attack detection result to be solved, σ a (t) represents the deception value of the switching signal in a spoofing attack, x a (t) represents the deception value of the system state signal in the deception attack;

[0041] S32, The state estimator is as shown in formula (9).

[0042]

[0043] in, It is an estimate of x(t). It is the observer gain. This represents the modal estimate; It is a mode Given the system matrix and input matrix, It is a mode The controller gain to be designed;

[0044] S33, The modal estimator is as shown in formula (10).

[0045]

[0046] in, This represents the estimated state partition corresponding to the i-th state estimator. yes The boundary, and α x (t),α σ (t) represents the deception attack detection result to be solved.

[0047] Furthermore, the Q-learning-based switching Q-learning method includes switching the Q-table training phase and the attack detection phase:

[0048] S34. Switch to Q-table training phase:

[0049] Construct a quintuple (A, S, γ) f ,π f For a Markov decision process of f(k), train the Q-table and update the Q-value, where f∈F={0,1,…,F} is the number of training iterations;

[0050] For the set of environment variables, where It is the r,m-th learning interval Φ in the f-th training round. r,m The modal estimator provides estimates of the subsystem modes. and It is the residual norm of the modal estimation at step r,n in the f-th round of training. and state estimation residual norm in It is in the fth round of training It is in the fth round of training

[0051] Action Set Indicates that the detector is based on environment variables. The detection alarm value, a1=(0,0), indicates that no status signal was detected. With switching signal Under attack, a2 = (0,1) indicates that only [the target] was detected. Under attack, a3 = (1,0) indicates that only [the target] was detected. When attacked, a4 = (1,1) indicates that an attack has been detected. and All were attacked;

[0052] reward function γ f (r,m) represents the environment variables at step r,m in the f-th training round. In performing the action The rewards obtained later, among which

[0053] Strategy This represents the environment variables at step r,m in the f-th training round. Execute action The probability of the discount parameter k∈[0,1] represents the degree of attention to long-term rewards;

[0054] S341. Initialize the actions for the f-th round of training. Environment variables at steps r and m According to strategy Select Action As shown in formula (11),

[0055]

[0056] The environmental variables at step r,m By performing actions Transform to environment variables in step r,m+1

[0057] According to the reward function γ f (r,m) Calculate the action to be executed The reward obtained later is shown in formula (12).

[0058]

[0059] Where c1 and c2 are positive constants. The probability of an attack occurring is known during training;

[0060] S342. Update the Q value according to the reward function and the value of the (f-1)th training round, as shown in formula (13).

[0061]

[0062] in, For the environment variable-action pair during the f-th training round The Q-value after the (f-1)th training round. Environment variables for the f-th round of training The Q-table value α corresponding to the optimal action after the (f-1)th round of training. f=0.1·(2+arctan(f-α) * )) is the learning step size, α * Let be the learning constant, where the initial value of the Q table is .

[0063] S343, Regarding the environment variables at step r,m+1 Repeat the process of S341-S342 to obtain the corresponding Q value until the system stops running, and the training of the fth round ends.

[0064] S344. Repeat steps S341-S343 to perform the next round of training until f = F. Training ends, and the system returns a list containing all environment variable-action pairs. Complete switching of the Q table;

[0065] S35. During the attack detection phase, the corresponding optimal action is selected from the complete switching Q table.

[0066] S351. Determine the environmental variables based on the modal estimates, state estimation residual norms, and modal estimation residual norms in the complete switching Q table:

[0067] Let r = 1, m = 0, when t ∈ Φ r,m At that time, observe the current environmental variables and the current modal estimates.

[0068] S352. Select the optimal action corresponding to the current environment variable from the complete switching Q table.

[0069] S353. Output the deception attack detection results, as shown in formula (14).

[0070] (α x (t),α σ (t))=a r,m ,t∈Φ r,m (14).

[0071] Furthermore, S4 obtains the controller mode based on the deception attack detection result, and substitutes the state estimate and controller mode into the controller model to obtain the security controller equation. By simultaneously solving the security controller equation, the state estimator and mode estimator of the deception attack detection mechanism, and the switching system equation, a closed-loop system equation is obtained, including:

[0072] Substituting the state estimate and the modal estimate into the controller model, the safety controller equation is obtained, as shown in formula (16).

[0073]

[0074] φ = φ(t), φ(t) ∈ L, representing the modal information of the controller, as shown in formula (16).

[0075]

[0076] in, This represents the modal transmission value, or modal signal, calculated based on the results of spoofing attack detection. This represents the modal estimate obtained through the modal estimator;

[0077] By simultaneously solving the safety controller equation, the state estimator equation, the modality estimator equation, and the switching system equation, the closed-loop system equation is obtained, as shown in formula (17).

[0078]

[0079] in,

[0080] Extended state vector of closed-loop system Extended System Matrix System parameter matrix System parameter matrix Initial state vector Initial time Subsystem and controller modes

[0081] Furthermore, S5 constructs switching rules based on spoofing attack detection. These switching rules are used to analyze the switching behavior of the closed-loop system under spoofing attacks based on the spoofing attack detection results, obtain judgment conditions for system security and stability, and perform security control on the closed-loop system, including:

[0082] S51. Construct a switching rule based on deception attack detection, as shown in formula (18).

[0083]

[0084] D i boundary and

[0085] S52. Analyze the deception attack detection results according to the switching rules and the event triggering mechanism, determine the operating status of the subsystem mode and the controller mode under the deception attack, obtain the judgment conditions for system security and stability, and use the judgment conditions to perform security control on the switching system.

[0086] This invention establishes a deception attack detection mechanism based on switching Q-learning, which obtains the optimal evaluation threshold to determine whether the state signals and modal values ​​transmitted by the system to the controller have been tampered with by a deception attack. It also provides deception attack detection results for the designed switching rules and event triggering mechanisms. By designing an event triggering mechanism based on deception attack detection, triggering is arranged according to error conditions based on the alarm value of the state signal when the signal has not been tampered with, thereby reducing the updating of erroneous data. Timely data updates are performed after subsystem switching, mitigating the asynchronous duration at the subsystem switching point. Timely data updates after the deception attack ends shorten the asynchronous duration between the subsystem and the controller. Switching rules based on deception attack detection are designed to mitigate multiple consecutive asynchronous operations between the subsystem and the controller caused by deception attacks. Finally, based on the event triggering mechanism and switching rules, a security control scheme for the switching system under deception attacks is constructed. Based on the attack detection results, security control judgment conditions are given to ensure the switching system's security under deception attacks, thus achieving secure control of the switching system. Attached Figure Description

[0087] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0088] Figure 1 This is a flowchart illustrating the security control method for a networked handover system based on handover Q-learning spoofing attack detection, as described in this invention.

[0089] Figure 2 This is a block diagram of the switching system structure of the present invention;

[0090] Figure 3 This diagram illustrates three scenarios for the switching signal in the closed-loop system of this invention.

[0091] Figure 4 This is a topology diagram of a switching system according to an embodiment of the present invention;

[0092] Figure 5 This is a trajectory curve diagram on the X and Y axes of an embodiment of the present invention;

[0093] Figure 6 This is a tracking and positioning error diagram of multiple unmanned vessels according to an embodiment of the present invention;

[0094] Figure 7 This is a graph showing the switching signal versus Lyapunov variation in one embodiment of the present invention.

[0095] Figure 8An attack detection result diagram for an embodiment of the invention that learns and estimates residuals from fixed states;

[0096] Figure 9 This is an event triggering interval diagram for a specific embodiment of the present invention, showing the presence / absence of deception attack detection.

[0097] Figure 10 This is an asynchronous duration diagram under the switching rule of having / not having deception attack detection in one embodiment of the present invention. Detailed Implementation

[0098] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0099] This embodiment provides a security control method for a networked handover system based on handover Q-learning spoofing attack detection, such as... Figure 1 As shown, it includes:

[0100] S1: Construct the switching system model and controller model to obtain state signals and modal signals;

[0101] S2: Construct an event triggering mechanism based on deception attack detection, the event triggering mechanism including error detection conditions, modality matching conditions and attack detection conditions; use the event triggering mechanism to process the state signal and modality signal to obtain the subsystem modality signal and state signal at the trigger time;

[0102] S3: Construct a deception attack detection mechanism based on switching Q-learning according to the Q-learning algorithm, use the event triggering mechanism and the deception attack detection mechanism to process the subsystem modal signal and state signal at the triggering time, obtain the state estimate, modal estimate, state estimate residual norm and modal estimate residual norm, and determine the detection result of the deception attack;

[0103] S4: Obtain the controller mode based on the deception attack detection result, and substitute the state estimate and controller mode into the controller model to obtain the security controller equation. Combine the security controller equation, the state estimator of the deception attack detection mechanism, the mode estimator, and the switching system equation to obtain the closed-loop system equation.

[0104] S5: Construct switching rules based on deception attack detection. The switching rules are used to analyze the switching behavior of the closed-loop system under deception attack based on the deception attack detection results, obtain the judgment conditions for system security and stability, and perform security control on the closed-loop system.

[0105] Specifically, a switching system model and a controller model are constructed to obtain state signals and modal signals, thus establishing an initial switching system and controller model to prepare for subsequent steps. Secondly, an event-triggered mechanism based on spoofing attack detection is constructed, including error detection conditions, modal matching conditions, and attack detection conditions. The event-triggered mechanism is used to process the state signals and modal signals to obtain the subsystem modal signals and state signals at the trigger time, reducing erroneous data updates, alleviating asynchronous duration at subsystem switching points, and shortening the asynchronous duration between the subsystem and the controller. Based on the Q-learning algorithm, a spoofing attack detection mechanism based on switching Q-learning is constructed. The event-triggered mechanism and the spoofing attack detection mechanism are used to process the subsystem modal signals and state signals at the trigger time to obtain state estimates, modal estimates, state estimation residual norms, and modal estimation residual norms, determining the detection result of the spoofing attack. Compared with the common method of setting a fixed evaluation threshold to detect attacks in switching systems, the attack detection mechanism based on switching Q-learning can obtain the optimal evaluation threshold. This allows for a more intelligent determination of whether a deception attack has occurred or the system is undergoing a normal switch, thereby reducing false alarms and improving the reliability of security control. Based on the deception attack detection results, the controller mode is obtained, and the state estimate and controller mode are substituted into the controller model to obtain the security controller equation. The security controller equation, the state estimator and mode estimator of the deception attack detection mechanism, and the switching system equation are combined to obtain the closed-loop system equation. The closed-loop system includes the state estimator and mode estimator, which, in conjunction with the attack detection algorithm, can effectively defend against deception attacks and reduce false alarms caused by normal system switching. Finally, a switching rule based on deception attack detection is constructed. This switching rule can determine whether the current state signal and the switching signal are historical signals replayed from a deception attack, keeping the currently active subsystem mode unchanged or switching to a subsystem mode that meets the state switching conditions. This alleviates the continuous asynchronous operation between the subsystem and subcontroller caused by deception attacks, determines whether a subsystem switch is needed, and ensures system stability.

[0106] A block diagram of a switching system according to an embodiment of the present invention is shown below. Figure 2 As shown,

[0107] In a specific embodiment, the scheme for constructing the switching system model and controller model to obtain state signals and modal signals is as follows:

[0108] S11. Construct the switching system model, as shown in formula (19).

[0109]

[0110] in, and To switch the system's state vector and control input; It is a modal signal, A σ B σ It is a given constant matrix with appropriate dimensions;

[0111] S12. Construct the controller model, as shown in formula (20).

[0112] u(t) = K φ x(t) (20)

[0113] Where φ = φ(t), φ(t) ∈ L are the modal information of the controller, and K φ This is the controller gain.

[0114] This approach establishes an initial switching system and controller model to prepare for subsequent steps.

[0115] In a specific embodiment, an event-triggered mechanism based on deception attack detection is constructed. This mechanism includes error detection conditions, modality matching conditions, and attack detection conditions. The scheme for processing the state signal and modality signal using this event-triggered mechanism to obtain the subsystem modality signal and state signal at the trigger moment is as follows:

[0116] S21. Construct an event triggering mechanism based on deception attack detection, as shown in formula (21).

[0117]

[0118] Wherein, the initial trigger time t1h = l0, These represent error detection conditions, modality matching conditions, and attack detection conditions, respectively.

[0119] S22. Construct error detection conditions, as shown in formula (22).

[0120]

[0121] in,

[0122] Indicates sampling time g r,m The parameter matrix at h, ε r (t)=x(g r,m h)-x(t r h), g r,m h = t r h+mh represents the r-th trigger time t.r At the m-th sampling time after h, δ σ >0 is the trigger threshold; x(t) r h) represents the subsystem state signal at the trigger time; ε(t) represents the error between the state at the trigger time and the state at the m-th sampling time thereafter; (α) x (g r,m h),α σ (g r,m h)) represents the deception attack detection result to be calculated at the trigger time;

[0123] S23. Construct modal matching conditions, as shown in formula (23).

[0124]

[0125] Wherein, σ(g) r,m h) represents the r-th trigger time t r The modal signal at the m-th sampling time after h, then t r The modal signal at time h is σ(t) r h);

[0126] S24. Construct attack detection conditions, as shown in formula (24).

[0127]

[0128] This solution constructs an event triggering mechanism based on deception attack detection. Under error detection conditions, triggering is arranged according to the error conditions based on the alarm value of the status signal when the signal has not been tampered with, thereby reducing the updating of erroneous data. Under modal matching conditions, data is updated in a timely manner after subsystem switching based on the alarm value of the status signal, thereby alleviating the asynchronous time at the subsystem switching point. Under attack detection conditions, data is updated in a timely manner after the deception attack ends based on the alarm values ​​of the status signal and the switching signal, shortening the asynchronous time between the subsystem and the controller.

[0129] In a specific embodiment, a spoofing attack detection mechanism based on switching Q-learning is constructed according to the Q-learning algorithm. The event triggering mechanism and the spoofing attack detection mechanism are used to process the subsystem modal signals and state signals at the triggering time to obtain state estimates, modal estimates, state estimation residual norms, and modal estimation residual norms. The scheme for determining the detection result of the spoofing attack is as follows:

[0130] S31. Construct a deception attack detection mechanism based on switching Q-learning, wherein the deception attack detection mechanism based on switching Q-learning includes state estimation measures, mode estimation measures, and a switching Q-learning method based on Q-learning:

[0131] Due to transmission delays in the communication network, the data x(t) transmitted at the trigger time... r h) and σ(t) r h) After passing through the network, The time arrives at the sub-controller, in This represents the transmission delay of the data triggered for the rth time, and therefore the signal's hold interval. It can be divided into Where Φ r,m The subinterval is represented as shown in formula (25).

[0132]

[0133] Represents the maximum number of subintervals, and the piecewise time delay function η. r (t)=tg r,m h satisfies the condition

[0134] S31. Under a deception attack, the modal signals and state signals of the switching system are as shown in formulas (26) and (27).

[0135]

[0136] ε(t) represents the error between the state at the trigger time and the state at the m-th sampling time thereafter, α σ (t) and α x (t) represents the deception attack detection result to be solved, σ a (t) represents the deception value of the switching signal in a spoofing attack, x a (t) represents the deception value of the system state signal in the deception attack;

[0137] S32. The state estimation measure is implemented by a state estimator, which is used to calculate the state estimate of the system state signal in the switching system. The state estimator equation is shown in formula (28).

[0138]

[0139] in, It is an estimate of x(t). It is the observer gain. This represents the modal estimate; It is a mode Given the system matrix and input matrix, It is a mode The controller gain to be designed;

[0140] S33. The modal estimation measure is implemented by a modal estimator, which is used to calculate the modal estimate of the modal signal in the switching system. The modal estimator equation is shown in formula (29).

[0141]

[0142] in, This represents the estimated state partition corresponding to the i-th state estimator. yes The boundary, and

[0143] This scheme is based on a spoofing attack detection mechanism using Q-learning. It uses state and mode estimators to calculate the estimated residuals of the system state and the estimated residuals of the system modes. These residual values ​​are used as environmental variables for the Q-learning algorithm to obtain the optimal evaluation threshold. This allows the scheme to determine whether the state signals and mode values ​​transmitted to the controller by the system have been tampered with by a spoofing attack, and to provide spoofing attack detection results for the switching rules and event triggering mechanisms to be designed.

[0144] In a specific embodiment, the Q-learning-based switching Q-learning method is used to calculate the difference between the state estimate and the state signal under a spoofing attack, i.e., the state estimate residual norm; the difference between the mode estimate and the mode signal under a spoofing attack, i.e., the mode estimate residual norm; and the spoofing attack detection result is determined based on the state estimate residual norm, the mode estimate residual norm, and the mode estimate, including a switching Q-table training phase and an attack detection phase.

[0145] S34. Switch to Q-table training phase:

[0146] Construct a quintuple (A, S, γ) f ,π f For a Markov decision process of f(k), train the Q-table and update the Q-value, where f∈F={0,1,…,F} is the number of training iterations;

[0147] For the set of environment variables, where It is the r,m-th learning interval Φ in the f-th training round. r,m The modal estimator provides estimates of the subsystem modes. and It is the residual norm of the modal estimation at step r,m in the f-th training round. and state estimation residual norm in It is in the fth round of training It is in the fth round of training

[0148] Action Set Indicates that the detector is based on environment variables. The detection alarm value, a1 = (0,0) indicates that no detection was detected. and Under attack, a2 = (0,1) indicates that only [the target] was detected. Under attack, a3 = (1,0) indicates that only [the target] was detected. When attacked, a4 = (1,1) indicates that an attack has been detected. and All were attacked;

[0149] reward function γ f (r,m) represents the environment variables at step r,m in the f-th training round. In performing the action The rewards obtained later, among which

[0150] Strategy This represents the environment variables at step r,m in the f-th training round. Execute action The probability of the discount parameter k∈[0,1] represents the degree of attention to long-term rewards;

[0151] S341. Initialize the actions for the f-th round of training. Environment variables at steps r and m According to strategy Select Action As shown in formula (30),

[0152]

[0153] Specifically, generate a random number c∈[0,1], if So otherwise in Exploration rate β * To explore constants;

[0154] The environmental variables at step r,m By performing actions Transform to environment variables in step r,m+1 According to the reward function γ f (r,m) Calculate the action to be executed The reward obtained later is shown in formula (31).

[0155]

[0156] Where c1 and c2 are positive constants. The probability of an attack occurring is known during training;

[0157] S342. Update the Q value according to the reward function and the value of the (f-1)th training round, as shown in formula (32).

[0158]

[0159] in, For the environment variable-action pair during the f-th training round The Q-value after the (f-1)th training round. Environment variables for the f-th round of training The Q-table value α corresponding to the optimal action after the (f-1)th round of training. f =0.1·(2+arctan(f-α) * )) is the learning step size, α * Let be the learning constant, where the initial value of the Q table is .

[0160] S343, Regarding the environment variables at step r,m+1 Repeat the process of S341-S342 to obtain the corresponding Q value until the system stops running, and the training of the fth round ends.

[0161] S344. Repeat steps S341-S343 to perform the next round of training until f = F. Training ends, and the system returns a list containing all environment variable-action pairs. Complete switching of the Q table;

[0162] S35. During the attack detection phase, the corresponding optimal action is selected from the complete switching Q table.

[0163] S351. Determine the environmental variables based on the modal estimates, state estimation residual norms, and modal estimation residual norms in the complete switching Q table:

[0164] Let r = 1, m = 0, when t ∈ Φ r,m At that time, observe the current environmental variables and the current modal estimates.

[0165] S352. Select the optimal action corresponding to the current environment variable from the complete switching Q table.

[0166] S353. Output the deception attack detection results, as shown in formula (33).

[0167] (α x (t),α σ (t))=a r,m,t∈Φ r,m (33)

[0168] This embodiment constructs a handover Q-learning method for detecting spoofing attacks based on Q-learning. Compared with the common method in handover systems that uses a fixed evaluation threshold to detect attacks, the attack detection mechanism based on handover Q-learning can more intelligently determine whether a spoofing attack has occurred. Compared with the intelligent algorithms used in attack detection in non-handover systems, this mechanism uses a Q-table containing modal information to more accurately detect attacks with the same estimated residual under different modalities.

[0169] In a specific embodiment, the controller mode is obtained based on the deception attack detection result, and the state estimate and controller mode are substituted into the controller model to obtain the security controller equation. The scheme for simultaneously solving the security controller equation, the state estimator of the deception attack detection mechanism, the mode estimator, and the switching system equation to obtain the closed-loop system equation is as follows:

[0170] Substituting the state estimate and the modal estimate into the controller model, the safety controller equation is obtained, as shown in formula (34).

[0171]

[0172] φ = φ(t), φ(t) ∈ L, representing the modal information of the controller, as shown in formula (35).

[0173]

[0174] in, This represents the modal transmission value, or modal signal, calculated based on the results of spoofing attack detection. This represents the modal estimate obtained through the modal estimator;

[0175] By simultaneously solving the safety controller equation, the state estimator equation, the modality estimator equation, and the switching system equation, the closed-loop system equation is obtained, as shown in formula (36).

[0176]

[0177] in,

[0178] Extended state vector of closed-loop system Extended System Matrix System parameter matrix System parameter matrix Initial state vector Initial time Subsystem and controller modes

[0179] This solution constructs a closed-loop system, which includes a state estimator and a modality estimator. It can work with attack detection algorithms to effectively defend against deception attacks and reduce false alarms caused by normal system switching.

[0180] In a specific embodiment, a switching rule based on deception attack detection is constructed. This switching rule is used to analyze the switching behavior of the closed-loop system under a deception attack based on the detection results, obtain the judgment conditions for system security and stability, and implement a security control scheme for the closed-loop system as follows:

[0181] S51. Construct a switching rule based on deception attack detection, as shown in formula (37).

[0182]

[0183] D i boundary and

[0184] S52. Analyze the deception attack detection results based on the switching rules and the event triggering mechanism, determine the operating status of the subsystem mode and controller mode under the deception attack, obtain the judgment conditions for system security and stability, and use the judgment conditions to perform security control on the switching system:

[0185] To facilitate the analysis of handover behavior, it is assumed that the j-th and i-th subsystems are in adjacent handover intervals [l q-1 ,l q ) and [l q ,l q+1 When activated within the subsystem, the subsystem mode σ may be tampered with due to the attack, and the subcontroller will adopt different modes φ according to formula (33). This depends on whether the attack covers the subsystem l. q or sub-controller switching time And different attack durations, any interval [l q ,l q+1 The changes in the upper modal signals σ and φ can be divided into Figure 3 The three cases shown:

[0186] In this embodiment, the network switching is affected by a spoofing attack, and the activation interval H of the p-th spoofing attack is defined. p =[h p ,h p +d p ) and dormancy zone Where h p Indicates the moment the attack was initiated, d pThis indicates the duration of the attack. The attacker injects an attack signal, v(t), into the data transmitted over the network. a (t), which is then altered to Where v∈{x,σ}, if t∈H p α v (t) = 1, otherwise α v (t) = 0. This attack also requires the following assumptions to be satisfied:

[0187] Assumption 1 (Attack Energy): For a deception attack, the attack signal x is injected into the system state signal x(t) and the modal signal σ(t). a (t), σ a (t) respectively satisfy ||x a (t)|| 2≤ ||Gx(t)||2, The given constant matrix G is determined by the characteristics of the deception attack;

[0188] Assumption 2 (Attack Frequency): There exists w p ≥0 and τ p ≥h, such that in For in the interval Number of internal deception attacks;

[0189] Assumption 3 (Attack Duration): For H p and Positive parameters exist and s Each satisfies and

[0190] Theorem 1: Under Assumptions 1-3, for a given positive constant h, η , parameter And i≠j, The closed-loop system is asymptotically stable under the switching Q-learning attack detection mechanism if the following conditions are met.

[0191] (1) There exists a matrix S 0 >0, S 1 >0, R 0 >0, R 1 >0, M and M satisfy formulas (38) and (39),

[0192]

[0193] in It consists of the following block matrices:

[0194]

[0195] (2) There exists an indeterminate matrix of appropriate dimension. Satisfying formulas (40), (41) and (42),

[0196]

[0197]

[0198] in It consists of the following block matrices:

[0199]

[0200] Proof: To achieve the control objective, we prove that the closed-loop system is asymptotically stable. First, we select candidate Lyapunov functions, as shown in equation (43).

[0201]

[0202] in,

[0203]

[0204] Calculate the time derivative of formula (43) along the closed-loop system and take the expected value to obtain...

[0205]

[0206] right Applying Jensen's integral inequality, Park's lemma, and formula (39) to the integral terms in the equations, we obtain formulas (44) and (45).

[0207]

[0208] in,

[0209] Will and Substituting into formula (37), we obtain the inequality.

[0210]

[0211] Substituting the inequalities and formulas (44) and (45) into Based on formula (38), we obtain formula (46).

[0212]

[0213] in,

[0214] Based on the following conditions, the changes in subsystem mode σ and controller mode φ in AC analysis are discussed, and the equation (43) is applied in any interval [l]. q ,l q+1 Changes within )

[0215] Case A: In the interval Above, (σ,φ)=(i,j), in the interval Above, (σ,φ)=(i,i), in the interval Above, (σ,φ)=(i,φ) p ), in the interval Above, (σ,φ)=(i,i), in the interval Above, (σ,φ)=(i,i);

[0216] By combining the modal estimator formula (29), the switching rule formula (37), and the formula (40), it can be derived that under case A, formula (43) is in the switching interval [l q ,l q+1 The evolution within ) is shown in formula (47).

[0217]

[0218] Case B: In the interval Above, (σ,φ)=(i,j), in the interval Above, (σ,φ)=(i,φ) p ), in the interval Above, (σ,φ)=(i,i);

[0219] By combining the modal estimator formula (29), the switching rule formulas (37) and (40), it can be derived that under case B, formula (43) is in the switching interval. The evolution within is shown in formula (48).

[0220]

[0221] Case C: In the interval Above, (σ,φ)=(i,j), in the interval Above, (σ,φ)=(i,i),

[0222] By combining the modal estimator formula (29), the switching rule formula (37), and the formula (40), it can be derived that under case C, formula (43) is in the switching interval [l q ,l q+1 The evolution within ) is as follows:

[0223] Since the alarm value for situation C is (α) x (t),α σ (t))≠(1,1), from which formula (49) can be derived.

[0224]

[0225] Therefore, the result obtained through formula (42) For case AC, formula (43) applies to any interval [l q ,l q+1 The changes in ) can be summarized by formula (50).

[0226]

[0227] Discussion of subsystem switching point l q Changes in formula (43) before and after:

[0228] Under the attack detection-based switching rule, formula (51) is obtained.

[0229]

[0230] Discuss V(t) and its operation on the interval [0,t). Relational derivation:

[0231] Based on formula (41) and formulas (50) and (51), we can obtain formula (52).

[0232]

[0233] This proves that the closed-loop system is asymptotically stable;

[0234] Formula (38) in Theorem 1 and formula (40) Linearization is performed on the nonlinear terms:

[0235] Theorem 2: Under Assumptions 1-3, for a given positive constant h, η , w and parameter And i≠j, If formulas (39), (41), and (42) in Theorem 1 hold, and there exists a matrix S 0 >0, S 1 >0, R 0 >0, R 1 >0, and the opposing matrix M 1 Satisfying formulas (53) and (54),

[0236]

[0237] The closed-loop system then becomes asymptotically stable, where equation (53) states... It consists of the following matrices:

[0238]

[0239] In formula (54) It consists of the following block matrices:

[0240] The other block matrices in the matrix are zero matrices of appropriate dimension;

[0241] Proof: First, multiply both sides of formula (38) by a matrix. and its transpose, in which It can be obtained that consists of 6×6 matrix blocks

[0242]

[0243] Multiply the left and right sides of formula (40) by a matrix and its transpose, in which It can be obtained that consists of 4×4 matrix blocks Applying Schur's complement lemma, we obtain -[(4-α x )R] -1 <w 2 (4-α x )R-2wI,[5α x R] -1 <-5w 2 α x R-2wI, thus having and Equivalent to formulas (38) and (40); with the help of Schur's complement lemma, formulas (38) to (40) can be derived from formulas (53) and (54), proving the stability of the closed-loop system.

[0244] This scheme constructs a switching rule based on deception attack detection. When the state signal is not tampered with by a deception attack, the system switches to the target subsystem mode; when the state signal is tampered with by a deception attack, the system maintains the currently active subsystem mode. Compared with switching rules that do not consider attack detection results, this switching rule can mitigate the multiple consecutive asynchronous operations between the subsystem and the controller caused by deception attacks. Treating the attack energy as a special external input, the scheme analyzes the system changes under AC conditions to reveal the relationship between the internal energy and external input of the switching system under deception attacks, describes the impact of deception attacks on the overall stability of the switching system, and constructs security conditions to ensure the system operates under these conditions, achieving secure control under deception attacks.

[0245] Simulation is performed using an unmanned surface vessel as an example:

[0246] The effectiveness of the method is verified using a case study of dynamic positioning control of unmanned surface vessels (USVs) on multiple water surfaces. The dynamics of the USV can be expressed as formula (55).

[0247]

[0248] in and Indicates the X-axis position and Y-axis position of the nth ship. Let n represent the ship's heading angle, where n ∈ {1, 2, 3}. Indicates the oscillation velocity. Indicates sway speed, This represents the yaw rate. u n To control the input, M n ,∏ n and for

[0249]

[0250] Where m n x n,g and I n,z Let represent the mass of the nth ship, its center of gravity on the X-axis, and its moment of inertia on the heading, respectively. It is the added mass in the longitudinal direction. and Indicates the additional mass in the sway direction. The additional moment of inertia in the heading direction, and The damping coefficient is linear.

[0251] Unmanned surface vessel (USV) 1 acts as the leader, responsible for receiving location information. USVs 2 and 3, acting as followers, maintain coordinated tracking and positioning by receiving relative position information from USV 1.

[0252] Specifically, let the state of the nth unmanned vessel be... In formula (55), the dynamic positioning objective of the unmanned vessel is to adjust the vessel's position and heading so that... in Represents the expected formation vector. For a single desired location point, ρ n For preset scalar;

[0253] The relevant parameters for the unmanned surface vessel are:

[0254] The initial state of the three unmanned boats is as follows: Other parameters are shown in Table 1.

[0255] Table 1 Parameters of the three unmanned vessels

[0256]

[0257] Figure 4 The simulation verification of the safety control method uses a switching system topology. Due to the limited communication channels of unmanned surface vessel 1, it can only communicate with one accompanying vessel at a time, and accompanying vessels 2 and 3 cannot communicate directly with each other. Specifically, when unmanned surface vessel 1 sends information only to unmanned surface vessel 2, it is topology 1, as shown below. Figure 4 As shown on the left. On the other hand, when unmanned surface vessel 1 only sends information to unmanned surface vessel 3, it is topology 2, as shown... Figure 4 As shown on the right. Since only one follower ship can receive information at a time, the other follower ship that is not being communicated with may deviate from the queue, making each topology unstable when operating alone. To ensure that the unmanned ships can achieve cooperative tracking and positioning, it is necessary to switch between the two topologies;

[0258] based on Figure 4 Due to changes in the topology, the cooperative localization problem of the three ships can be modeled as a switching system containing two unstable subsystems, as shown in equation (56).

[0259]

[0260] in In the formula K n For the gain of the controller to be designed, The communication topology matrix is ​​σ∈{1,2}, and the other coefficient matrices are as follows.

[0261]

[0262] The control objective of formula (55) is then transformed into the stabilization problem of formula (56);

[0263] Select a deception attack that satisfies Assumptions 1-3: The input parameters of the algorithm are: discount parameter κ = 0.42, training iterations F = 50, action set A = {(0,0),(0,1),(1,0),(1,1)}, and exploration constant β. * =3, learning constant α * =5, in the reward function c1=10, c2=5; partition parameters of the state estimation residual norm: θ0=0, θ1=0.05, θ2=0.1, θ3=0.11, θ4=0.12, θ5=0.13, θ6=0.14, θ7=0.15; partition parameters of the modal estimation residual norm: The training data consists of subsystem modes generated during system runtime. State estimation residual norm and modal estimation residual norm choose η = 0.003, δ1=δ2=0.1, n 12 =n 21 =-1, w=0.9, solving Theorem 2 yields the controller gains for the three ships.

[0264]

[0265]

[0266] In this embodiment, the position and state of the unmanned vessel are as follows: Figures 5-10 As shown,

[0267] Figure 5 and Figure 6 The image shows that the three unmanned vessels have reached the desired location.

[0268] Figure 7 The description of the switching signal and V(t) change curves under the switching rule based on attack detection shows that when (0,1) and (1,0) type attacks occur, the switching rule keeps the current active subsystem mode unchanged.

[0269] Figure 8 A switching Q-learning method is presented, that is, switching the deception attack detection result (α) output by the Q-learning algorithm. x ,α σ Table 2 shows the attack detection results comparing the learned and fixed-state estimated residuals. Specifically, in the switching Q-learning algorithm for detecting deception attacks in this invention, the state residuals are divided into {θ0,…,θ7}, and the modal residuals are divided into… The training process consisted of 50 rounds; a fixed evaluation threshold was selected based on the minimum non-zero state residual. intermediate item partitioning with maximum state residuals Selecting non-zero modal residuals for partitioning The comparison results are shown in Table 2;

[0270] Table 2 Attack detection results of the learned and fixed-state estimation residuals (T=60s)

[0271]

[0272] It can be seen that as the set threshold gradually increases, both the detection accuracy and the false alarm rate show a downward trend. The switching Q-learning algorithm can effectively control the false alarm rate at a low level while ensuring high detection accuracy. On the other hand, by providing more accurate detection results for the switching rules and event triggering mechanism, the switching Q-learning algorithm reduces the asynchronous duration of the switching rules and the number of times erroneous data is triggered under attack. Figure 8 The switching Q-learning algorithm and the fixed estimation residual detection algorithm, which has the lowest false alarm rate, were compared separately. The test results α x (t) and α σ (t).

[0273] Figure 9 The triggering interval of the attack detection-based event triggering mechanism is described. In the figure, the purple shaded area represents (α). x ,α σ () = (0,1) type deception attack, pink shading is (α) x ,α σ (1,0) type deception attack, yellow shading is (α) x ,α σ () = (1,1) type deception attack. From Figure 9 It can be seen that, due to the introduction of attack detection results, the attack is triggered normally when a (0,1) type attack occurs within 5s-7s and 50.1s-51.05s, but is not triggered when a (1,0) or (1,1) type attack occurs within 11.24s-13.79s, 19.98s-22.45s, 29.09s-31.32s, 35.7s-36.81s, and 39.84s-41.5s.

[0274] Table 3 and Figure 9 The number of event triggers with and without deception attack detection is compared with the switching point and the asynchronous duration after the attack ends. This invention, based on the event triggering mechanism for deception attack detection, introduces the attack detection result (α) into each triggering condition. x (t),α σ (t)); the multi-condition event triggering mechanism in comparison was not introduced (α) x (t),α σ(t)), for ease of comparison, the memory step size is set to 0, at which point the triggering mechanism becomes static triggering. The comparison results are shown in Table 3 and Figure 9 As shown,

[0275] Table 3. Number of events triggered and asynchronous duration (T=60s) with / without spoofing attack detection

[0276] Event triggering mechanism Trigger count Switching point and asynchronous duration after attack ends Event-triggered mechanism based on attack detection 128 3.46s Event triggering mechanism without attack detection 138 5.10s

[0277] It can be seen that the introduction of (α) x (t),α σ The event triggering mechanism of (t) can optimize the number of data triggers and effectively reduce the asynchronous duration after the switching point and the end of the attack.

[0278] Figure 10 Table 4 compares the asynchronous duration under different switching rules with... Figure 10 The asynchronous duration and the number of consecutive asynchronous switches were compared under switching rules with and without deception attack detection. The switching rule based on deception attack detection introduces the attack detection result (α). x (t),α σ (t)); The switching rules without attack detection, except for the lack of introduction of (α) x (t),α σ Except for (t)), all other switching conditions are consistent with the switching rules.

[0279] Table 4. Asynchronous duration and number of consecutive asynchronous switches with / without spoofing attack detection (T=60s)

[0280] Switching rules Total asynchronous duration Number of consecutive asynchronous switching Switching rules based on attack detection 10.07s 0 Switching rules without attack detection 12.58s 3

[0281] It can be seen that, under the same deception attack, the introduction of (α) x (t),α σ The switching rules of (t) effectively shorten the asynchronous time caused by deception attacks and effectively avoid multiple consecutive asynchronous switching.

[0282] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A security control method for a networked switching system based on switching Q-learning deception attack detection, characterized in that, The application relates to a method for constructing a secure controller of a switching system under deception attack. The method comprises the following steps: S1: constructing a switching system model and a controller model to obtain state signals and mode signals; the specific steps are as follows: (1) wherein and is a state vector of the switching system and the control input; is a modal signal, , is a given constant matrix; S11: constructing a switching system model, as shown in formula (1), (2) wherein, is modal information of the controller, is a controller gain; S12: constructing a controller model, as shown in formula (2), S2: constructing an event-triggering mechanism based on deception attack detection, wherein the event-triggering mechanism comprises an error detection condition, a mode matching condition and an attack detection condition; the state signals and the mode signals are processed by using the event-triggering mechanism to obtain the mode signals and the state signals of the subsystem at the triggering moment; the specific steps comprise: (3) wherein the initial triggering moment , , , respectively represent the error detection condition, the modal matching condition and the attack detection condition; S21: constructing an event-triggering mechanism based on deception attack detection, as shown in formula (3), (4) S22: constructing an error detection condition, as shown in formula (4), denotes the parameter matrix at the sampling instant , , is the sampling instant after m , is a triggering threshold; denotes a subsystem status signal at the triggering instant; denotes the error between the status at the triggering instant and the status at the m-th sampling instant after the triggering instant; denotes the result of the fraud attack detection to be computed at the triggering instant; wherein, (5) wherein, represents the modality signal at the sampling time point m ; and the modality signal at the sampling time point ; and S23: constructing a mode matching condition, as shown in formula (5), (6); S24: constructing an attack detection condition, as shown in formula (6), S3: constructing a deception attack detection mechanism based on switching Q learning according to a Q learning algorithm; the mode signals and the state signals of the subsystem at the triggering moment are processed by using the event-triggering mechanism and the deception attack detection mechanism to obtain state estimation values, mode estimation values, state estimation residual norm and mode estimation residual norm, and to determine a deception attack detection result; S4: obtaining a controller mode according to the deception attack detection result, and substituting the state estimation values and the controller mode into the controller model to obtain a safe controller equation; the safe controller equation, the state estimator and the mode estimator of the deception attack detection mechanism and the switching system equation are combined to obtain a closed-loop system equation; 2. The security control method of networked handoff system based on handoff Q-learning spoofing attack detection according to claim 1, characterized in that, S5: constructing a switching rule based on deception attack detection, wherein the switching rule is used for analyzing the switching behavior of the closed-loop system under deception attack according to the deception attack detection result, obtaining a judgment condition of system safety and stability, and performing safety control on the closed-loop system. The deception attack detection mechanism based on switching Q learning comprises a state estimation measure, a mode estimation measure and a switching Q learning method based on Q learning; The state estimation measure is realized by a state estimator and is used for calculating the state estimation values of the state signals of the switching system; The mode estimation measure is realized by a mode estimator and is used for calculating the mode estimation values of the mode signals of the switching system; 3. The security control method of networked handoff system based on handoff Q-learning spoofing attack detection according to claim 2, characterized in that, The switching Q learning method based on Q learning is used for calculating the state estimation residual norm, which is the difference between the state estimation values and the state signals under deception attack, and the mode estimation residual norm, which is the difference between the mode estimation values and the mode signals under deception attack; the deception attack detection result is determined according to the state estimation residual norm, the mode estimation residual norm and the mode estimation values. The mode signals and the state signals of the switching system under deception attack are as shown in formulas (7) and (8), (7) (8) denotes an error between the state at the triggering time instant and the state at the mth sampling time instant thereafter, denote the modal signal spoofing attack detection result to be solved and the state signal spoofing attack detection result to be solved, respectively, denotes a spoofing value of the switching signal by the spoofing attack, denotes a spoofing value of the system state signal by the spoofing attack; The state estimator is as shown in formula (9), (9) in, yes The estimated value, It is the observer gain. , representing the modal estimate; It is a mode Given the system matrix and input matrix, It is a mode The controller gain to be designed The mode estimator is as shown in formula (10), (10) wherein, represents the estimation state partition corresponding to the i state estimator of the k-th state estimator, , , is the boundary of, and , denotes the fraud attack detection result to be solved.

4. The security control method of networked handoff system based on handoff Q-learning spoofing attack detection according to claim 3, characterized in that, The switching Q-learning method based on Q-learning includes a switching Q-table training phase and an attack detection phase: S34, switching Q-table training phase: Construct a Markov decision process training Q-table containing a quintuple Update the Q-value, where is the number of training times; is the set of environment variables, where is the f-th training round, f is the i-th learning interval in the f-th training round, is the j-th learning interval in the i-th learning interval in the f-th training round, is the estimate of the subsystem modal by the upper modal estimator, and is the f-th training round, f is the i-th learning interval in the f-th training round, is the j-th learning interval in the i-th learning interval in the f-th training round, is the upper modal estimation residual norm in the i-th learning interval in the f-th training round, is the state estimation residual norm in the i-th learning interval in the f-th training round, is the f-th training round, , is the f-th training round, ; set of actions represents that the detector based on the environmental variable of the detection alarm value, represents that no state signal with the switching signal is attacked, represents that only is attacked, represents that only is attacked, represents that and are both attacked; reward function Indicates the first f In the first round of training, Step environment variables In performing the action The rewards obtained later, among which , ; Policy representing in the first f round of training, the first step environment variable probability of performing an action discount parameter representing the degree of focus on long-term rewards; S341. Initialize actions of the f-th round of training ; the f-th step environment variable According to the strategy Select actions As shown in equation (11), (11) the first step environment variable by performing an action transition to the first step environment variable ; According to the reward function Computing to perform an action The reward obtained later, as shown in equation (12), (12) wherein, is a normal number, is the probability of attack occurrence known in the training; S342、according to the reward function and the first f-1 The Q value is updated according to the training value, as shown in equation (13) ,(13) wherein, is the f environmental variable-action pair for the th training round f-1 corresponding Q-value after the th training round f environmental variable for the th training round f-1 optimal action after the th training round is the learning step size, is the learning constant, wherein the initial value of the Q-table is ; S343、to the first step environment variable The process of S341-S342 is repeated to obtain the corresponding Q value until the system stops running, the first f round of training ends; S344, repeat the steps of S341-S343 for the next round of training until the training is finished, return the complete switching Q-table containing all environment variable-action pairs ; S35, in the attack detection stage, selecting a corresponding optimal action in the complete switching Q table : S351, determine the environment variable according to the modal estimation value, state estimation residual norm and modal estimation residual norm in the complete switching Q-table: Let , , when the current environment variable, the current modal estimate value is observed; S352, selecting an optimal action corresponding to the current environmental variable in the complete switching Q table ; S353, output the deception attack detection result, as shown in formula (14), (14)。 5. The security control method of networked handoff system based on handoff Q-learning spoofing attack detection according to claim 4, characterized in that, S4, obtain the controller mode according to the deception attack detection result, and substitute the state estimation value and the controller mode into the controller model to obtain a safety controller equation, and combine the safety controller equation, the state estimator of the deception attack detection mechanism, the modal estimator and the switching system equation to obtain a closed-loop system equation, including: Substitute the state estimation value and the modal estimation value into the controller model to obtain a safety controller equation, as shown in formula (16), , (15) , represents the modal information of the controller as shown in equation (16), (16) wherein, represents a modal transmission value calculated based on the result of the spoofing attack detection, i.e. the modal signal, represents a modal estimate value obtained by the modal estimator; Combine the safety controller equation, the state estimator equation, the modal estimator equation and the switching system equation to obtain a closed-loop system equation, as shown in formula (17), (17) Wherein, closed loop system extended state vector , extended system matrix , system parameter matrix , system parameter matrix , initial state vector , initial time , subsystem and controller modes .

6. The security control method of networked handoff system based on handoff Q-learning spoofing attack detection according to claim 5, characterized in that, S5, construct a switching rule based on deception attack detection, which is used to analyze the switching behavior of the closed-loop system under deception attack according to the deception attack detection result, obtain the judgment condition of system safety and stability, and perform safety control on the closed-loop system, including: S51, construct a switching rule based on deception attack detection, as shown in formula (18), (18) The initial value of the switching rule is , the state division of the i th subsystem , the boundary of the i th subsystem , and ; S52, analyze the deception attack detection result according to the switching rule and the event triggering mechanism, determine the running state of the subsystem mode and the controller mode under deception attack, obtain the judgment condition of system safety and stability, and use the judgment condition to perform safety control on the switching system.

Citation Information

Patent Citations

  • Security control method of event-driven network control system under multi-network attack

    CN110213115A

  • Network control method under spoofing attack based on interval type-2 T-S fuzzy

    CN113625558A