A train control defense method and system based on an infectious disease model

Through an infectious disease model-based method, an APT attack propagation model is established and machine learning algorithms are combined to dynamically detect and optimize train operation strategies, which solves the problem of neglecting the impact of the physical domain in the existing technology, realizes efficient train control and defense, and improves the real-time and robustness of the system.

CN119814438BActive Publication Date: 2025-07-18SOUTHWEST JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411966528.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-07-18
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

When facing APT attacks and pseudo-data injection attacks, existing train control systems ignore the direct impact of information domain attacks on the physical domain operation behavior, resulting in the inability to respond quickly, which may lead to train accidents or delays.

Method used

A method based on infectious disease model is adopted to establish an APT attack propagation model, combine random forest algorithms and long-term memory networks for dynamic detection, optimize train operation paths and speed control strategies, and reduce attack impact by adjusting mobile authorized information flow paths.

Benefits of technology

Comprehensive modeling and dynamic protection of the information physics domain have been achieved, the attack detection accuracy has been improved to 96%, the communication path recovery capability has been improved to 80%, the train scheduling delay has been reduced by 70%, and the real-time and robustness have been significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119814438B_ABST
    Figure CN119814438B_ABST
Patent Text Reader

Abstract

The present invention discloses a train control defense method and system based on an infectious disease model. The method includes: S1. Establish an APT attack propagation model to describe the node state transition and propagation process of the attack in the train control system; S2. Based on the APT attack propagation model, obtain the performance of the information domain network of the train control system and the train operation performance in the physical domain; S3. Collect the physical state data of train operation and the communication state data of the information domain to construct a multi-dimensional feature vector; S4. Use the random forest algorithm and the long short-term memory network to dynamically detect the train operation data and identify the false data injection attack; S5. After the APT attack occurs, optimize the train operation path and speed control strategy, and reduce the impact of the attack on the system by adjusting the mobile authorization information flow path. Through the dynamic interaction between the information domain and the physical domain, the present invention makes up for the neglect of the impact on the physical domain in the traditional defense method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of rail transit control systems, and particularly to a train control defense method and system based on an epidemic model. Background Art

[0002] With the gradual development of train control systems towards automation and intelligence, communication-based train control (CBTC) and cyber-physical systems (CPS) are widely used globally. These systems complete train dispatching, traction, braking, and safety guarantee through real-time communication among sensors, networks, and control units. However, such efficient intelligent systems are also threatened by APT attacks and false data injection attacks. APT attacks usually penetrate multiple nodes in the system step by step and damage the normal operation of critical nodes in a lateral expansion manner. False data injection attacks inject forged data packets into the train control system, such as tampering with speed, displacement, and acceleration information, causing the train to decelerate, overspeed, or stop abnormally, thereby affecting the overall traffic dispatching and operation safety.

[0003] On the one hand, existing technologies mainly focus on attack detection in the information domain. Although they can effectively detect and protect against single attack types, when facing complex attacks, they ignore the direct impact of information domain attacks on the operation behavior in the physical domain, and the defense mechanism cannot restore the normal operation of the system. On the other hand, most current defense systems have a certain delay in attack detection and response, which is particularly fatal when facing high-speed trains. When the train encounters a cyber attack or false data injection, the existing defense systems cannot make an effective response in a short time, which may lead to train accidents or delays. Therefore, it is necessary to propose a train control defense method that can comprehensively model and dynamically protect the cyber-physical domain. Summary of the Invention

[0004] To solve the problem of ignoring the impact on the physical domain in existing train control defense methods, the present invention proposes a train control defense method and system based on an epidemic model to solve the above problems.

[0005] The present application discloses a train control defense method based on an epidemic model, including the following steps:

[0006] S1. Establish an APT attack propagation model to describe the node state transition and propagation process of the attack in the train control system;

[0007] S2. Based on the APT attack propagation model, obtain the performance of the information domain network of the train control system and the train operation performance in the physical domain;

[0008] S3. Collect the physical state data of train operation and the communication state data in the information domain, and construct a multi-dimensional feature vector;

[0009] S4. Use the random forest algorithm and long short-term memory network to dynamically detect the train operation data and identify false data injection attacks;

[0010] S5. After the APT attack occurs, optimize the train operation path and speed control strategy, and reduce the impact of the attack on the system by adjusting the mobile authorization information flow path.

[0011] Preferably, the S1 includes the following steps:

[0012] Temporal analysis of multiple attack stages, and the time series is modeled as:

[0013]

[0014] Among them, P(i, j) represents the state transition probability, and i, j represent the device nodes in the train control system network;

[0015] Define the relationship between the propagation rate and the attack intensity as:

[0016] β = β0·(1 + α·I a )

[0017] Among them, β is the propagation rate, β0 represents the default propagation rate of the system without the influence of the attack intensity, α is the attack intensity coefficient, and I a is the intensity of the resources invested by the attacker;

[0018] The probability of successful infection of the virus software is:

[0019] λ = f(β) = 1 - e -β ;

[0020] Establish the APT attack propagation model as follows:

[0021]

[0022] Among them, N(t) is the number of normal nodes at time t, I(t) is the number of infected nodes at time t, P(t) is the number of propagation nodes at time t, C(t) is the number of completed nodes at time t, S(t) is the number of secure nodes at time t, θ is the probability that the attacker obtains the device control permission after intrusion, is the probability that the attacker launches an attack, δ is the false alarm rate of the intrusion product, and M is the total number of devices in the system.

[0023] Preferably, the APT attack propagation model is constrained as follows:

[0024] The direct information flow between nodes i and j in the train control system network is:

[0025]

[0026] Among them, I i,j (t) is the traffic volume on the edge from node i to node j under normal operating conditions, r i,j (t) is the abnormal information flow generated by the attacker's behavior, δ i (t) is the false alarm rate of the intrusion product at node i at time t;

[0027] The link capacity constraint is:

[0028] 0 ≤ f i,j (t) ≤ c i,j (t);

[0029] Among them, c i,j (t) is the link capacity;

[0030] The node traffic conservation constraint is as follows:

[0031]

[0032] Among them, E is the set of edges connecting any two nodes i and j in the network, S r is the source node, T r is the transfer node, S i is the sink node.

[0033] Preferably, the performance of the train control system information domain network is as follows:

[0034]

[0035] ε is the set of all feasible flow paths from S r to S i , N is the number of feasible flows in the set, e l is the l-th feasible path;

[0036] Normalize C P (t) to get:

[0037]

[0038] Among them, represents the minimum value of c p (t), is 0, indicating that there is no feasible path in the network, is the maximum value of C p (t), indicating that all information transmission paths in the network are available, Cy(t) ∈ [0, 1], and the closer it is to 0, the greater the impact of the attack on the train control system information domain network;

[0039] Taking the speed deviation of train operation and the train delay time as indicators to measure the train operation performance, the loss of train operation performance under APT attack is as follows:

[0040] Δp(t) = Δv + Δt = [v(d) - v norm (d)] + [t(d) - t schedule (d)];

[0041] Among them, Δv is the speed deviation of train operation, Δt is the train delay time, v(d) is the actual running speed of the train at position d, t(d) is the time when the train arrives at position d, v norm (d) is the running speed of the train at position d obtained according to the optimal speed - distance curve and the train timetable, t schedule (d) is the time when the train arrives at position d obtained according to the optimal speed - distance curve and the train timetable;

[0042] The train operation performance in the physical domain is obtained through normalization as follows:

[0043]

[0044] Among them, the smaller Py(t) is, the smaller the impact of APT attack on train operation performance. Among them, the smaller Py(t) is, the smaller the impact of APT attack on train operation performance. Δv(d) is the actual speed deviation of the train at position d, Δv min (d) is the minimum value of the actual speed deviation of the train at position d, Δv max (d) is the maximum value of the actual speed deviation of the train at position d, Δt(d) is the train delay time when the train arrives at position d, Δt min (d) is the minimum value of the train delay time when the train arrives at position d, Δt max (d) is the maximum value of the train delay time when the train arrives at position d.

[0045] Preferably, the S3 includes the following steps:

[0046] Using the analytic hierarchy process to solve the network performance in the information domain and the train operation performance in the physical domain, the information security analysis model of the train control system is obtained as:

[0047] Sy(t) = w1Cy(t) + w2Py(t);

[0048] Among them, w1 is the weight of the network performance Cy(t) in the information domain, and w2 is the weight of the train operation performance Py(t) in the physical domain.

[0049] Preferably, the S4 includes the following steps:

[0050] Multi-classifier fusion, realizing the output fusion of the random forest algorithm and the long short-term memory network through weighted average:

[0051] M final = η1M RF + η2M LSTM ;

[0052] Among them, M final represents the outputs of the random forest algorithm and the long short-term memory network. M RF is the classification probability of the random forest, and M LSTM is the classification probability of the long short-term memory network. η1 and η2 are adjustable weights used to balance the contributions of the random forest algorithm and the long short-term memory network.

[0053] Using a statistical model to verify the fusion result and reduce the false alarm rate:

[0054]

[0055] Among them, S represents the false alarm rate, a t is the current acceleration value, is the average value of the acceleration, and σ a is the standard deviation of the acceleration. If S exceeds the threshold, it is marked as abnormal;

[0056] Multi-modal feature fusion, fusing the input data of the random forest algorithm and the long short-term memory network through multi-modal features:

[0057] X = [X time , X network , X behavior ;

[0058] Among them, X time is the time series feature including speed, acceleration, and displacement, X network is the network feature including communication delay and packet loss rate, and X behavior is the device behavior feature including the CPU occupancy rate and memory usage rate of the device;

[0059] Using a deep learning architecture based on Transformer to model time series data:

[0060] H t = Transformer(X t );

[0061] Among them, X t is the input feature of the current time window, and H t is the high-dimensional feature extracted by the Transformer model;

[0062] Adaptive learning and update:

[0063] Update the weights of the classifier through online learning to adapt to the changing attack patterns. The online learning includes the following steps:

[0064] Store the detected abnormal samples in a temporary cache;

[0065] Retrain the model at fixed intervals and incorporate the latest attack samples.

[0066] Preferably, the S5 includes the following steps:

[0067] S51. Dynamic weight allocation for path optimization:

[0068] In path optimization, introduce a dynamic adjustment mechanism based on weights:

[0069]

[0070] where r 容量 is the remaining capacity of the path, is the probability that the path node is attacked, ρ1 and ρ2 are adjustable weights, and W 路径 is the weight of the path.

[0071] The weights are dynamically adjusted as follows:

[0072]

[0073] where, represents the time change rate of the path capacity, represents the change rate of the probability that the path node is attacked, α is a coefficient for weighing the capacity change, and β is a coefficient for weighing the change in the attack probability.

[0074] The path optimization objective is to minimize the path weight:

[0075]

[0076] where J 路径 is the path optimization objective, K is the number of all optional paths, and W 路径,k is the weight of the k-th path.

[0077] Model the remaining capacity of the path as a function of the link capacity,

[0078]

[0079] where C l is the total capacity of link l, and F l is the current traffic of link l;

[0080] Model the probability of a path node being attacked as a time - series module based on historical attack behaviors, and use a long - short - term memory network for prediction:

[0081]

[0082] Represents the probability that the path is attacked;

[0083] By dynamically adjusting ρ1 and ρ2, real - time path optimization is achieved. When the system detects a rapid decrease in the remaining capacity of the path, increase the weight of ρ1.

[0084] S52. Cooperative train scheduling strategy:

[0085] In the scenario of multi - train operation, coordinate the speeds and paths among trains to reduce the global scheduling disorder caused by attacks.

[0086] Under APT attacks, train scheduling optimization needs to consider speed, time, and path simultaneously. The optimization objective is:

[0087]

[0088] Among them, J 调度 is the scheduling optimization objective, Δv i is the deviation between the actual speed and the optimal speed of train i, Δt I is the deviation between the actual arrival time and the planned time of train i, φ i is the probability that the path where train i is located is attacked, and λ1, λ2, λ3 are weight parameters.

[0089] The actual running time of the train is determined by speed and acceleration:

[0090]

[0091] Among them, t i is the actual running time of train i, L i is the total running distance of train i, v i (s) represents the speed of train i at position s. By adjusting the train speed v i Time optimization is achieved, and combined with the train delay time constraint:

[0092]

[0093] Among them, T 容忍 is the maximum tolerable delay time, is the planned arrival time of train i.

[0094] S53. Attack isolation mechanism:

[0095] After an APT attack occurs, while isolating the infected nodes, communication recovery is achieved through redundant paths and standby devices, and the standby signal devices beside the track are used to take over the communication tasks.

[0096] Considering the communication capabilities of standby devices, the communication recovery ability is defined as:

[0097]

[0098] Among them, C m is the communication capacity of the device, φ m is the probability that the device is attacked. By maximizing the communication recovery ability R 恢复 , the allocation of optimal redundant devices is determined.

[0099] The objective function of isolation optimization is defined as:

[0100]

[0101] Among them, J 隔离 is the isolation optimization objective, is the combination of redundant paths, is the attack probability of path k, R 恢复,m is the recovery ability of the standby device, and γ is the coefficient used to balance the weights between path selection and device recovery.

[0102] To ensure the effectiveness of attack isolation, path selection needs to meet the following constraints:

[0103] R 容量 (P) ≥ R 阈值 ;

[0104] R 容量 (P) represents the remaining capacity of path P, and R 阈值 is the minimum path capacity.

[0105] S54. Real-time Defense and Prediction Model:

[0106] Use LSTM to dynamically predict the attack probability:

[0107]

[0108] Combining the characteristics of the physical domain and the information domain, use the Transformer model to analyze multi-modal data:

[0109] H = Transformer(X);

[0110] Among them, X is the multi-modal feature vector, and H is the high-dimensional feature.

[0111] To cope with the changes in the APT attack pattern, the defense model needs to store the monitored abnormal data in the cache and retrain the defense model every ΔT time:

[0112]

[0113] Among them, L is the loss function, and f(x i ; θ) is the predicted value of the defense model.

[0114] S55. Comprehensive optimization framework:

[0115] Combining path optimization, collaborative train scheduling, and attack isolation mechanisms to form a complete optimization framework. The overall optimization goal is:

[0116] minJ 总 = τ1·J 路径 + τ2·J 调度 + τ3·J 隔离 ;

[0117] Among them, τ1, τ2, and τ3 are weight parameters.

[0118] This application also discloses a train control defense system based on an infectious disease model, which is used to implement the train control defense method based on the infectious disease model, including:

[0119] Data acquisition module: used to collect train physical domain and information domain data in real time. The physical domain acquisition content includes dynamic operation information such as the real-time speed, acceleration, and displacement of the train; the information domain acquisition content includes the state of control signals, the communication state of data packets, and the communication health of the sensor network;

[0120] Detection module: based on the random forest algorithm and long short-term memory network, it detects false data injection attacks in real time;

[0121] Analysis module: used to simulate the APT attack propagation process and evaluate the potential impact of the attack on the train control system;

[0122] Defense module: used to dynamically adjust the train path and control strategy to improve the robustness of the train control system.

[0123] Advantages of the present invention:

[0124] (1) Full-coverage defense framework: provides an end-to-end system defense method from APT propagation modeling to false data injection detection.

[0125] (2) Cyber-physical coupling analysis: by dynamically interacting between the information domain and the physical domain, it makes up for the neglect of the impact on the physical domain in traditional defense methods.

[0126] (3) Real-time performance and robustness: Based on the dynamic detection and rapid response mechanism of machine learning algorithms, the defense is triggered within 0.2 seconds after an attack occurs.

[0127] (4) Efficient optimization: Adopt path optimization algorithms to reduce the impact of attacks on communication paths and ensure the safe operation of trains to the greatest extent.

[0128] (5) The attack detection accuracy reaches 96%, the communication path recovery ability is increased to 80%, and the train scheduling delay is reduced by 70%. Description of the Drawings

[0129] Figure 1 It is a schematic flow diagram of the train control defense method based on the infectious disease model according to the embodiment of the present invention;

[0130] Figure 2 It is a physical information coupling model diagram according to the embodiment of the present invention;

[0131] Figure 3 It is a dynamic response diagram of the defense strategy according to the embodiment of the present invention;

[0132] Figure 4 It is a train system speed change diagram according to the embodiment of the present invention;

[0133] Figure 5 It is a schematic structural diagram of the train control defense system based on the infectious disease model according to the embodiment of the present invention;

[0134] Figure 6 It is a schematic diagram of the defense effect according to the embodiment of the present invention. Detailed Embodiments

[0135] To make the objectives, technical solutions and advantages of the present application clearer, the following examples are given with reference to the accompanying drawings to further elaborate on the present application in detail.

[0136] The embodiment of the present invention discloses a train control defense method based on an infectious disease model. As Figure 1 shown, by establishing an APT attack propagation model, the node state transition and propagation process of the attack in the train control system are described. At the same time, the physical state data of train operation and the communication state data of the information domain are collected to construct a multi-dimensional feature vector. And machine learning algorithms are used to dynamically detect the train operation data to identify false data injection attacks. Finally, after an APT attack occurs, the train operation path and speed control strategy are optimized. By adjusting the path of the moving authorization information flow (MA), the impact of the attack on the system is reduced, and the integrated modeling of the cyber-physical domain is realized to dynamically protect the train control system. Specifically, it includes the following steps:

[0137] S1. Establish an APT attack propagation model to describe the node state transition and propagation process of the attack in the train control system.

[0138] Timing analysis of multiple attack phases, with the time series modeled as:

[0139]

[0140] Among them, P(i, j) represents the state transition probability, where i and j represent device nodes in the train control system network, and the Markov chain is used to analyze the dynamic behavior of state transitions.

[0141] Analyze the impact of the attack intensity on the propagation rate, and define the relationship between the propagation rate and the attack intensity as:

[0142] β = β0·(1 + α·I a )

[0143] Among them, β is the propagation rate, β0 represents the default propagation rate of the system without the influence of the attack intensity, and its value in this embodiment is from 0.01 to 0.1, α is the attack intensity coefficient, and I a is the intensity of the resources invested by the attacker.

[0144] The train control system can simulate the changes in the propagation rate under low, medium, and high-intensity attacks. The simulation is based on the propagation rate β0. When β0 takes the value of 0.01, it is low intensity; when β0 takes the value of 0.05, it is medium intensity; when β0 takes the value of 0.1, it is high intensity. As the attacker's resource investment increases, the propagation rate increases significantly, enhancing the adaptability to the real scenario. Then the probability of successful infection by the virus software is:

[0145] λ = f(β) = 1 - e -β ;

[0146] Establish the APT attack propagation model as follows:

[0147]

[0148] Among them, N(t) is the number of normal nodes at time t, I(t) is the number of infected nodes at time t, P(t) is the number of propagating nodes at time t, C(t) is the number of completed nodes at time t, S(t) is the number of secure nodes at time t, θ is the probability that the attacker obtains the device control authority after intrusion, is the probability that the attacker launches an attack, δ is the false alarm rate of the intrusion product, and M is the total number of devices in the system.

[0149] The direct information flow between node i and node j in the train control system network is:

[0150]

[0151] Among them, I i,jThe flow on the edge from node i to node j under normal operating conditions is (t), and r i,j The abnormal information flow generated by the attacker's behavior is (t), and δ i The false alarm rate of the intrusion product at node i at time t is;

[0152] The link capacity constraint is:

[0153] 0 ≤ f i,j (t) ≤ c i,j (t);

[0154] Among them, c i,j (t) is the link capacity.

[0155] The node flow conservation constraint is as follows:

[0156]

[0157] Among them, E is the set of edges connecting any two nodes i and j in the network, S r is the source node, T r is the transfer node, and S i is the sink node.

[0158] 2. Based on the APT attack propagation model, the performance of the information domain network of the train control system and the train operation performance in the physical domain are obtained.

[0159] The performance of the information domain network of the train control system is the product of the maximum feasible flow existing in the network at time t and its quantity, and the expression is as follows:

[0160]

[0161] ε is the set of all feasible flow paths from S r to S i , N is the number of feasible flows in the set, and e l is the l-th feasible path.

[0162] Normalize C P (t) to obtain:

[0163]

[0164] Among them, represents the minimum value of C p (t), is 0, indicating that there is no feasible path in the network, is the maximum value of C p (t), indicating that all information transmission paths in the network are available, Cy(t) ∈ [0, 1], and the closer it is to 0, the greater the impact of the attack on the information domain network of the train control system.

[0165] Taking the speed deviation of train operation and the train delay time as the indicators to measure the train operation performance, the loss of train operation performance under APT attack is as follows:

[0166] Δp(t) = Δv + Δt = [v(d) - v norm (d)] + [t(d) - t schedule (d)];

[0167] where, Δv is the speed deviation of train operation, Δt is the train delay time, v(d) is the actual running speed of the train at position d, t(d) is the time when the train arrives at position d, and v norm (d) is the running speed of the train at position d obtained according to the optimal speed - distance curve and the train timetable, and t schedule (d) is the time when the train arrives at position d obtained according to the optimal speed - distance curve and the train timetable;

[0168] The train operation performance in the physical domain after normalization is as follows:

[0169]

[0170] where, the smaller Py(t) is, the smaller the impact of APT attack on train operation performance. Δv(d) is the actual running speed deviation of the train at position d, Δv min (d) is the minimum value of the actual running speed deviation of the train at position d, Δv max (d) is the maximum value of the actual running speed deviation of the train at position d, Δt(d) is the train delay time when the train arrives at position d, and Δt min (d) is the minimum value of the train delay time when the train arrives at position d, and Δd max (d) is the maximum value of the train delay time when the train arrives at position d.

[0171] S3. Collect the physical state data of train operation and the communication state data in the information domain, and construct a multi - dimensional feature vector. The interaction between the physical domain and the information domain is as Figure 2 shown.

[0172] Using the analytic hierarchy process to solve the information domain network performance and the physical domain train operation performance, the information security analysis model of the train control system is obtained as:

[0173] Sy(t) = w1Cy(t) + w2Py(t);

[0174] where, w1 is the weight of the information domain network performance Cy(t), and w2 is the weight of the physical domain train operation performance Py(t).

[0175] S4. Use the random forest algorithm and the long short-term memory network (LSTM) to dynamically detect train operation data and identify false data injection attacks. The random forest algorithm can quickly detect specific patterns of false data injection, and the long short-term memory network can process time series data and identify the time dependence in attack behaviors.

[0176] The random forest is a machine learning algorithm based on the integration of multiple decision trees, which can quickly detect specific false data injection patterns.

[0177] Define the prediction loss of each decision tree as:

[0178]

[0179] where Loss represents the loss function, which is the cross-entropy loss, y i is the true value, is the predicted value of the g-th tree, and N is the number of data points.

[0180] The final predicted value is:

[0181]

[0182] where G is the number of trees in the forest.

[0183] The long short-term memory network is an improved model of the recurrent neural network that can capture long-term and short-term dependencies in time series data. The key of the long short-term memory network lies in its unique gate mechanism, including the forget gate, the input gate, and the output gate, which jointly act on the cell state.

[0184] Control the information to be discarded through the forget gate:

[0185] f t = σ(W f · [h t-1 , x t + b f );

[0186] where f t is the output of the forget gate, σ is the activation function, and in this embodiment, the activation function is adopted: W f is the weight matrix of the forget gate, b f is the bias of the forget gate, h t-1 is the hidden state at the previous moment, and x t is the current input.

[0187] Determine the information to be added through the input gate:

[0188]

[0189] where it is the output of the input gate, is the candidate cell state, tanh is the hyperbolic tangent activation function, W i is the weight matrix of the input gate, W C is the weight matrix of the candidate cell state, b C is the bias term of the candidate cell state.

[0190] Update the cell state:

[0191]

[0192] where C t is the cell state at the current time.

[0193] Output gate and hidden state update:

[0194]

[0195] where o t is the output of the output gate, h t is the hidden state at the current time, W o is the weight matrix of the output gate, b o is the bias of the output gate.

[0196] Using a long short-term memory network to capture complex dependencies in time series, it can capture the dynamic changes of acceleration, velocity, displacement, and the change patterns of communication delays.

[0197] Multi-classifier fusion, by combining the random forest algorithm and the long short-term memory network, improves the accuracy of detecting false data injection and attack behaviors. The outputs of the random forest algorithm and the long short-term memory network are fused through weighted averaging:

[0198] M final = η1M RF + η2M LSTM ;

[0199] where M final represents the outputs of the random forest algorithm and the long short-term memory network, M RF is the classification probability of the random forest, M LSTM is the classification probability of the long short-term memory network, and η1 and η2 are adjustable weights used to balance the contributions of the random forest algorithm and the long short-term memory network.

[0200] Adopt a statistical model to verify the fusion result and reduce the false alarm rate:

[0201]

[0202] where S represents the false alarm rate, a t is the current acceleration value, is the average value of acceleration, and σ a is the standard deviation of acceleration. If S exceeds the threshold, it is marked as abnormal;

[0203] Multi-modal feature fusion, input data for the random forest algorithm and long short-term memory network through multi-modal feature fusion:

[0204] X = [X time , X network , X behavior ;

[0205] where X time is the time series feature including speed, acceleration and displacement, and X network is the network feature including communication delay and packet loss rate, and X behavior is the device behavior feature including the CPU occupancy rate and memory usage rate of the device.

[0206] Introduce the Transformer model, and use a deep learning architecture based on Transformer (such as BERT or Time Series Transformer) to model time series data:

[0207] H t = Transformer(X t );

[0208] where X t is the input feature of the current time window, and H t is the high-dimensional feature extracted by the Transformer model. Based on this, capture the data dependence across time steps for monitoring the multi-stage behavior patterns in APT attacks.

[0209] Adaptive learning and update, update the weights of the random forest algorithm and long short-term memory network through online learning to adapt to the changing attack patterns. Online learning includes the following steps:

[0210] Store the detected abnormal samples in a temporary cache;

[0211] Retrain the model every fixed time (24 hours in this embodiment), that is, repeat the training based on the propagation rates under low, medium, and high-intensity attacks, and incorporate the latest attack samples.

[0212] S5. After an APT attack occurs, optimize the train operation path and speed control strategy, and reduce the impact of the attack on the system by adjusting the mobile authorization information flow path.

[0213] S51. Dynamic weight allocation for path optimization:

[0214] In path optimization, introduce a dynamic adjustment mechanism based on weights:

[0215]

[0216] Among them, R 容量 is the remaining capacity of the path, is the probability that the path node is attacked, ρ1 and ρ2 are adjustable weights, and W 路径 is the weight of the path.

[0217] The weights are dynamically adjusted as follows:

[0218]

[0219] Among them, represents the time change rate of the path capacity, represents the change rate of the probability that the path node is attacked, α is the coefficient for weighing the capacity change, and β is the coefficient for weighing the change in the attack probability.

[0220] The path optimization objective is to minimize the path weight:

[0221]

[0222] Among them, J 路径 is the path optimization objective, K is the number of all optional paths, and W 路径,k is the weight of the k-th path.

[0223] The remaining capacity of the path is modeled as a function of the link capacity,

[0224]

[0225] Among them, C l is the total capacity of link l, and F l is the current traffic of link l;

[0226] The probability that the path node is attacked is modeled as a time series module based on historical attack behaviors, and a long short-term memory network is used for prediction:

[0227]

[0228] represents the probability that the path is attacked.

[0229] By dynamically adjusting ρ1 and ρ2, real-time path optimization is achieved. When the system detects that the remaining capacity of the path decreases rapidly, the weight of ρ1 is increased.

[0230] S52. Cooperative train scheduling strategy:

[0231] In the scenario of multi-train operation, coordinate the speeds and paths among trains to reduce the global scheduling disorder caused by attacks.

[0232] Under APT attacks, train scheduling optimization needs to consider speed, time, and path simultaneously. The optimization objective is as follows:

[0233]

[0234] Among them, J 调度 is the scheduling optimization objective, Δv i is the deviation between the actual speed and the optimal speed of train i, Δt i is the deviation between the actual arrival time and the planned time of train i, φ i is the attack probability of the path where train i is located, and λ1, λ2, and λ3 are weight parameters.

[0235] The actual running time of the train is determined by speed and acceleration:

[0236]

[0237] Among them, t i is the actual running time of train i, L i is the total running distance of train i, v i (s) represents the speed of train i at position s. By adjusting the train speed v i time optimization is achieved, and combined with the train delay time constraint:

[0238]

[0239] Among them, T 容忍 is the maximum tolerable delay time, is the planned arrival time of train i.

[0240] S53. Attack isolation mechanism:

[0241] After an APT attack occurs, while isolating the infected nodes, communication recovery is achieved through redundant paths and standby devices, and the on-rail standby signal devices are used to take over the communication tasks.

[0242] Considering the communication capacity of the standby devices, the communication recovery ability is defined as:

[0243]

[0244] Among them, C m is the communication capacity of the device, φ m is the probability that the device is attacked. By maximizing the communication recovery ability R 恢复 , the allocation of the optimal redundant devices is determined.

[0245] The objective function of isolation optimization is defined as:

[0246]

[0247] Among them, J 隔离 is the isolation optimization objective, is the redundant path combination, is the attack probability of path k, R 恢复,m is the response ability of the standby device, and γ is the coefficient used to balance the weights between path selection and device recovery.

[0248] To ensure the effectiveness of attack isolation, the path selection needs to meet the following constraints:

[0249] R 容量 (P) ≥ R 阈值 ;

[0250] R 容量 (P) represents the remaining capacity of path P, and R 阈值 is the minimum path capacity.

[0251] S54. Real-time Defense and Prediction Model:

[0252] Dynamically predict the attack probability using LSTM:

[0253]

[0254] Combine the characteristics of the physical domain and the information domain, and use the Transformer model to analyze multi-modal data:

[0255] H = Transformer(X);

[0256] Among them, X is the multi-modal feature vector, and H is the high-dimensional feature.

[0257] To cope with the changes in the APT attack pattern, the defense model needs to store the monitored abnormal data in the cache and retrain the defense model every ΔT time:

[0258]

[0259] Among them, L is the loss function, and f(x i ; θ) is the predicted value of the defense model.

[0260] S55. Comprehensive Optimization Framework:

[0261] Combine path optimization, collaborative train scheduling, and attack isolation mechanisms to form a complete optimization framework. The overall optimization objective is:

[0262] min J 总 = τ1·J 路径 + τ2·J 调度 + τ3·J 隔离 ;

[0263] Among them, τ1, τ2, and τ3 are weight parameters.

[0264] The dynamic response of the defense strategy is as Figure 3 shown. The train detects an attack at 0.2 seconds, and the system has recovered at 1.2 seconds. The picture shows the dynamic response of the defense strategy, demonstrating that the system has fast real-time performance and robustness.

[0265] The change in the train's defense speed is as Figure 4 shown. As Figure 4 can be seen, the train speed under normal operating conditions varies between 55 km / h and 65 km / h. This indicates that during normal operation, the train travels according to a predetermined speed pattern, and the speed change is regular and predictable. After the train is attacked, the train speed shows significant fluctuations and anomalies. At about 3 seconds, the train speed drops sharply and even drops to 0 km / h, indicating that the attack has a serious impact on the train's operation and may cause the train to stop running or other failures. After taking defense measures, the train speed gradually returns to the normal operating state. Although there are still some fluctuations in the speed at the initial stage of defense, as time goes by, the speed gradually stabilizes and returns to the normal periodic fluctuation pattern. The following conclusions can be drawn: When the train is operating normally, the speed change is regular and predictable. When the train is attacked, the speed will show serious abnormal fluctuations, which may lead to operation failures. After taking effective defense measures, the train speed can gradually return to the normal operating state, ensuring the safety and normal operation of the train.

[0266] Another embodiment of this application also discloses a train control and defense system based on an infectious disease model for implementing the above-mentioned train control and defense method based on an infectious disease model, as Figure 5 shown, including:

[0267] Data acquisition module: used to collect train physical domain and information domain data in real time. The physical domain acquisition content includes dynamic operation information such as the real-time speed, acceleration, and displacement of the train; the information domain acquisition content includes the status of control signals, the communication status of data packets, and the communication health of the sensor network;

[0268] Detection module: based on the random forest algorithm and long short-term memory network to detect false data injection attacks in real time;

[0269] Analysis module: used to simulate the APT attack propagation process and evaluate the potential impact of the attack on the train control system;

[0270] Defense module: used to dynamically adjust the train path and control strategy to improve the robustness of the train control system.

[0271] In a specific embodiment, the SUMO simulation tool is used to conduct a joint test of APT attacks and fake data injection to evaluate the detection accuracy, response time, and recovery ability of the defense system. The results are as Figure 6 shown. The communication path recovery ability is increased to 80%, the response time triggers the defense 0.2 seconds after the attack occurs, and the monitoring accuracy rate is as high as 96%.

[0272] This application takes cyber-physical coupling as the core and comprehensively solves the security problems of modern rail transit systems under advanced threat attacks through innovative APT propagation modeling, machine learning detection, path optimization, and dynamic defense strategies. Through modular design and cross-system integration, it demonstrates extremely high adaptability and practicality, and can significantly improve the robustness and scheduling efficiency of rail transit.

[0273] The above has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A train control defense method based on an infectious disease model, characterized in that, It includes the following steps: S1. Establish an APT attack propagation model to describe the node state transition and propagation process of the attack in the train control system; S2. Based on the APT attack propagation model, obtain the performance of the information domain network and the train operation performance of the physical domain of the train control system; S3. Collect the physical state data of train operation and the communication state data of the information domain, and construct a multi-dimensional feature vector; S4. Use the random forest algorithm and the long short-term memory network to dynamically detect the train operation data and identify the false data injection attack; S5. After the APT attack occurs, optimize the train operation path and speed control strategy, and reduce the impact of the attack on the system by adjusting the mobile authorization information flow path; S51. The path optimization goal is to minimize the path weight: Among them, J 路径 is the path optimization objective, K is the number of all optional paths, and W 路径,k is the weight of the k-th path; Model the remaining path capacity as a function of the link capacity, Among them, C l is the total capacity of link l, and F l is the current traffic of link l; Model the probability of a path node being attacked as a time series module based on historical attack behaviors, and use the long short-term memory network for prediction: Indicates the probability that the path is attacked; Achieve real-time path optimization by dynamically adjusting ρ1 and ρ2; S52. Under the APT attack, considering speed, time, and path simultaneously, the optimization goal is: Among them, J 调度 is the scheduling optimization objective, Δv i is the deviation between the actual speed and the optimal speed of train i, Δt i is the deviation between the actual arrival time and the planned time of train i, φ i is the attack probability of the path where train i is located, and λ1, λ2, and λ3 are weight parameters; S53. Define the objective function of isolation optimization as: Among them, J 隔离 is the isolation optimization objective, is the redundant path combination, is the attack probability of path k, R 恢复,m is the response ability of the standby device, and γ is the coefficient used to balance the weights between path selection and device recovery; To ensure the effectiveness of attack isolation, the path selection needs to meet the following constraints: R 容量 (P) ≥ R 阈值 ; R 容量 (P) represents the remaining capacity of path P, R 阈值 is the minimum path capacity; S54. Real-time defense and prediction model: Use LSTM to dynamically predict the attack probability: Combining the characteristics of the physical domain and the information domain, use the Transformer model to analyze multi-modal data: H = Transformer(X); where X is the multi-dimensional feature vector and H is the high-dimensional feature; To cope with the change of the APT attack mode, the defense model needs to store the monitored abnormal data in the cache and retrain the defense model every ΔT time; Among them, L is the loss function, and f(x i ; θ) is the predicted value of the defense model; S55. Comprehensive optimization framework: Combining path optimization, collaborative train scheduling, and attack isolation mechanisms to form a complete optimization framework, and the overall optimization goal is: minJ 总 = τ1·J 路径 + τ2·J 调度 + τ3·J 隔离 ; where τ1, τ2, and τ3 are weight parameters.

2. The train control defense method based on the infectious disease model according to claim 1, wherein The S1 includes the following steps: Temporal analysis of multiple attack stages, and the time series is modeled as: where P(i,j) represents the state transition probability, and i and j represent the device nodes in the train control system network; Define the relationship between the propagation rate and the attack intensity as: β=β0·(1+α·I a ); Among them, β is the transmission rate, β0 represents the default transmission rate of the system without the influence of the attack intensity, α is the attack intensity coefficient, and I a is the intensity of resources invested by the attacker; The probability of successful infection of the virus software is: λ = f(β) = 1 - e -β ; Establish the APT attack propagation model as follows: Among them, N(t) is the number of normal nodes at time t, I(t) is the number of infected nodes at time t, P(t) is the number of spreading nodes at time t, C(t) is the number of completed nodes at time t, S(t) is the number of secure nodes at time t, θ is the probability that the attacker obtains the device control permission after intrusion, is the probability that the attacker launches an attack, δ is the false alarm rate of the intruded product, and M is the total number of devices in the system.

3. The train control defense method based on the infectious disease model according to claim 2, characterized in that, The constraints of the APT attack propagation model are as follows: The direct information flow between nodes i and j in the train control system network is: Among them, I i,j (t) is the traffic on the edge from node i to node j under normal operating conditions, r i,j (t) is the abnormal information flow generated by the attacker's behavior, δ i (t) is the false alarm rate of the intrusion product at node i at time t; The link capacity constraint is: 0 ≤ f i,j (t) ≤ c i,j (t); where c i,j (t) is the link capacity; The node flow conservation constraint is as follows: Among them, E is the set of edges connecting any two nodes i and j in the network, S r is the source node, T r is the transfer node, S i is the sink node.

4. The train control defense method based on the infectious disease model according to claim 3, wherein, The performance of the information domain network of the train control system is as follows: ε is the set of all feasible flow paths from S r to S i , N is the number of feasible flows in the set, and e l is the l-th feasible path; For C P (t) is normalized to obtain: Among them, represents the minimum value of C p (t), being 0 indicates that there is no feasible path in the network, being C p (t) maximum value indicates that all information transmission paths in the network are available. Cy(t) ∈ [0, 1]. The closer it is to 0, the greater the impact of the attack on the information domain network of the train control system; Take the speed deviation of train operation and the train delay time as indicators to measure the train operation performance. Under the APT attack, the loss of train operation performance is as follows: Δp(t) = Δv + Δt = [v(D) - v norm (d)] + [t(d) - t schedule (d)]; Among them, Δv is the speed deviation of the train during travel, Δt is the train delay time, v(d) is the actual travel speed of the train at position d, t(d) is the time when the train arrives at position d, and v norm (d) is the travel speed of the train at position d obtained according to the optimal speed - distance curve and the train schedule, and t schedule (d) is the time when the train arrives at position d obtained according to the optimal speed - distance curve and the train schedule; Obtain the train operation performance of the physical domain through normalization as follows: Among them, the smaller Py(t) is, the smaller the impact of the APT attack on the train operation performance. Among them, the smaller Py(t) is, the smaller the impact of the APT attack on the train operation performance. Δv(d) is the actual driving speed deviation of the train at position d, and Δv min (d) is the minimum value of the actual driving speed deviation of the train at position d, and Δv max (d) is the maximum value of the actual driving speed deviation of the train at position d, and Δt(d) is the late arrival time of the train at position d, and Δt min (d) is the minimum value of the late arrival time of the train at position d, and Δt max (d) is the maximum value of the late arrival time of the train at position d.

5. The train control defense method based on the infectious disease model according to claim 4, wherein The S3 includes the following steps: Use the analytic hierarchy process to solve the information domain network performance and the train operation performance of the physical domain, and obtain the information security analysis model of the train control system as: Sy(t) = w1Cy(t) + w2Py(t); Among them, w1 is the weight of the information domain network performance Cy(t), and w2 is the weight of the physical domain train operation performance Py(t).

6. The train control defense method based on the infectious disease model according to claim 5, wherein The S4 includes the following steps: Multi-classifier fusion, achieving the output fusion of the random forest algorithm and the long short-term memory network through weighted average: M final = η1M RF + η2M LSTM ; Among them, M final represents the outputs of the random forest algorithm and the long short-term memory network. M RF is the classification probability of the random forest, and M LSTM is the classification probability of the long short-term memory network. η1 and η2 are adjustable weights; Using a statistical model to verify the fusion result and reduce the false alarm rate: Among them, S represents the false alarm rate, and a t is the current acceleration value, is the average value of the acceleration, and σ a is the standard deviation of the acceleration. If S exceeds the threshold, it is marked as abnormal; Multi-modal feature fusion, fusing the input data of the random forest algorithm and the long short-term memory network through multi-modal features: X = [X time , X network , X behavior ; Among them, X time is a time series feature including speed, acceleration, and displacement, and X network is a network feature including communication delay and packet loss rate, and X behavior is a device behavior feature including the CPU occupancy rate and memory usage rate of the device; Using a deep learning architecture based on Transformer to model time series data: H t = Transformer(X t ); Among them, X t is the input feature of the current time window, and H t is the high-dimensional feature extracted by the Transformer model; Adaptive learning update, updating the weights of the classifier through online learning to adapt to the constantly changing attack patterns. The online learning includes the following steps: Storing the detected abnormal samples in a temporary cache; Retraining the model every fixed time and integrating the latest attack samples.

7. A train control and defense system based on an infectious disease model, characterized in that, For implementing the train control defense method based on the infectious disease model according to any one of claims 1-6, including: Data acquisition module: used to collect train physical domain and information domain data in real time. The physical domain acquisition content includes dynamic operation information such as the real-time speed, acceleration, and displacement of the train; the information domain acquisition content includes the state of control signals, the communication state of data packets, and the communication health of the sensor network; Detection module: based on the random forest algorithm and the long short-term memory network, detecting false data injection attacks in real time; Analysis module: used to simulate the APT attack propagation process and evaluate the potential impact of the attack on the train control system; Defense module: used to dynamically adjust the train path and control strategy to improve the robustness of the train control system.

Citation Information

Patent Citations

  • Network security situation awareness and prediction method and system based on LSTM and random forest

    CN115378653A