Injection type threat detection system and method of industrial control system
By building an attention-enhanced graph neural network model, the blind spots and data imbalance of injection threat detection in the OT domain in the industrial control system are solved, real-time detection and traceability of injection threats are realized, false alarm rate is reduced, complex industrial environment is adapted to, and security protection needs of low latency are met.
Patent Information
- Application Number
- CN202510707118.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art is difficult to effectively identify and prevent injection threats in industrial control systems, especially in the detection of injection threats in the OT domain, which leads to lagging malicious behavior discovery, increasing economic losses and security risks, and the detection delay problem of existing detection models for data imbalance and covert attacks has not been effectively solved.
An injection-type threat detection system and method of industrial control system is adopted. Through data acquisition module, dynamic feature modeling module, threat detection module and multi-dimensional sub-model training and entropy-based judgment module, an attention-enhancing graph neural network model is built to realize real-time detection and traceability of injected threats in the OT domain. Combined with residual connection, layer normalization and attention mechanism, the detection threshold is dynamically adjusted to reduce the false alarm rate.
It realizes efficient detection of threats injected within the OT domain, improves detection accuracy, reduces false alarm rate, shortens traceability time, adapts to complex industrial environments, meets the security protection needs of low-latency, is suitable for edge deployment, and is suitable for security protection needs of industrial control systems.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial Internet security, and particularly relates to an injection threat detection system and method for an industrial control system. Background Art
[0002] The rapid development of the industrial Internet has promoted the deep integration of information technology (IT) and industrial control (IC), making the industrial control network face a new security challenge of the detection blind area of OT domain compromise threats. 52% of vulnerability exploitations are directly related to the initial access of attackers, and these vulnerabilities often achieve system penetration through injection threats (such as unpatched cloud service vulnerabilities and industrial control system vulnerabilities). For example, attackers use protocol parsing vulnerabilities or cloud environment configuration defects in the industrial Internet to implant malicious code to achieve lateral movement.
[0003] Through research and analysis of a large number of relevant domestic and foreign literatures, the research on injection threat detection in industrial control Internet mainly focuses on IT domain protection, and rarely considers the direct threats brought by OT domain compromise. Although the existing technologies provide a certain boundary protection mechanism for industrial Internet security, due to the complexity and uniqueness of the industrial control system structure, it is usually difficult to effectively identify the injection threat characteristics from the IT domain to the OT domain. With the continuous evolution of advanced persistent threats (APT) and zero-day vulnerabilities, the risk of breaking through the traditional boundary security protection mechanism is becoming increasingly serious, and relying solely on boundary traffic detection can no longer meet the deep protection requirements of industrial control systems.
[0004] For the detection of injection threats, current research tends to adopt the Traffic Behavior Analysis method to identify potential injection attack types by actively identifying abnormal feature behaviors in traffic. However, the current security protection strategies mainly focus on intercepting the compromise threats in the IT field, and pay insufficient attention to the injection threats directly caused in the OT field after industrial control devices are connected to the network. These strategies have limitations in actively discovering the compromise threats inside the industrial control network, resulting in a delay in the discovery time of malicious behaviors, bringing incalculable economic losses and strategic risks to the national industrial Internet. Therefore, it has become a key problem to be solved urgently to efficiently and accurately detect and evaluate the suspected compromised intranet of industrial control to discover abnormal behaviors and locate the compromised nodes.
[0005] In view of this, it is necessary to adopt injection threat detection in industrial control systems to prevent various threats caused by injection malicious behaviors.
[0006] Currently, the detection research on injection threats in industrial control systems mainly adopts the Traffic Behavior Analysis (TBA) method, which identifies potential attack behaviors by mining abnormal features in network traffic. However, the existing security protection system has the following limitations:
[0007] 1) IT / OT security protection fragmentation: The current defense measures mainly focus on threat interception in the IT domain, while lacking effective control over the injection threats directly exposed after industrial control devices in the OT domain are connected to the network. This security defense gap leads to a significant lag in the discovery of malicious behaviors, bringing major economic losses and national security risks to the industrial Internet.
[0008] 2) Detection model relying on imbalanced data: Most existing injection threat detection methods rely on a large number of malicious traffic samples for model training. In the actual operation and maintenance environment of industrial systems, the proportion of normal traffic far exceeds that of abnormal traffic (for example, the abnormal proportion in log data is <0.1%). The extreme inter-class distribution imbalance of the dataset causes the trained model to be extremely sensitive to regular traffic, resulting in a soaring false alarm rate.
[0009] 3) Detection delay of covert attack behaviors: Modern network attacks usually adopt multi-level proxy hopping or traffic obfuscation technologies, making abnormal behaviors show highly non-linear characteristics in terms of time sequence, and it is difficult for the detection engine to respond in real time. Such delays may lead to the interruption of core industrial automation control processes, further expanding economic losses.
[0010] Therefore, quickly detecting anomalies in the internal traffic of the OT domain, accurately tracing the compromised nodes, and blocking their further lateral movement have become the key technical challenges that need to be solved urgently. Summary of the Invention
[0011] To overcome the above deficiencies of the prior art, the purpose of the present invention is to provide an injection threat detection system and method for industrial control systems, which mainly discovers anomalies through feature dimensionality reduction extraction and detailed detection of injection attacks that bypass the IT domain protection and reach the OT domain, and has the characteristics of high detection accuracy, strong controllability of false alarm rate, and excellent calculation efficiency.
[0012] To achieve the above purpose, the technical solution adopted by the present invention is: an injection threat detection system for industrial control systems, including a data acquisition module, a dynamic feature modeling module, a threat detection module, a multi-dimensional sub-model training and entropy-based determination module, and an abnormal tracking function module;
[0013] The data acquisition module converts the original industrial control network traffic into a graph structure representation and extracts node features;
[0014] The described dynamic feature modeling module cleans, transforms, and structures the collected heterogeneous data to generate a dynamic traceability graph suitable for input to the graph neural network, and extracts multi-dimensional features of nodes and edges from the graph data for subsequent model training and detection;
[0015] The described threat detection module is a three-layer network structure based on the improved SAGE network, which is used to capture the global and local associations of threat events;
[0016] The described multi-dimensional sub-model training and entropy-based determination module is used to deal with stealth attacks by recursively traversing and expanding the associated subgraphs of suspicious nodes.
[0017] An injection threat detection method for an industrial control system includes the following steps:
[0018] Step 1, data collection and dynamic heterogeneous graph feature modeling
[0019] Data is collected from network traffic, system logs, and device status respectively, and the collected data is stored in a time series form to ensure that it can reflect the dynamic changes of system behavior; the data collection module preprocesses the collected data, converts the streaming collected data into a traceability graph to obtain the graph data for model training and recognition; through the dynamic feature modeling module, abstract feature extraction is performed on the obtained graph data, and the abstract feature extraction is to mine multi-dimensional features from the preprocessed graph data to enhance the model's ability to identify injection threats;
[0020] Step 2, construct an injection threat detection model for a graph neural network with enhanced attention
[0021] The described RASM model constructs a heterogeneous graph for more in-depth and comprehensive learning of the enhanced abstract features of the logs, strengthens the model's ability to focus on high-value features in the traffic logs, and introduces key technologies such as attention mechanism, residual connection, and layer normalization on the basis of the SAGE model. The node features are generated by aggregating in-edge and out-edge statistical information, and local attributes and relationship information are fused to lay a foundation for subsequent detection;
[0022] Step 3, multi-dimensional sub-model training and entropy threshold evaluation
[0023] Abnormal behaviors in complex systems usually cover multiple different dimensions. A single detection model often cannot fully capture the information of these dimensions. Multidimensional sub-model training and entropy threshold judgment modules are used to solve the problem of insufficient in-depth learning of traceability graph dimensions. The multidimensional sub-model training architecture divides the original graph data into multiple sub-graphs according to different dimensions, builds a special sub-model for each sub-graph, and calculates the information entropy value of the prediction results of each sub-model at the same time, and constructs a threshold judgment mechanism based on information entropy. Specifically, the predicted probability distribution of the model is first converted into a Shannon entropy quantification indicator. When the entropy value is lower than the preset threshold, the prediction result is adopted, otherwise it is eliminated. Finally, the anomaly detection results are verified by multi-model fusion to improve the overall detection performance and generalization ability.
[0024] Step 4: Implement exception tracking
[0025] When the RASM model finds a suspicious node, due to its anomaly-based detection characteristics, in order to reduce the false alarm rate and track possible malicious nodes, the waiting queue algorithm of the anomaly tracking function module is adopted to reasonably extend the suspicious node detection cycle to effectively control the occurrence of false alarms.
[0026] The data preprocessing and dynamic feature modeling described above, the goal of data preprocessing is to convert the original data into a graph structure suitable for RASM model input, and the goal of dynamic feature modeling is to extract the abstract features of the graph structure, specifically including the following steps:
[0027] Step 1, data cleaning, removing duplicate records and invalid data (such as null value packets), using interpolation to fill missing values, and eliminating outliers (based on the 3σ criterion);
[0028] Step 2: Entity recognition and relationship extraction: Identify entities in streaming data as nodes of a graph; extract relationships between entities as edges of the graph;
[0029] Step 3: Dynamic traceability graph construction, defining node type set and edge type set Count the number of node types and the number of edge types
[0030] Step 4, type mapping, construct node type mapping function and edge type mapping function Map node and edge types to a range of integers, where N v Number of node types, N e Number of edge types;
[0031] Step 5: Feature initialization: assign a feature vector to each node v∈V Including in-edge and out-edge statistical features:
[0032] Feature components of in-edges:
[0033] a i = |{e ∈ In(v): i = M e (χ e (e))}|, i ∈ {0, …, N e - 1}
[0034] Feature components of out-edges:
[0035]
[0036] where In(v) represents all in-edges with v as the target node, Out(v) represents all out-edges with v as the source node, χ e (e) represents the type of edge e, and M e is the edge type mapping function;
[0037] Step6, construct the complete feature representation of node v, where a i represents the in-out edge feature components of node v.
[0038] Step7, for each node v, extract local structure features based on three-hop neighborhood sampling:
[0039]
[0040] where N3(v) represents the three-hop neighborhood of node v, AGGREGATE uses mean aggregation, and x u represents a certain part of the node features,
[0041] Step8, assign a weight ω e to each edge e ∈ E, based on the type and frequency of the edge: ω e = log(1 + freq(e)) where freq(e) represents the number of occurrences of edge e within the time window;
[0042] Step9, introduce time dimension features Record the timestamp distribution of node behaviors:
[0043]
[0044] where Δt is the time interval of events associated with node v, and mean and std represent the mean and standard deviation;
[0045] Step10, output the finally abstracted node features:
[0046]
[0047] The described attention-enhanced graph neural network injection-based threat detection model mines multi-dimensional information from the graph data after abstract feature preprocessing, introduces residual connections and attention mechanisms, and significantly improves feature aggregation and detection performance. It specifically includes the following steps:
[0048] Step-1: Adopt the GraphSAGE graph convolution framework to construct a three-layer graph convolution network (SAGE), and layer by layer realize the aggregation and update of node features, providing a hierarchical processing architecture for the feature extraction of heterogeneous graph data.
[0049] Step-2: Embed residual connections (Residual Connection) in the three-layer graph convolution network. Through cross-layer feature direct connection, retain the original input information of the previous layer network, avoid the problems of gradient disappearance and performance degradation caused by the increase in network depth, and ensure the effectiveness of feature transmission in the deep network.
[0050] Step-3: Introduce an attention mechanism during the neighbor node aggregation process. By calculating the importance weights between nodes, dynamically screen the key neighbor nodes related to threat detection, suppress the interference of noise features, and enhance the model's ability to capture abnormal patterns.
[0051] Step-4: Add layer normalization (Layer Normalization) after each layer of graph convolution operation to standardize the node features. By stabilizing the feature distribution, reduce the risk of gradient disappearance and improve the training stability.
[0052] Step-5: Use the Leaky ReLU activation function to replace the traditional ReLU to solve the problem of neuron inactivation in the negative interval in the sparse data scenario, and enhance the model's non-linear expression ability for low-frequency abnormal features (such as rare attack patterns).
[0053] Step-6: Apply the cosine annealing strategy to dynamically adjust the learning rate during the training phase. By periodically decaying the learning rate, balance the global exploration and local fine-tuning capabilities of the model in the solution space.
[0054] Step-7: Monitor the detection metrics (such as F1 score, AUC) on the validation set. When the metrics do not improve within N consecutive epochs, trigger the early stopping mechanism to terminate the training in advance. This mechanism avoids overfitting and reduces redundant computational consumption by dynamically evaluating the model's generalization performance.
[0055] The described multi-dimensional sub-model training and entropy threshold determination specifically includes the following steps:
[0056] Step 3-1: First, divide the original graph data G=(V, E) into multiple subgraphs G according to the row dimensiond =(V, E d ), where V is the set of nodes and E is the set of edges. The behavior dimensions include: system calls, process communication, and network interaction. E d represents the set of edges for a specific behavior dimension; through the edge type mapping function transform the behavior features into one-dimensional indices. R is the set of edge types for easy processing. For each subgraph G d , calculate the node feature vector featd(v). This node feature vector featd(v) contains statistical information about incoming and outgoing edges. The statistical information includes: the number or weight of the edges; to focus on unrecognized abnormal patterns, reconstruct the complex features, remove the known parts, and ensure that the subsequent steps of training are carried out for features specific to a particular dimension;
[0057] Step 3-2, for each subgraph G d , train a dedicated sub-model M d . The architecture of the sub-model adopts the enhanced graph neural network architecture RASM. The goal of the sub-model M d is to learn the normal behavior patterns in this dimension and detect abnormal behaviors that deviate from the normal system operation behavior patterns. By focusing on the signal capture of a single dimension, the sub-model M d can improve the accuracy of overall detection;
[0058] Step 3-3, multi-dimensional label calculation and entropy threshold determination
[0059] Each sub-model M d outputs the predicted probability distribution P d (v)=[p d,1 , p d,2 ,…, p d,K , where K is the number of classes. To measure the uncertainty of the prediction, calculate the entropy value:
[0060]
[0061] where, p d,k represents the node prediction probability. Normalize H(v) to get:
[0062]
[0063] In the formula, 0 represents complete certainty and 1 represents maximum uncertainty. In the pre-training stage, use the entropy value distribution H val of the validation set V d ={H d (v i )|v i ∈V val}, select the 80% quantile Q 0.8 (H d)As the initial entropy threshold R t,d , and then adjust the threshold through the dynamic update formula:
[0064] R t_new = λR t_old +(1 - λ)R t_estimated
[0065] where λ is the smoothing factor to ensure that the threshold adapts to data changes, R t_estimated represents the current entropy threshold estimated based on the latest data, and R t_old represents the threshold that was in effect at the previous moment;
[0066] Step 3 - 4, Node classification and dynamic mask evolution, classify according to the comparison between the normalized entropy and the threshold obtained in Step 3 - 3:
[0067] If H norm,d (v) ≤ R t,d , accept the prediction of M d , and select the class with the highest probability argmax i p d,i as the result; if H norm,d (v) > R t,d , mark the node v as unclassified, with a predicted value of -1, to be processed later; initialize the behavior status matrix S (0) = [1, 1, …, 1], indicating that all nodes participate in training. After each round of training, update the status according to the confusion matrix CM (t) : If node i belongs to false positive FP or true negative TN, then it will no longer participate in training; retain the true positive TP and false negative FN nodes to make the model gradually focus on key behaviors;
[0068] Step 3 - 5, Construct an optimized training sub - graph through a two - stage neighborhood sampling strategy. In the first stage, sample the direct neighbors of the nodes, and in the second stage, further sample highly relevant neighbors from these neighbors to form a sub - graph containing multi - layer neighborhood information, and provide the sub - graph structure to the sub - model M d for training to enhance the model's ability to understand and capture complex relationships;
[0069] Step 3 - 6, Calculate the performance evaluation score for each sub - model M d :
[0070]
[0071] Set the double - threshold condition: Precision P(M d ) ≥ γ p , Recall R(M d ) ≥ γ r , and select Q(M d)The largest sub-model As the optimal sub-model, the prediction results of multiple optimal sub-models are further fused to comprehensively judge the abnormal state of the node and improve the detection comprehensiveness.
[0072] For the waiting queue mentioned above, its algorithm includes the following steps:
[0073] Step4-1, Initialize the priority queue Q, which contains all nodes v with abnormal scores S(v) > θ, where θ is a preset threshold and S(v) is the node scoring function. At the same time, initialize the abnormal node set V abnormal to be empty, which is used to store the finally confirmed abnormal nodes;
[0074] Step4-2, Traverse each node v in the priority queue Q, and set its timestamp t(v) to the current time current_time to record the moment when the node enters the queue, providing a basis for subsequent time window judgment;
[0075] Step4-3, As long as the priority queue Q is not empty, continuously perform the following operations: Take out a node v from the priority queue Q, and check whether the k-hop neighborhood N k (v) has an intersection with the normal node set V normal ; If there exists at least one node u ∈ N k (v) ∩ V normal , it means that the neighborhood of v contains normal behaviors, then remove v from the queue Q and do not process it further;
[0076] Step4-4, If the k-hop neighborhood N k (v) of node v has no intersection with V normal , expand its neighborhood by one level and update it to N k+1 (v) to check for more extensive associated behaviors; Then, judge whether the difference between the current time t current and the timestamp t(v) of node v exceeds the preset time window T. If t current - t(v) > T, it means that v has not been classified as normal within the time window, then add it to the abnormal node set V abnormal , otherwise, re-add v to the queue Q and wait for the next round of inspection;
[0077] Step4-5, When the priority queue Q is empty, the algorithm terminates and returns the abnormal node set V abnormal , which contains all nodes that have not been classified as normal within the time window T, and these nodes are confirmed as malicious nodes.
[0078] The described RASM model adopts a multi-level neighbor sampling strategy (such as sampling widths [30, 20, 10]), restricting the number of neighbors in each layer, which not only preserves key structural information but also improves computational efficiency and adapts to the large-scale data of industrial systems;
[0079] The described RASM model uses a three-layer graph convolutional network (SAGE), retaining the information of the previous layer through residual connections to avoid performance degradation caused by increased depth;
[0080] The described RASM model introduces an adaptive threshold mechanism based on Shannon entropy to dynamically adjust the detection criteria: by calculating the entropy of the predicted probability distribution of each node, low-entropy nodes (high confidence) are used to estimate the optimal threshold for distinguishing normal and abnormal behaviors;
[0081] Compared with the prior art, the beneficial effects of the present invention are:
[0082] Based on the graph network architecture, the present invention designs an injection threat detection method for industrial control systems. By innovatively integrating the residual attention mechanism and the GraphSAGE framework, the RASM model is constructed to achieve efficient detection of OT domain compromise threats. This solution introduces an organic combination of layer normalization, residual connections, and attention mechanism at the system architecture level, effectively solving the problems of gradient disappearance and overfitting in the training of large-scale heterogeneous graph data. At the algorithm design level, through a multi-level neighborhood sampling strategy and gradient accumulation technology, the computational efficiency and feature representation ability are successfully balanced. In particular, the adaptive decision-making mechanism based on entropy threshold designed by the present invention can dynamically adjust the detection parameters, optimize the detection strategy according to the enterprise data distribution characteristics, and significantly reduce the alarm fatigue effect caused by the high false alarm rate of the intrusion prevention system. Experimental results show that compared with traditional methods, this solution shows significant advantages in detection accuracy, false alarm rate control, and computational efficiency.
[0083] Aiming at the injection threat detection of industrial control systems, the present invention has achieved breakthrough optimizations in the following aspects, effectively solving the OT domain security blind spot, data imbalance, and detection delay problems described in the background art:
[0084] 1) Cross-IT / OT global collaborative detection to fill the OT domain security blind spot
[0085] An innovative detection model based on the RASM architecture (residual attention graph network) is constructed to realize the real-time detection of injection threats in the internal traffic of the OT domain for the first time. Compared with traditional IT domain protection solutions, this method increases the detection rate of compromised nodes in the OT domain by 10.12%, and the traceability time is shortened to 172 seconds (traditional method > 304 seconds), meeting the low-latency requirements of key industrial control scenarios.
[0086] Through the dynamic heterogeneous graph construction technology, integrate IT domain logs and OT domain industrial control protocol traffic (such as Modbus, OPCUA), realize cross-domain correlation analysis of attack links, block the lateral movement path of APT attacks, and increase the defense coverage by 2.8 times.
[0087] 2) Small-sample anomaly recognition against data imbalance to optimize model robustness
[0088] Propose an attention-enhanced SAGE network, combine layer normalization (LayerNorm), residual connection and attention mechanism, solve the problems of gradient disappearance and overfitting in large-scale graph data training, and significantly enhance the model's ability to recognize abnormal patterns. In an extreme scenario where the normal traffic accounts for 97.21%, the false alarm rate of the model is reduced to 3.2% (traditional method > 20.39%), and the alarm fatigue is reduced by 56%.
[0089] The adaptive decision-making mechanism based on entropy threshold dynamically adjusts the detection threshold to improve detection accuracy and robustness and adapt to the complex and changeable industrial environment. Quantify the traffic distribution shift through KL divergence, so that the model still maintains a stability of F1-score ≥ 0.973 when the data distribution changes (such as device upgrade or network topology adjustment).
[0090] 3) Real-time detection of highly concealed attacks to break through the bottleneck of temporal behavior analysis
[0091] Adopt multi-level neighborhood sampling and gradient accumulation techniques to balance computational efficiency and representation ability, and improve the performance of the model on large-scale heterogeneous traceability graphs. Double the training efficiency of large-scale industrial heterogeneous graphs, and at the same time ensure the detection accuracy (Recall = 98.3%) of microsecond-level injection instructions (such as malicious PLC code).
[0092] Innovatively introduce an elastic time window mechanism, recursively verify abnormal nodes through a sliding window, reduce the recognition delay of multi-hop relay attacks from hours to within 180s, and avoid the risk of critical service interruption. The elastic time window mechanism designed by the present invention performs a complete verification of abnormal nodes, further reducing the false alarm rate and improving the detection reliability.
[0093] 4) Improvement of computational efficiency and industrial scenario adaptability
[0094] Through the design of lightweight graph convolution kernels, the inference time of the model on typical industrial control devices (such as Siemens S7-1500 PLC) is only 8ms / sample, and the memory occupancy is controlled within 700MB, suitable for edge deployment.
[0095] Tests on 3 industrial-level datasets (including power grid, industrial control IT domain, and intelligent manufacturing scenarios) show that the comprehensive performance of this solution has significant advantages compared with traditional methods (such as ResGCN, GraphSAGE baselines):
[0096] Index Traditional method The present invention Improvement range Detection accuracy 82.11% 98.69% +16.58% False alarm rate 20.39% 3.2% -17.19% Tracing delay 304s 172s -43.42% Model training speed 1× 2× +100%
[0097] 5) Innovation of Security Protection Paradigm
[0098] Provide a new "detection - traceability - blocking" closed - loop paradigm for injection threat detection in the industrial Internet. By generating an attack knowledge graph (Attack KG) in real - time, it supports visual threat hunting and improves the security operation and maintenance efficiency by 60%.
[0099] In the deployment practice of a national - level critical infrastructure (a smart grid system), 7 APT attacks (including 2 zero - day vulnerability exploitations) were successfully intercepted, verifying the practical value of this method.
[0100] Through core technologies such as the RASM architecture, adaptive entropy threshold, and elastic time window, the present invention systematically solves three major problems faced by industrial control systems: the detection blind area of OT domain compromise, model failure caused by data imbalance, and lag in highly concealed attacks, providing a feasible technical path for constructing a highly reliable, low - latency, and adaptive industrial Internet in - depth defense system.
[0101] In summary, the injection threat detection method designed by the present invention has significant technological innovation and practical value. Through multi - dimensional sub - model training and integrated decision - making mechanism, it solves the blind - area problem of the existing technology in OT domain compromise threat detection; based on the recursive traversal and priority queue management of the real - time anomaly tracking mechanism, it realizes the accurate identification and traceability of concealed attack chains; the introduction of the elastic time window mechanism further improves the reliability of detection results. These technological innovations effectively fill the key gaps in the industrial control system security protection system, and have certain significance for blocking OT domain compromise threats and ensuring the security of industrial control network critical infrastructure. This solution is expected to provide a reference idea for the current security protection requirements of industrial control systems, and may also provide certain technical inspiration for the research and construction of the industrial Internet in - depth security defense system. Brief Description of the Drawings
[0102] Figure 1 It is the design architecture diagram of the injection threat detection model of the present invention.
[0103] Figure 2 It is the demonstration diagram of the injection threat invading through PLC of the present invention.
[0104] Figure 3 It is the traceability diagram of attack data in the present invention.
[0105] Figure 4 It is the design architecture diagram of the RASM model in the present invention. Detailed Description of the Invention
[0106] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0107] The industrial control system injection threat detection method proposed by the present invention designs an enhanced graph neural network architecture specifically for capturing abnormal traffic patterns and injection threats in industrial control networks. The system architecture of this method is shown in Figure 1 as follows.
[0108] Refer to Figure 1 , an injection threat detection system for an industrial control system, including a data acquisition module, a dynamic feature modeling module, a threat detection module, a multi-dimensional sub-model training and entropy-based determination module, and an abnormal tracking function module;
[0109] The data acquisition module converts the original industrial control network traffic into a graph structure representation and extracts node features;
[0110] The dynamic feature modeling module cleans, transforms, and structures the collected heterogeneous data to generate a dynamic traceability graph suitable for input to the graph neural network, and extracts multi-dimensional features of nodes and edges from the graph data for subsequent model training and detection;
[0111] The threat detection module is based on a three-layer network structure constructed by an improved SAGE network and is used to capture the global and local associations of threat events;
[0112] The multi-dimensional sub-model training and entropy-based determination module is used to deal with stealth attacks and recursively traverses and expands the associated subgraphs of suspicious nodes.
[0113] In this embodiment, the injection threat characteristics in the OT domain in the current industrial control Internet scenario are deeply analyzed, and an injection threat detection model is designed to actively discover internal compromised terminals. This model mainly discovers anomalies by performing feature dimensionality reduction extraction and detailed detection on injection attacks that bypass the IT domain protection and reach the OT domain. The scenarios that the model deals with are as shown in Figure 2 as follows:
[0114] Figure 2Shows the typical industrial control system (ICS) architecture and the intrusion path of the PLC. The bottom - layer field device layer consists of cutting devices, robots, processing devices, data acquisition devices, pressure sensors, motors, fans, etc. These devices are connected to the real PLC through the field bus; the control layer where the PLC is located is interconnected with the historical database and the engineer workstation through the industrial Ethernet, and then connected to the Web server through the Ethernet; at the same time, to achieve remote monitoring and management of the production environment, a remote management node is introduced in the IT / OT convergence architecture. This node is connected to both the field PLC, the historical database, and the engineer station, and also accesses the Internet. It is this practice of directly exposing the PLC to the public network that greatly expands the attack surface of the system - attackers can use PLC - CVE vulnerabilities, implant malicious code, or 0 - day attack methods to bypass the boundary protection and invade the field device layer. Once breached, it may lead to production interruption, equipment damage, and even serious physical security incidents.
[0115] An injection - type threat detection method for an industrial control system, comprising the following steps:
[0116] Step 1, data acquisition and dynamic heterogeneous graph feature modeling
[0117] In an industrial control system, injection - type threats may be manifested through network traffic or device logs. Therefore, data is collected separately from network traffic, system logs, and device status. The collected data is stored in a time - series form to ensure that it can reflect the dynamic changes of system behavior; the data acquisition module pre - processes the collected data, converting the streaming collected data into a traceability graph to obtain the graph data for model training and recognition; through the dynamic feature modeling module, abstract feature extraction is performed on the obtained graph data. Abstract feature extraction is to mine multi - dimensional features from the pre - processed graph data to enhance the model's ability to identify injection - type threats;
[0118] Data acquisition targets the following sources: network traffic, system logs, device status; through traffic capture devices deployed at key network nodes (such as zero - trust gateways), communication data of industrial protocols such as Modbus and OPC UA is collected; operation logs are extracted from devices such as PLCs, SCADA systems, and DCSs, recording events such as process calls, file accesses, and network connections; the operating status of industrial terminals (IITs), such as CPU usage rate and memory occupancy, is monitored as auxiliary data; the graph structure of the data is shown in Figure 3As shown in the figure, it shows a schematic diagram of a malicious link based on Nginx: the ellipse on the left represents three remote clients, which establish communication with the Nginx process through sockets respectively; the Nginx process reads the local / etc / group and / etc / passwd files, and writes the payload into / tmp / vUgefal; subsequently, Nginx starts the vUgefal process with executable permission. The vUgefal process first writes logs into / var / log / devc, and then initiates Sendto socket connections to three different remote C2 servers to complete data leakage or obtain subsequent instructions. Static nodes are represented by thick solid lines or solid ellipses, and dynamic nodes are represented by thin rhombuses. The whole link clearly depicts the entire process from initial intrusion to local persistence and then to communication with external control servers.
[0119] For the data preprocessing mentioned above, the goal of data preprocessing is to convert the original data into a graph structure suitable for input to the RASM model, which specifically includes the following steps:
[0120] Step1, data cleaning, removing duplicate records and invalid data (such as null value packets), filling in missing values using interpolation method, and removing outliers (based on the 3σ criterion);
[0121] Step2, entity recognition and relationship extraction, identifying entities in the stream - collected data, including processes, files, network connections, and device terminals, as nodes of the graph; extracting relationships between entities, such as inter - process communication (IPC), file reading and writing, network packet transmission, etc., as edges of the graph;
[0122] Step3, constructing a dynamic traceability graph, defining a set of node types and a set of edge types counting the number of node types and the number of edge types
[0123] Step4, type mapping, constructing a node type mapping function and an edge type mapping function to map node and edge types into the integer range, where N v is the number of node types, N e is the number of edge types;
[0124] Step5, feature initialization, assigning a feature vector to each node v ∈ V including in - edge and out - edge statistical features:
[0125] The feature components of in - edges:
[0126] a i = |{e ∈ In(v): i = Me (χ e (e))}|, i ∈ {0, …, N e -1}
[0127] Characteristic components of the outgoing edge:
[0128]
[0129] where In(v) represents all incoming edges with v as the target node, Out(v) represents all outgoing edges with v as the source node, χ e (e) represents the type of edge e, and M e is the edge type mapping function;
[0130] Step6, construct the complete feature representation of node v, where a i represents the in-out edge feature components of node v.
[0131] Step7, for each node v, extract local structure features based on three-hop neighborhood sampling:
[0132]
[0133] where N3(v) represents the three-hop neighborhood of node v, AGGREGATE uses mean aggregation, and x u represents a certain part of the node's features,
[0134] Step8, assign a weight ω e to each edge e ∈ E, based on the type and frequency of the edge: ω e = log(1 + freq(e)), where freq(e) represents the number of occurrences of edge e within the time window;
[0135] Step9, introduce time dimension features Record the timestamp distribution of node behaviors:
[0136]
[0137] where Δt is the time interval of events associated with node v, and mean and std represent the mean and standard deviation;
[0138] Step10, output the finally abstracted node features:
[0139]
[0140] Step 2, construct an attention-enhanced graph neural network injection threat detection model
[0141] To more comprehensively and deeply learn the heterogeneous graph constructed from logs with enhanced abstract features and strengthen the module's ability to focus on high-value features in traffic logs, key technologies such as the attention mechanism, residual connection, and layer normalization are introduced based on the SAGE model to construct the RASM model (see Figure 4 ). Node features are generated by aggregating in-edge and out-edge statistical information. An adaptive threshold mechanism based on Shannon entropy is introduced by integrating local attributes and relationship information to dynamically adjust the detection criteria: by calculating the entropy of the predicted probability distribution of each node, low-entropy nodes (high confidence) are used to estimate the optimal threshold for distinguishing normal and abnormal behaviors. This method enhances the model's adaptability to different data distributions and improves its generality. In real-time detection, RASM expands the k-hop neighborhood of suspicious nodes by recursively traversing the adjacency list to capture attack chains across entities, which is particularly suitable for discovering hidden coordinated malicious behaviors. At the same time, a priority queue manages suspicious nodes, combined with a time window (such as 10 minutes) to wait for more context information, reducing the false alarm rate and ensuring that only continuous anomalies trigger alerts.
[0142] The architecture of the RASM model can not only efficiently process large-scale complex data but also detect multi-step injection threats in real time. Through dynamic neighborhood expansion and entropy-based threshold adjustment, RASM effectively responds to complex attack patterns, ensuring high accuracy and a low false alarm rate.
[0143] The described attention-enhanced graph neural network injection threat detection model mines multi-dimensional information from graph data after abstract feature preprocessing, introduces residual connection and attention mechanism, and significantly improves feature aggregation and detection performance, specifically including the following steps:
[0144] Step-1, adopt the GraphSAGE graph convolution framework to construct a three-layer graph convolution network (SAGE), and layer by layer realize the aggregation and update of node features, providing a hierarchical processing architecture for feature extraction of heterogeneous graph data.
[0145] Step-2, embed a residual connection (Residual Connection) in the three-layer graph convolution network. Through cross-layer feature direct connection, the original input information of the previous layer network is retained, avoiding the problems of gradient disappearance and performance degradation caused by the increase in network depth, and ensuring the effectiveness of feature transmission in the deep network.
[0146] Step-3, introduce an attention mechanism during the neighbor node aggregation process. By calculating the importance weights between nodes, key neighbor nodes related to threat detection are dynamically screened, suppressing the interference of noise features and enhancing the model's ability to capture abnormal patterns.
[0147] Step-4. After each layer of graph convolution operation, add layer normalization to standardize the node features, reduce the risk of gradient vanishing by stabilizing the feature distribution, and improve the training stability.
[0148] Step-5. Replace the traditional ReLU with the Leaky ReLU activation function to solve the problem of neuron inactivation in the negative interval in the sparse data scenario, and enhance the model's non-linear expression ability for low-frequency abnormal features (such as rare attack patterns).
[0149] Step-6. Apply the cosine annealing strategy to dynamically adjust the learning rate during the training phase. By periodically decaying the learning rate, balance the model's global exploration and local fine-tuning capabilities in the solution space.
[0150] Step-7. Monitor the detection metrics (such as F1 score, AUC) on the validation set. When the metrics do not improve for N consecutive epochs, trigger the early stopping mechanism to terminate the training in advance. This mechanism avoids overfitting and reduces redundant computational consumption by dynamically evaluating the model's generalization performance.
[0151] Step 3. Multi-dimensional sub-model training and entropy threshold judgment
[0152] In view of the fact that abnormal behaviors in complex systems usually cover multiple different dimensions (such as system calls, process communication, and network interaction, etc.), a single detection model is often difficult to comprehensively capture the information of these dimensions. The multi-dimensional sub-model training and entropy threshold judgment module is used to solve the problem of insufficient in-depth learning of the traceability graph dimension. The multi-dimensional sub-model training architecture divides the original graph data into multiple subgraphs according to different dimensions, constructs a dedicated sub-model for each subgraph, calculates the information entropy value of the prediction results of each sub-model at the same time, and constructs a threshold discrimination mechanism based on information entropy. Specifically, first convert the prediction probability distribution of the model into a Shannon entropy quantization index. When the entropy value is lower than the preset threshold, adopt the prediction result, otherwise reject it. Finally, verify the anomaly detection results through the multi-model fusion method to improve the overall detection performance and generalization ability.
[0153] The above-mentioned multi-dimensional sub-model training and entropy threshold determination specifically include the following steps:
[0154] Step 3-1. First, divide the original graph data G=(V, E) into multiple subgraphs G d =(V, E d ) according to the behavior dimension, where V is the node set and E is the edge set. The behavior dimension includes: system call, process communication, network interaction, and E d represents the edge set of a specific behavior dimension; through the edge type mapping function transform the behavior features into one-dimensional indexes, and R is the set of edge types for easy processing. For each subgraph Gd , calculate the node feature vector featd(v). The node feature vector featd(v) contains the statistical information of the incoming and outgoing edges. The statistical information includes: the number or weight of the edges; in order to focus on the unrecognized abnormal patterns, reconstruct the complex features, remove the known parts, and ensure that the subsequent steps of training are carried out for the features of specific dimensions;
[0155] Step 3-2, for each subgraph G d , train a dedicated sub-model M d , the architecture of the sub-model adopts the enhanced graph neural network architecture RASM. The goal of the sub-model M d is to learn the normal behavior pattern in this dimension and detect abnormal behaviors that deviate from the normal system operation behavior pattern. By focusing on the signal capture of a single dimension, the sub-model M d can improve the accuracy of the overall detection;
[0156] Step 3-3, multi-dimensional label calculation and entropy threshold determination
[0157] Each sub-model M d outputs the predicted probability distribution P d (v)=[p d,1 ,p d,2 ,…,p d,K , where K is the number of classes. To measure the uncertainty of the prediction, calculate the entropy value:
[0158]
[0159] where, p d,k represents the node prediction probability. Normalize it to get:
[0160]
[0161] In the formula, 0 represents complete certainty and 1 represents the maximum uncertainty. In the pre-training stage, use the entropy value distribution H val of the validation set V d ={H d (v i )∣v i ∈V val}, select the 80th percentile Q 0.8 (H d ) as the initial entropy threshold R t,d , and then adjust the threshold through the dynamic update formula:
[0162] R t_new =λR t_old +(1 - λ)R t_estimated
[0163] where λ is the smoothing factor to ensure that the threshold adapts to data changes, and R t_estimated represents the entropy threshold estimated based on the latest data currently, and R t_old represents the threshold that was in effect at the previous moment;
[0164] Step3-4, Node classification and dynamic mask evolution. Classification is performed based on the comparison between the normalized entropy and the threshold obtained in Step 3-3:
[0165] If H norm,d (v) ≤ R t,d , accept the prediction of M d , and select the class with the highest probability argmax i p d,i as the result; if H norm,d (v) > R t,d , mark the node v as unclassified, with a predicted value of -1, to be processed later; initialize the behavior status matrix S (0) = [1, 1, …, 1], indicating that all nodes participate in training. After each round of training, update the status according to the confusion matrix CM (t) : If node i belongs to false positive FP or true negative TN, then it will no longer participate in training; retain the true positive TP and false negative FN nodes to make the model gradually focus on key behaviors;
[0166] Step3-5, Construct an optimized training subgraph through a two-stage neighborhood sampling strategy. In the first stage, sample the direct neighbors of the nodes, and in the second stage, further sample highly relevant neighbors from these neighbors to form a subgraph containing multi-layer neighborhood information, and provide the subgraph structure to the sub-model M d for training to enhance the model's ability to understand and capture complex relationships;
[0167] Step3-6, Calculate the performance evaluation score for each sub-model M d :
[0168]
[0169] Set a double-threshold condition: Precision P(M d ) ≥ γ p , Recall R(M d ) ≥ γ r , select the sub-model with the largest Q(M d ) as the optimal sub-model, and then fuse the prediction results of multiple optimal sub-models to comprehensively judge the abnormal state of the nodes and improve the detection comprehensiveness;
[0170] Step 4, Implement anomaly tracking
[0171] When the RASM model discovers a suspicious node, due to the anomaly-based detection feature, in order to reduce the false alarm rate and at the same time track possible malicious nodes, the waiting queue algorithm of the anomaly tracking function module is adopted. By reasonably extending the detection period of the suspicious node, the occurrence of false alarm reports is effectively controlled; the algorithm of the waiting queue includes the following steps:
[0172] Step4-1, Initialize the priority queue Q, which contains all nodes v with abnormal scores S(v)>θ, where θ is a preset threshold and S(v) is the node scoring function. At the same time, initialize the abnormal node set V abnormal to be empty, which is used to store the finally confirmed abnormal nodes;
[0173] Step4-2, Traverse each node v in the priority queue Q, set its timestamp t(v) to the current time current_time to record the moment when the node enters the queue, providing a basis for subsequent time window judgment;
[0174] Step4-3, As long as the priority queue QQ is not empty, continuously perform the following operations: Take out a node v from the priority queue QQ, and check the k-hop neighborhood N k (v) of the node v to see if it has an intersection with the normal node set V normal ; If there exists at least one node u∈N k (v)∩V normal , it means that the neighborhood of v contains normal behavior, then remove v from the queue Q and do not process it further;
[0175] Step4-4, If the k-hop neighborhood N k (v) of the node v has no intersection with V normal , expand its neighborhood by one level and update it to N k+1 (v) to check for a wider range of associated behaviors; Then, judge whether the difference between the current time t current and the timestamp t(v) of the node v exceeds the preset time window T. If t current -t(v)>T, it means that v has not been classified as normal within the time window, then add it to the abnormal node set V abnormal , otherwise, add v back to the queue Q and wait for the next round of inspection;
[0176] Step4-5, When the priority queue Q is empty, the algorithm terminates and returns the abnormal node set V abnormal , which contains all nodes that have not been classified as normal within the time window T, and these nodes are confirmed as malicious nodes.
[0177] List of abbreviations and definitions:
[0178] IT: Information Technology, Information Technology;
[0179] IC: Industrial Control, Industrial Control;
[0180] OT: Operational Technology, Operational Technology;
[0181] IIT: Industrial Internet Terminal, Industrial Internet Terminal;
[0182] APT: Advanced Persistent Threat, Advanced Persistent Threat;
[0183] IPS: Intrusion Prevention System, Intrusion Prevention System;
[0184] SAGE: Strategic Automated Ground Environment, Strategic Automated Ground Environment;
[0185] PLC: Programmable Logic Controller, Industrial Automation Control;
[0186] DCS: Distributed control system, Distributed control system;
[0187] RASM: Residual Attention SAGE Misuse-detection, Residual Attention SAGE Misuse-detection.
Claims
1. An injection threat detection system for an industrial control system, characterized in that It includes a data acquisition module, a dynamic feature modeling module, a threat detection module, a multi-dimensional sub-model training and entropy-based determination module, and an anomaly tracking function module; The data acquisition module converts the original industrial control network traffic into a graph structure representation and extracts node features; The dynamic feature modeling module cleans, transforms, and structurally processes the collected heterogeneous data to generate a dynamic traceability graph suitable for input to a graph neural network, and extracts multi-dimensional features of nodes and edges from the graph data for subsequent model training and detection; The threat detection module is based on a three-layer network structure constructed by an improved SAGE network and is used to capture the global and local associations of threat events; The multi-dimensional sub-model training and entropy-based determination module is used to deal with stealth attacks and recursively traverses and expands the associated subgraphs of suspicious nodes.
2. An injection threat detection method for an industrial control system, characterized in that, It includes the following steps: Step 1, Data acquisition and dynamic heterogeneous graph feature modeling Data is collected from network traffic, system logs, and device status respectively, and the collected data is stored in a time series form to ensure that it can reflect the dynamic changes of system behavior; The data acquisition module preprocesses the collected data, converts the streaming collected data into a traceability graph to obtain the graph data for model training and recognition; The dynamic feature modeling module extracts abstract features from the obtained graph data. The abstract feature extraction is to mine multi-dimensional features from the preprocessed graph data to enhance the model's ability to identify injection threats; Step 2, Build an injection threat detection model for a graph neural network with enhanced attention On the basis of the SAGE model, key technologies such as attention mechanism, residual connection, and layer normalization are introduced. Node features are generated by aggregating in-edge and out-edge statistical information, and local attributes and relationship information are fused to lay the foundation for subsequent detection; Step 3, Multi-dimensional sub-model training and entropy threshold evaluation The multi-dimensional sub-model training and entropy threshold evaluation module is used to solve the problem that the dimension learning of the traceability graph is not deep enough. The multi-dimensional sub-model training architecture divides the original graph data into multiple subgraphs according to different dimensions, constructs a dedicated sub-model for each subgraph, and calculates the information entropy value of the prediction results of each sub-model at the same time, and constructs a threshold discrimination mechanism based on information entropy. Specifically, first convert the prediction probability distribution of the model into a Shannon entropy quantization index. When the entropy value is lower than the preset threshold, adopt the prediction result, otherwise reject it. Finally, verify the anomaly detection results through the method of multi-model fusion to improve the overall detection performance and generalization ability; Step 4, Implement anomaly tracking When the RASM model discovers a suspicious node, due to the anomaly-based detection characteristics, in order to reduce the false alarm rate and at the same time track possible malicious nodes, the waiting queue algorithm of the anomaly tracking function module is adopted, and the detection period of the suspicious node is reasonably extended to effectively control the occurrence of false alarm reports.
3. The injection threat detection method for an industrial control system according to claim 2, characterized in that, In the said Step 1, the goal of data preprocessing is to convert the original data into a graph structure suitable for input to the RASM model, and the goal of dynamic feature modeling is to extract the abstract features of the graph structure. Specifically, it further includes the following steps: Step1, Data cleaning, removing duplicate records and invalid data, using interpolation method to fill in missing values, and removing outliers; Step 2, entity recognition and relationship extraction, identify entities in the streaming data as the nodes of the graph; extract the relationships between entities as the edges of the graph; Step3, Construction of dynamic traceability graph, defining a set of node types and a set of edge types Count the number of node types and the number of edge types Step4, Type mapping, constructing node type mapping function and edge type mapping function Map node and edge types to the integer range, where N v is the number of node types, N e is the number of edge types; Step5, Feature initialization, assign a feature vector to each node v∈V Including in-edge and out-edge statistical features: Feature components of incoming edges: a i = |{e ∈ In(v) : i = M e (χ e (e))}|, i ∈ {0, …, N e - 1} Feature components of outgoing edges: where In(v) represents all incoming edges with v as the target node, Out(v) represents all outgoing edges with v as the source node, and χ e (e) represents the type of edge e, and M e is the edge type mapping function; Step6, construct the complete feature representation of node v, where a i represents the in-out edge feature component of node v; Step 7, for each node v, extract local structural features based on three-hop neighborhood sampling: Among them, N3(v) represents the three-hop neighborhood of node v, AGGREGATE adopts mean aggregation, and x u represents a certain part of the node's features; Step 8, assign a weight ω to each edge e ∈ E e , based on the type and frequency of the edge: ω e = log(1 + freq(e)), where freq(e) represents the number of occurrences of edge e within the time window; Step9, introduce time dimension features Record the timestamp distribution of node behaviors: where Δt is the time interval of the events associated with node v, and mean and std represent the mean and standard deviation; Step 10, output the final node features:
4. The injection threat detection method for an industrial control system according to claim 2, characterized in that, The described Step 2 specifically includes the following steps: Step - 1, adopt the GraphSAGE graph convolution framework to construct a three-layer graph convolution network (SAGE), and gradually realize the aggregation and update of node features, providing a hierarchical processing architecture for the feature extraction of heterogeneous graph data; Step - 2, embed residual connections in the three-layer graph convolution network, and directly connect the cross-layer features to retain the original input information of the previous layer network, avoiding the problems of gradient disappearance and performance degradation caused by the increase in network depth, and ensuring the effectiveness of feature transmission in the deep network; Step - 3, introduce an attention mechanism in the neighbor node aggregation process, dynamically screen key neighbor nodes related to threat detection by calculating the importance weights between nodes, suppress the interference of noise features, and enhance the model's ability to capture abnormal patterns; Step - 4, add layer normalization after each layer of graph convolution operation to standardize the node features, reduce the risk of gradient disappearance by stabilizing the feature distribution, and improve the training stability; Step - 5, use the Leaky ReLU activation function to replace the traditional ReLU to solve the problem of neuron inactivation in the negative interval in the sparse data scenario, and enhance the model's non-linear expression ability for low-frequency abnormal features; Step - 6, apply the cosine annealing strategy to dynamically adjust the learning rate during the training stage, balance the global exploration and local fine-tuning capabilities of the model in the solution space by periodically decaying the learning rate; Step - 7, monitor the detection metrics on the validation set, and trigger the early stopping mechanism when the metrics have not improved for N consecutive epochs, and terminate the training in advance.
5. The injection threat detection method for an industrial control system according to claim 2, characterized in that The described multi-dimensional sub-model training and entropy threshold determination module specifically includes the following steps: Step 3-1: First, split the original graph data G = (V, E) into multiple sub-graphs G according to the behavior dimension. d =(V,E d ), where V is the node set, E is the edge set, and the behavior dimensions include: system calls, process communications, network interactions, and E d Represents a set of edges of a specific behavior dimension; through edge type mapping function The behavior features are converted into one-dimensional indexes, R is a set of edge types, which is easy to process, and for each subgraph G d , calculate the node feature vector feat d(v), which contains the statistical information of the incoming and outgoing edges, including the number or weight of the edges; in order to focus on the unidentified abnormal patterns, reconstruct the complex features and remove the known parts to ensure that the subsequent steps of training are carried out for the features of specific dimensions; Step 3-2, for each sub-graph G d , train a dedicated sub-model M d . The architecture of all sub-models adopts the enhanced graph neural network architecture RASM. The goal of sub-model M d is to learn the normal behavior pattern in this dimension and detect abnormal behaviors that deviate from the normal system operation behavior pattern. By focusing on signal capture in a single dimension, sub-model M d can improve the accuracy of overall detection; Step 3 - 3, multi-dimensional label calculation and entropy threshold determination Each sub-model M d outputs the predicted probability distribution P d (v) = [p d,1 , p d,2 , …, p d,K , where K is the number of classes. To measure the uncertainty of the prediction, the entropy value is calculated as follows: where p d,k represents the node prediction probability, which is obtained by normalizing H(v): where 0 represents complete certainty and 1 represents maximum uncertainty. During the pre-training phase, the entropy value distribution H val of the validation set V d ={H d (v i )|v i ∈V val}, and the 80th percentile Q 0.8 (H d ) is selected as the initial entropy threshold R t,d . Subsequently, the threshold is adjusted through a dynamic update formula: R t_new = λR t_old + (1 - λ)R t_estimated where λ is the smoothing factor to ensure that the threshold adapts to data changes, and R t_estimated represents the entropy threshold estimated based on the latest data currently, and R t_old represents the threshold that was in effect at the previous moment; Step 3 - 4, node classification and dynamic mask evolution, classify according to the comparison between the normalized entropy and the threshold obtained in Step 3 - 3: If H norm,d (v) ≤ R t,d , accept the prediction of M d and select the class with the highest probability argmax i p d,i as the result; if H norm,d (v) > R t,d , mark the node v as unclassified, with a predicted value of -1, to be processed later; initialize the behavior status matrix S (0) = [1, 1, …, 1], indicating that all nodes participate in training. After each round of training, update the status according to the confusion matrix CM (t) : if node i belongs to false positive FP or true negative TN, then it will no longer participate in training; retain the true positive TP and false negative FN nodes to make the model gradually focus on key behaviors; Step 3-5: Construct an optimized training subgraph through a two-stage neighborhood sampling strategy. In the first stage, directly sample the neighbors of the nodes. In the second stage, further sample highly relevant neighbors from these direct neighbors to form a training subgraph containing multi-layer neighborhood information, and provide the training subgraph structure to the sub-model M d Train to enhance the model's ability to understand and capture complex relationships; Step 3-6, for each sub-model M d Calculate the performance evaluation score: Setting dual threshold conditions: Precision P(M d )≥γ p , recall rate R(M d )≥γ r , select Q(M d )The largest sub-model As the optimal sub-model, the prediction results of multiple optimal sub-models are integrated to comprehensively judge the abnormal status of the node and improve the comprehensiveness of detection.
6. The injection threat detection method for an industrial control system according to claim 2, characterized in that The described waiting queue, its algorithm includes the following steps: Step4-1, Initialize the priority queue Q, which contains all nodes v with abnormal scores S(v) > θ, where θ is a preset threshold and S(v) is the node scoring function. At the same time, initialize the set V of abnormal nodes abnormal to be empty, which is used to store the finally confirmed abnormal nodes; Step 4 - 2, traverse each node v in the priority queue Q, set its timestamp t(v) to the current time current_time to record the moment when the node enters the queue, providing a basis for subsequent time window judgment; Step 4-3, as long as the priority queue Q is not empty, continuously perform the following operations: Take out a node v from the priority queue Q and check whether the k-hop neighborhood N k (v) has an intersection with the set of normal nodes V normal ; if there exists at least one node u ∈ N k (v) ∩ V normal , it means that the neighborhood of v contains normal behavior, then remove v from the queue Q and do not process it further; Step4-4, if the k-hop neighborhood N k (v) of node v has no intersection with V normal , expand its neighborhood by one level and update it to N k+1 (v) to check for association behavior in a wider range; then, determine whether the difference between the current time t current and the timestamp t(v) of node v exceeds the preset time window T. If t current - t(v) > T, it means that v has not been classified as normal within the time window, so add it to the abnormal node set V abnormal , otherwise, add v back to the queue Q and wait for the next round of checks; Step 4-5, when the priority queue Q is empty, the algorithm terminates and returns the set of abnormal nodes V abnormal , which contains all the nodes that have not been classified as normal within the time window T, and these nodes are identified as malicious nodes.
7. The injection threat detection method for an industrial control system according to claim 2, characterized in that The RASM model adopts a multi-level neighbor sampling strategy (such as sampling widths [30, 20, 10]), limits the number of neighbors in each layer, not only retains the key structural information but also improves the calculation efficiency, adapting to the large-scale data of industrial systems; The RASM model adopts a three-layer graph convolution network, retains the information of the previous layer through residual connections, and avoids performance degradation caused by the increase in depth; The described RASM model introduces an adaptive threshold mechanism based on Shannon entropy to dynamically adjust the detection criteria: by calculating the entropy of the predicted probability distribution of each node, low-entropy nodes are used to estimate the optimal threshold for distinguishing normal and abnormal behaviors.
Citation Information
Cited By
Data annotation method and system based on user behavior and attention tracking
CN120929832A
Data transmission security method based on graph nerve detection
CN121462302A
A data transmission security method based on graph neural detection
CN121462302B