Industrial internet security situation analysis method and system based on support vector regression

Through time window sliding analysis and support vector regression model, homologous alarm events are identified and attack scene diagrams are constructed, which solves the problem of inaccurate situation evaluation in the existing technology, realizes the generation of precise protection strategies and resource optimization, and improves the security of the industrial Internet.

CN120498903AActive Publication Date: 2025-08-15NAT IND INFORMATION SECURITY DEV RES CENT

Patent Information

Application Number
CN202510971762.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-08-15
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

The existing industrial Internet security situation analysis methods lack the ability to deeply mine and correlate the analysis of alarm data, making it difficult to accurately identify homologous alarm events, resulting in insufficient accuracy of situation evaluation results, insufficient targeted protection strategies, and ineffective quantification of security situations and automatically generate targeted protection strategies.

Method used

The time window sliding method is used to conduct correlation analysis of the alarm data, identify homologous alarm events, establish a time series propagation matrix and attack scene diagram, calculate the security situation index using the support vector regression model, and select the optimal protection node combination based on node centrality and protection decision tree to generate protection instructions.

Benefits of technology

Accurate analysis and evaluation of the security situation of the industrial Internet has been achieved, targeted and efficient protection strategies have been improved, protection resource allocation has been optimized, and overall security has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120498903A_ABST
    Figure CN120498903A_ABST
Patent Text Reader

Abstract

The invention provides an industrial internet security situation analysis method and system based on support vector regression, and relates to the technical field of industrial internet security, and the method comprises the steps: collecting alarm data, and recognizing a homologous alarm event through a time window sliding mode; establishing a time sequence propagation matrix, calculating a propagation probability, constructing an attack scene graph and extracting features; calculating a security situation index by using a support vector regression model; and selecting an optimal protection node combination based on the node centrality and the protection decision tree. According to the method, the threat propagation path can be accurately identified, the security situation can be quantified, and the precise protection of the industrial internet environment can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to industrial Internet security technology, and in particular to an industrial Internet security situation analysis method and system based on support vector regression. Background Art

[0002] Existing industrial internet security situation analysis methods primarily rely on expert experience and static rules, lacking the ability to deeply mine and analyze alarm data, making it difficult to accurately identify causally related homologous alarm events. Furthermore, traditional methods fail to fully consider network topology and alarm propagation characteristics when constructing attack scenario diagrams and assessing security situation, resulting in inaccurate situation assessment results and weakly targeted protection strategies.

[0003] In the Industrial Internet, the complex inter-device connections and diverse attack vectors make security assessment and defense decision-making extremely complex. Existing technologies are unable to effectively quantify security trends and automatically generate targeted defense strategies, making it difficult to rapidly respond to attacks and proactively defend against them. Furthermore, with limited defense resources, optimizing defense node selection to maximize protection effectiveness remains a pressing issue. Therefore, a method that can intelligently analyze alert data, accurately assess security trends, and automatically generate defense strategies is urgently needed to enhance the security capabilities of the Industrial Internet. Summary of the Invention

[0004] The embodiments of the present invention provide an industrial Internet security situation analysis method and system based on support vector regression, which can solve the problems in the existing technology.

[0005] A first aspect of an embodiment of the present invention provides an industrial Internet security situation analysis method based on support vector regression, comprising: Collect alarm data from industrial Internet devices and, based on the time series properties of the alarm data and the network topology between industrial Internet devices, perform correlation analysis on the alarm data using a time window sliding method to identify homologous alarm events with causal relationships. Based on homologous alarm events, a time-series propagation matrix is established to calculate the alarm propagation probability between industrial Internet devices and construct a diffusion path. An attack scenario graph is constructed based on the overlap and coverage of multiple diffusion paths. The propagation characteristics of the alarm data and the number of affected devices are extracted from the attack scenario graph to generate a situation assessment dataset. Input the situation assessment dataset into the pre-trained support vector regression model to calculate the security situation index, determine the protection priority, and formulate a protection strategy based on the diffusion path in the attack scenario graph; Based on the attack scenario graph, the node centrality of each industrial Internet device is calculated, and a protection decision tree is constructed in combination with the protection strategy. The optimal protection node combination is selected based on the node centrality and the protection decision tree, and protection instructions are generated. The protection instructions are sent to the affected industrial Internet devices to implement active protection.

[0006] In an optional embodiment, Collect alarm data from industrial Internet devices. Based on the time series properties of the alarm data and the network topology between industrial Internet devices, use a time window sliding method to perform correlation analysis on the alarm data to identify homologous alarm events with causal relationships, including: Collect alarm data from industrial Internet devices, set a corresponding reference time window according to the alarm level of the alarm data, and construct a network connection diagram based on the network topology relationship between industrial Internet devices, the network connection diagram including device nodes and communication links; An attention-based density clustering algorithm is used to calculate the alarm density of the alarm data, and a reference time window is adaptively adjusted according to the alarm density and the number of shortest path hops between device nodes in a network connection graph to obtain a time window. Within the time window, calculating the time series similarity based on the time series attributes of the alarm data, calculating the path similarity based on the shortest path between the device nodes in the network connection graph, and calculating the feature similarity based on the attribute features of the alarm data; Calculating a propagation weight coefficient of the alarm data, performing a weighted combination of the time series similarity, path similarity, and feature similarity according to the propagation weight coefficient to obtain an alarm propagation probability, and constructing an alarm propagation graph, wherein nodes are alarm data and edges are alarm propagation relationships with alarm propagation probabilities; Based on the recursive search algorithm, the alarm propagation graph is dynamically verified for paths, the alarm propagation relationships that do not meet the propagation timing constraints are identified and eliminated, and multiple connected subgraphs are obtained. The connected subgraphs are then verified for timing consistency to identify homologous alarm events with causal relationships.

[0007] In an optional embodiment, Dynamic path verification is performed on the alarm propagation graph based on a recursive search algorithm to identify and eliminate alarm propagation relationships that do not meet propagation timing constraints, thereby obtaining multiple connected subgraphs. Temporal consistency verification is performed on the connected subgraphs to identify homologous alarm events with causal relationships, including: Calculating a propagation delay attenuation coefficient according to the number of network hops between the alarm data, and determining a propagation timing constraint between adjacent alarm data based on the propagation delay attenuation coefficient; Constructing an alarm propagation rule based on the alarm type of the alarm data, preprocessing the alarm propagation graph according to the alarm propagation rule, deleting the alarm propagation relationship that does not meet the alarm propagation rule, and obtaining a preprocessed alarm propagation graph; Perform a depth-first search on the source alarm data in the pre-processed alarm propagation graph. During the search, determine whether the time intervals between adjacent alarm data satisfy the propagation timing constraints, and mark the alarm data that violates the propagation timing constraints as invalid alarm data. Skipping the path search starting from the invalid alarm data, updating the reachability mark of the alarm data, and saving the alarm propagation relationship that meets the propagation timing constraint as a valid alarm propagation relationship; According to the connectivity of the effective alarm propagation relationship, the preprocessed alarm propagation graph is decomposed into multiple connected sub-graphs, and the propagation probability of the alarm propagation relationship in the connected sub-graph and the timing consistency index of the alarm timestamp are calculated. Based on the timing consistency index, the connected sub-graph is verified for timing consistency to identify homologous alarm events with causal relationships.

[0008] In an optional embodiment, Based on the same-source alarm events, a time series propagation matrix is established to calculate the alarm propagation probability between industrial Internet devices. The diffusion path is constructed, including: Construct an initial time series propagation matrix based on homologous alarm events. Use the industrial Internet devices in the homologous alarm events as row vectors and column vectors. Count the number of alarm propagations between each pair of devices within a specified time window. For any two industrial Internet devices, calculate the alarm propagation probability based on the ratio of the number of target device alarms generated after the source device generates an alarm to the total number of source device alarms. Fill the alarm propagation probability into the corresponding matrix element position to generate a time series propagation matrix. The alarm propagation probability less than the preset probability threshold is set to zero to obtain a filtered time series propagation matrix. Based on the filtered time series propagation matrix, a depth-first search method is used to search for the alarm propagation path starting from the first alarm device. The alarm propagation probability on each path is multiplied as the path weight, and a weight threshold is set for the path weight. The paths greater than the weight threshold are retained as diffusion paths.

[0009] In an optional embodiment, An attack scenario graph is constructed based on the overlap and coverage of multiple diffusion paths. The propagation characteristics of the alarm data and the number of affected devices are extracted from the attack scenario graph to generate a situation assessment dataset including: Constructing a path connection relationship table based on the diffusion path, wherein the path connection relationship table records the device node sequence passed by each diffusion path, and counting the number of repeated device nodes in the path connection relationship table to obtain the path overlap degree; The device node sequence of each diffusion path is traversed to obtain the total number of nodes accessible by the path, and the ratio of the total number of accessible nodes to the total number of network nodes is used as the path coverage; Divide the diffusion paths with the same repeated nodes into the same group according to the degree of path overlap, calculate the coverage of the paths in each group, and select the path with the largest coverage to construct an attack scenario graph; Identify the propagation feature nodes in the attack scenario graph, extract the propagation features based on the connection relationship between the propagation feature nodes, count the number of device nodes reachable by the propagation feature nodes as the number of affected devices, and generate a situation assessment data set.

[0010] In an optional embodiment, The situation assessment dataset is input into the pre-trained support vector regression model to calculate the security situation index, determine the protection priority, and formulate a protection strategy based on the diffusion path in the attack scenario graph, including: Extract the alarm propagation feature sequence and the affected device number sequence from the situation assessment dataset, calculate the standardized feature based on the maximum and minimum values of the alarm propagation feature sequence, and calculate the normalized feature based on the ratio of the affected device number sequence to the total number of devices; The standardized features and the normalized features are combined to construct a feature matrix, the mutual information values between the features in the feature matrix are calculated to obtain feature redundancy, features are selected based on the feature redundancy and information gain is calculated, and training features are obtained when the information gain is less than a preset gain threshold; Dividing the training features into a training set and a validation set and performing cross-validation to obtain optimal kernel function parameters, inputting the training features and kernel function parameters into a support vector regression model, calculating an error value between a predicted result and a true result, updating the model parameters according to the error value, and obtaining a trained support vector regression model when the error value is less than a preset error threshold; Inputting the situation assessment data set into the trained support vector regression model to obtain a security situation index, and performing weighted smoothing processing on the security situation index within a preset time window to obtain a smoothed situation index; The difference in smoothed situation index between adjacent time points is calculated to obtain the situation change rate. The protection priority is determined based on the situation change rate. The node weight is calculated based on the time sequence of the nodes visited on the diffusion path in the attack scenario graph. The protection priority is adjusted to obtain a protection node sequence. A protection strategy is formulated based on the protection node sequence and the node access timing of the diffusion path.

[0011] In an optional embodiment, Based on the attack scenario graph, the node centrality of each industrial Internet device is calculated. A protection decision tree is constructed in combination with the protection strategy. The optimal protection node combination is selected based on the node centrality and the protection decision tree. The generated protection instructions include: The degree centrality is calculated by obtaining the alarm reception and propagation volume of each industrial Internet device node in the attack scenario graph. The betweenness centrality is obtained by calculating the number of alarm forwarding times of the node based on the attack scenario graph. The closeness centrality is obtained by calculating the shortest hop count between nodes. Extracting the alarm processing time sequence of the node in the attack scenario graph, calculating the alarm propagation delay of the node, adaptively weighting the degree centrality, betweenness centrality, and closeness centrality according to the propagation delay, and combining them to obtain the node centrality; Based on the node centrality, the industrial Internet devices are divided into importance levels, the protection measures are divided into different protection levels according to the protection requirements corresponding to the importance levels, a protection decision tree is constructed, and decision rules based on node centrality and propagation delay are set at the branch nodes of the protection decision tree; Selecting nodes to be protected at each protection level according to the decision rule, calculating the alarm processing efficiency of the selected nodes, calculating the protection priority based on the position of the nodes in the attack scenario graph, and matching the protection priority with the protection level to generate a candidate protection node combination; Calculate the alarm blocking effect of each candidate protection node combination, select the protection node combination with the best blocking effect and meeting the resource constraints, and generate protection instructions based on the protection node combination.

[0012] A second aspect of an embodiment of the present invention provides an industrial Internet security situation analysis system based on support vector regression, comprising: The first unit is used to collect alarm data from industrial Internet devices. Based on the time series properties of the alarm data and the network topology between industrial Internet devices, it uses a time window sliding method to perform correlation analysis on the alarm data and identify homologous alarm events with causal relationships. The second unit is used to establish a time-series propagation matrix based on homologous alarm events, calculate the probability of alarm propagation between industrial Internet devices, construct diffusion paths, and build an attack scenario graph based on the overlap and coverage of multiple diffusion paths. In the attack scenario graph, the propagation characteristics of the alarm data and the number of affected devices are extracted to generate a situation assessment dataset. The third unit is used to input the situation assessment dataset into the pre-trained support vector regression model, calculate the security situation index, determine the protection priority, and formulate the protection strategy based on the diffusion path in the attack scenario graph; The fourth unit is used to calculate the node centrality of each industrial Internet device based on the attack scenario graph, build a protection decision tree in combination with the protection strategy, select the optimal protection node combination based on the node centrality and the protection decision tree, generate protection instructions, and send the protection instructions to the affected industrial Internet devices to implement active protection.

[0013] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0014] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0015] In this embodiment, a time window sliding method is used to perform correlation analysis on the alarm data, identify homologous alarm events with causal relationships, establish a time series propagation matrix to calculate the alarm propagation probability, construct an attack scenario graph, extract the alarm propagation characteristics and the number of affected devices, and generate a situation assessment data set, effectively realizing accurate analysis and evaluation of the security situation of the industrial Internet. The situation assessment data set is input into the trained support vector regression model, the security situation index is calculated, the protection priority is determined, and the protection strategy is formulated in combination with the diffusion path in the attack scenario graph, realizing data-driven situation assessment and protection decision-making, and improving the accuracy of security situation assessment and the effectiveness of protection decision-making. Based on the attack scenario graph, the node centrality is calculated, and the protection decision tree is constructed in combination with the protection strategy. The optimal protection node combination is selected, and the protection instructions are generated and issued to realize active protection. Compared with the traditional passive protection method, it can more accurately identify key nodes, optimize the allocation of protection resources, and improve the protection efficiency and overall security. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 Schematic diagram of the process of industrial Internet security situation analysis method based on support vector regression according to an embodiment of the present invention; Figure 2 This is a comparison diagram of the effective alarm propagation relationship before and after the alarm propagation rule preprocessing according to an embodiment of the present invention; Figure 3 Construct a performance heat map for the attack scenario graph of an embodiment of the present invention. DETAILED DESCRIPTION

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0018] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0019] Figure 1 FIG. 1 is a flow chart of an industrial Internet security situation analysis method based on support vector regression according to an embodiment of the present invention. Figure 1 As shown, the method includes: Collect alarm data from industrial Internet devices and, based on the time series properties of the alarm data and the network topology between industrial Internet devices, perform correlation analysis on the alarm data using a time window sliding method to identify homologous alarm events with causal relationships. Based on homologous alarm events, a time-series propagation matrix is established to calculate the alarm propagation probability between industrial Internet devices and construct a diffusion path. An attack scenario graph is constructed based on the overlap and coverage of multiple diffusion paths. The propagation characteristics of the alarm data and the number of affected devices are extracted from the attack scenario graph to generate a situation assessment dataset. Input the situation assessment dataset into the pre-trained support vector regression model to calculate the security situation index, determine the protection priority, and formulate a protection strategy based on the diffusion path in the attack scenario graph; Based on the attack scenario graph, the node centrality of each industrial Internet device is calculated, and a protection decision tree is constructed in combination with the protection strategy. The optimal protection node combination is selected based on the node centrality and the protection decision tree, and protection instructions are generated. The protection instructions are sent to the affected industrial Internet devices to implement active protection.

[0020] In an optional embodiment, alarm data from industrial Internet devices is collected, and based on the time series attributes of the alarm data and the network topology relationship between the industrial Internet devices, a time window sliding method is used to perform correlation analysis on the alarm data to identify homologous alarm events with causal relationships, including: Collect alarm data from industrial Internet devices, set a corresponding reference time window according to the alarm level of the alarm data, and construct a network connection diagram based on the network topology relationship between industrial Internet devices, the network connection diagram including device nodes and communication links; An attention-based density clustering algorithm is used to calculate the alarm density of the alarm data, and a reference time window is adaptively adjusted according to the alarm density and the number of shortest path hops between device nodes in a network connection graph to obtain a time window. Within the time window, calculating the time series similarity based on the time series attributes of the alarm data, calculating the path similarity based on the shortest path between the device nodes in the network connection graph, and calculating the feature similarity based on the attribute features of the alarm data; Calculating a propagation weight coefficient of the alarm data, performing a weighted combination of the time series similarity, path similarity, and feature similarity according to the propagation weight coefficient to obtain an alarm propagation probability, and constructing an alarm propagation graph, wherein nodes are alarm data and edges are alarm propagation relationships with alarm propagation probabilities; Based on the recursive search algorithm, the alarm propagation graph is dynamically verified for paths, the alarm propagation relationships that do not meet the propagation timing constraints are identified and eliminated, and multiple connected subgraphs are obtained. The connected subgraphs are then verified for timing consistency to identify homologous alarm events with causal relationships.

[0021] This embodiment provides a method for identifying homologous alarm events based on timing attributes and network topology relationships. The method first collects alarm data from industrial Internet devices. Each piece of alarm data contains attributes such as device identification, alarm time, alarm level, alarm type, and alarm description. A corresponding reference time window is set according to the alarm level. For example, a reference time window of 300 seconds is set for emergency-level alarms, a reference time window of 180 seconds is set for important-level alarms, and a reference time window of 120 seconds is set for minor-level alarms. At the same time, a network connection graph is constructed based on the connection relationship between industrial Internet devices, where nodes represent devices and edges represent communication links between devices. For example, on an industrial production line, the communication connection relationship between PLC controllers, sensors, and actuators is represented as an edge in the network connection graph.

[0022] After constructing the network connection graph, a density clustering algorithm based on the attention mechanism is used to calculate the alarm density of the alarm data. Specifically, for each time point t, the number of alarms within a certain time range before and after (e.g., 60 seconds before and after) is counted, and alarms closer to time point t are given a higher weight. For example, if there are 20 alarms within 60 seconds before and after time point t, and 10 of them appear within 10 seconds before and after t, the alarm density at time point t will be higher. The calculated alarm density reflects the degree of abnormality of the system at a specific time point.

[0023] The benchmark time window is adaptively adjusted based on the calculated alarm density and the number of hops in the shortest path between device nodes in the network connection diagram. When the alarm density is high, it indicates that system anomalies are more concentrated, and the time window can be appropriately narrowed; when the alarm density is low, the time window is expanded to capture more potentially related alarms. At the same time, the number of hops in the shortest path between devices is taken into account. For alarm correlations between devices with a large number of hops, the time window is appropriately expanded to account for alarm propagation delays. For example, for a critical-level alarm with a benchmark time window of 180 seconds, if the alarm density is 0.8 (high density) and the shortest path between devices is 1 hop, the time window is adjusted to 150 seconds; if the alarm density is 0.3 (low density) and the shortest path between devices is 3 hops, the time window is adjusted to 220 seconds.

[0024] After determining the time window, three types of similarity are calculated for the alarm data within the window. Temporal similarity measures the proximity of alarm occurrence times. The difference between two alarm timestamps is calculated and normalized. For example, if the time difference between two alarms is 30 seconds and the time window is 180 seconds, the temporal similarity is 1-30 / 180 = 0.833. Path similarity is calculated based on the shortest path between device nodes in the network connection graph, using the inverse of the shortest path length as the path similarity. For example, if the shortest path between device A and device B is 2 hops, the path similarity between them is 1 / 2 = 0.5. Feature similarity is calculated based on the attribute characteristics of the alarm data, comparing the similarity of fields such as the alarm type and alarm description. For example, using a text similarity algorithm to calculate the similarity of two alarm descriptions, the feature similarity of "network connection interrupted" and "network connection timed out" might be 0.8.

[0025] The propagation weight coefficient for alarm data is calculated based on the alarm data's impact range, alarm severity, and historical statistical data. Alarms with wider impact ranges and higher severity levels receive greater propagation weights. For example, an alarm affecting the entire network subsystem might have a propagation weight coefficient of 0.9, while an alarm affecting only a single device might have a propagation weight coefficient of 0.3. These propagation weight coefficients are used to weight the timing similarity, path similarity, and feature similarity to obtain the alarm propagation probability. For example, for two alarms A and B, assuming the timing similarity is 0.8, the path similarity is 0.6, and the feature similarity is 0.7, and the propagation weight coefficient for alarm A is 0.8, then the alarm propagation probability might be 0.8 × 0.8 + 0.6 × 0.1 + 0.7 × 0.1 = 0.77. Based on the calculated alarm propagation probabilities, an alarm propagation graph is constructed, where nodes represent alarm data and edges represent alarm propagation relationships with alarm propagation probabilities.

[0026] Dynamic path verification is performed on the constructed alarm propagation graph, using a recursive search algorithm to identify and remove alarm propagation relationships that do not meet propagation timing constraints. Propagation timing constraints require that the cause alarm in a causal relationship must occur before the effect alarm. For example, if an edge exists between alarms A and B, but alarm B occurs earlier than alarm A, the edge violates the propagation timing constraint and should be removed. Recursive search checks all paths in the alarm propagation graph and removes edges that violate timing constraints.

[0027] After dynamic path verification, the alarm propagation graph is divided into multiple connected subgraphs, each of which contains a set of alarms that may have a causal relationship. These connected subgraphs are subjected to temporal consistency verification to check whether there are loops or timing contradictions in the subgraphs. Specifically, in each connected subgraph, the alarms are sorted according to the time sequence of their occurrence to check whether the temporal relationship of cause and effect is met. For example, in a connected subgraph containing alarms A, B, and C, if their occurrence time sequence is A→B→C, and there are edges A→B, B→C, and A→C, then the subgraph has temporal consistency; conversely, if the temporal relationship between the alarms is inconsistent with the direction of the edge, then the subgraph does not meet temporal consistency. Through temporal consistency verification, homologous alarm events with causal relationships are eventually identified. These events can help operation and maintenance personnel quickly locate the root cause of the fault and improve the operation and maintenance efficiency of the industrial Internet system.

[0028] In this embodiment, by introducing adaptive time windows driven by alarm levels and device network topology information, the timeliness and accuracy of identifying homologous alarm events are improved; the density clustering algorithm based on the attention mechanism is used to enhance the perception of local alarm-intensive areas; the integration of timing, path and feature similarity and the introduction of propagation weights for weighted calculation effectively improve the rationality and credibility of alarm propagation relationship modeling; combined with recursive search and timing consistency verification, it is possible to eliminate pseudo-correlated alarm paths, accurately extract the true causal propagation chain, and realize intelligent identification of homologous events in complex alarm scenarios, supporting operation and maintenance decision-making and security situation awareness.

[0029] In an optional embodiment, dynamic path verification is performed on the alarm propagation graph based on a recursive search algorithm to identify and eliminate alarm propagation relationships that do not meet propagation timing constraints, thereby obtaining multiple connected subgraphs, performing temporal consistency verification on the connected subgraphs, and identifying homologous alarm events with causal relationships, including: Calculating a propagation delay attenuation coefficient according to the number of network hops between the alarm data, and determining a propagation timing constraint between adjacent alarm data based on the propagation delay attenuation coefficient; Constructing an alarm propagation rule based on the alarm type of the alarm data, preprocessing the alarm propagation graph according to the alarm propagation rule, deleting the alarm propagation relationship that does not meet the alarm propagation rule, and obtaining a preprocessed alarm propagation graph; Perform a depth-first search on the source alarm data in the pre-processed alarm propagation graph. During the search, determine whether the time intervals between adjacent alarm data satisfy the propagation timing constraints, and mark the alarm data that violates the propagation timing constraints as invalid alarm data. Skipping the path search starting from the invalid alarm data, updating the reachability mark of the alarm data, and saving the alarm propagation relationship that meets the propagation timing constraint as a valid alarm propagation relationship; According to the connectivity of the effective alarm propagation relationship, the preprocessed alarm propagation graph is decomposed into multiple connected sub-graphs, and the propagation probability of the alarm propagation relationship in the connected sub-graph and the timing consistency index of the alarm timestamp are calculated. Based on the timing consistency index, the connected sub-graph is verified for timing consistency to identify homologous alarm events with causal relationships.

[0030] In this embodiment, the propagation delay attenuation coefficient is first determined by calculating the number of network hops between alarm data. The number of network hops refers to the number of network devices passed from one alarm node to another. For example, for alarm nodes that are 1 hop apart, the propagation delay attenuation coefficient is set to 0.8; for alarm nodes that are 2 hops apart, the propagation delay attenuation coefficient is set to 0.6; and for alarm nodes that are 3 hops apart, the propagation delay attenuation coefficient is set to 0.4. Based on these propagation delay attenuation coefficients, the propagation timing constraints between adjacent alarm data are further determined. The propagation timing constraints are expressed as a time threshold. For example, for alarm nodes that are 1 hop apart, the time threshold is set to 5 seconds; for alarm nodes that are 2 hops apart, the time threshold is set to 10 seconds; and for alarm nodes that are 3 hops apart, the time threshold is set to 15 seconds.

[0031] Alarm propagation rules are constructed based on the alarm types of the alarm data. Alarm propagation rules describe the causal relationships that may exist between different types of alarms. For example, an alarm of type "network link interruption" may cause an alarm of type "service unavailable," but an alarm of type "service unavailable" is unlikely to cause an alarm of type "network link interruption." Based on these rules, the alarm propagation graph is preprocessed, removing alarm propagation relationships that do not meet the alarm propagation rules. This results in a preprocessed alarm propagation graph. For example, in the original alarm propagation graph, if there is an edge from an alarm of type "service unavailable" to an alarm of type "network link interruption," this edge will be deleted.

[0032] A depth-first search is performed on the preprocessed alarm propagation graph for source alarm data. Source alarm data refers to alarm nodes without incoming edges, meaning that they are not caused by other alarms. During the search, the time intervals between adjacent alarm data are determined to meet the propagation timing constraints. For example, if alarm A occurs at time t1 and alarm B occurs at time t2, if t2 - t1 is less than the time threshold determined by the number of network hops, the propagation relationship between alarms A and B is considered to meet the timing constraints. Otherwise, alarm B is marked as invalid alarm data.

[0033] During the search, the path search starting from invalid alarm data is skipped to improve the search efficiency. At the same time, the reachability mark of the alarm data is updated, and the alarm propagation relationship that meets the propagation timing constraint is saved as a valid alarm propagation relationship. The reachability mark is used to indicate whether there is a valid propagation path from the source alarm data to the current alarm data. For example, for a path in the alarm propagation graph: alarm A → alarm B → alarm C, if the propagation relationship from alarm A to alarm B meets the timing constraint, and the propagation relationship from alarm B to alarm C does not meet the timing constraint, then alarm C will be marked as invalid alarm data, and all paths starting from alarm C will not be searched further.

[0034] Based on the connectivity of valid alarm propagation relationships, the preprocessed alarm propagation graph is decomposed into multiple connected subgraphs. A connected subgraph is one in which all nodes are mutually reachable. For each connected subgraph, the propagation probability of the alarm propagation relationship and the temporal consistency index of the alarm timestamp are calculated. The propagation probability is the probability, based on historical data statistics, that one alarm type will cause another alarm type. For example, the probability of "network link interruption" causing "service unavailability" may be 0.85. The temporal consistency index is a comprehensive metric calculated based on the time intervals of all alarm propagation relationships in the connected subgraph. It is used to measure the rationality of the temporal logic of alarm propagation in the connected subgraph.

[0035] Based on the temporal consistency index, the connected subgraph is verified for temporal consistency, thereby identifying homologous alarm events with causal relationships. Specifically, if the temporal consistency index of a connected subgraph exceeds a preset threshold (for example, 0.75), the alarms in the connected subgraph are considered to be caused by the same root event, that is, they constitute a homologous alarm event. For example, a connected subgraph contains alarms A, B, C, and D, where alarm A is the source alarm. If the temporal consistency index of the connected subgraph is 0.85, which exceeds the preset threshold of 0.75, then alarms A, B, C, and D are considered to be caused by the same root event and constitute a homologous alarm event.

[0036] In this embodiment, it is possible to effectively eliminate abnormal paths in the alarm propagation chain that do not conform to the temporal logic or propagation rules, thereby improving the accuracy and reliability of causal relationship identification. By introducing the propagation delay attenuation coefficient, the time propagation constraints under different network hops are reasonably limited to avoid misjudging long-distance nodes as homologous events; propagation rules are constructed in combination with alarm types to effectively filter out logically invalid alarm relationships; a depth-first search is used in combination with propagation timing judgment and reachability marking strategies to improve the efficiency and accuracy of propagation path verification; and finally, by jointly evaluating the propagation probability and temporal consistency of the connected subgraph, the precise identification and tracing of homologous alarm events is achieved, enhancing the intelligence and practicality of alarm processing.

[0037] Figure 2 This is a comparison diagram of the effective alarm propagation relationship before and after the alarm propagation rule preprocessing according to the embodiment of the present invention. Figure 2 The figure shows a comparative analysis of the number of valid alarm propagation relationships before and after alarm propagation rule preprocessing, and also compares the effectiveness of this technical solution with a graph neural network-based GNN alarm association method. As can be seen from the figure, the original alarm propagation relationships before preprocessing were relatively large, such as 85 relationships for the network link interruption type and 95 relationships for the security attack type. After alarm propagation rule preprocessing using this technical solution, the number of valid propagation relationships for each alarm type was significantly reduced: network link interruption was reduced to 64 (a 25% reduction), service unavailability was reduced to 50 (a 36% reduction), CPU overload was reduced to 57 (a 38% reduction), memory overflow was reduced to 43 (a 39% reduction), and security attack was reduced to 60 (a 37% reduction). In comparison, the GNN-based alarm association method performed slightly worse, with the corresponding valid relationship numbers being 49, 39, 46, 32, and 42, respectively. This demonstrates that this technical solution, through its alarm propagation rules based on alarm type, can more accurately identify and retain causal alarm propagation paths while effectively eliminating illogical propagation relationships. For example, "service unavailability" is unlikely to cause "network link interruption." This preprocessing mechanism significantly improves the accuracy and efficiency of subsequent alarm correlation analysis.

[0038] In an optional implementation, a time series propagation matrix is established based on homologous alarm events, the alarm propagation probability between industrial Internet devices is calculated, and the diffusion path is constructed, including: Construct an initial time series propagation matrix based on homologous alarm events. Use the industrial Internet devices in the homologous alarm events as row vectors and column vectors. Count the number of alarm propagations between each pair of devices within a specified time window. For any two industrial Internet devices, calculate the alarm propagation probability based on the ratio of the number of target device alarms generated after the source device generates an alarm to the total number of source device alarms. Fill the alarm propagation probability into the corresponding matrix element position to generate a time series propagation matrix. The alarm propagation probability less than the preset probability threshold is set to zero to obtain a filtered time series propagation matrix. Based on the filtered time series propagation matrix, a depth-first search method is used to search for the alarm propagation path starting from the first alarm device. The alarm propagation probability on each path is multiplied as the path weight, and a weight threshold is set for the path weight. The paths greater than the weight threshold are retained as diffusion paths.

[0039] In this implementation, alarm event data from the Industrial Internet environment is first collected. Each alarm event data entry contains information such as the device ID, alarm type, and alarm occurrence time. To construct the initial time-series propagation matrix, the collected alarm events need to be preprocessed and filtered to identify alarm events with the same source. These events are referred to as a series of device alarms caused by the same cause.

[0040] In a specific embodiment, when constructing an initial time-series propagation matrix based on homologous alarm events, the Industrial Internet devices involved in the homologous alarm events need to be used as the matrix's row and column vectors. For example, in an Industrial Internet environment, there are five key devices: Device A, Device B, Device C, Device D, and Device E. These devices generate multiple alarms within a certain period of time, and the propagation relationships between these alarms need to be analyzed. To this end, an initial 5×5 matrix is constructed, with the rows and columns representing the five devices, respectively.

[0041] When counting the number of alarm propagation events between each pair of devices within a specified time window, set the time window to 10 minutes. Within this time window, if device A generates an alarm first and then device B generates an alarm within the same time window, alarm propagation from device A to device B is considered possible. Analysis of historical alarm logs reveals that within this 10-minute window, device B generated an alarm 15 times after device A generated an alarm, for a total of 20 alarms generated by device A.

[0042] When calculating the alarm propagation probability from device A to device B, the number of times device B generates an alarm within the specified time window after device A generates an alarm (15 times) divided by the total number of alarms generated by device A (20 times) yields a probability of 0.75. Similarly, the alarm propagation probability between all device pairs is calculated. For example, the alarm propagation probability from device A to device C is 0.6, the alarm propagation probability from device B to device D is 0.8, the alarm propagation probability from device C to device E is 0.7, the alarm propagation probability from device B to device C is 0.3, and the alarm propagation probability from device D to device E is 0.5.

[0043] These calculated alarm propagation probabilities are entered into the corresponding matrix element positions to generate a time series propagation matrix. In this matrix, rows represent source devices, columns represent destination devices, and matrix element values represent the alarm propagation probabilities between corresponding device pairs. For example, the value of 0.75 in the first row and second column of the matrix indicates that the alarm propagation probability from device A to device B is 0.75. To filter out low-probability alarm propagation relationships, a preset probability threshold of 0.4 is set. Alarm propagation probabilities less than 0.4 in the time series propagation matrix are set to 0, resulting in a filtered time series propagation matrix. In this example, the alarm propagation probability of 0.3 from device B to device C is less than the threshold of 0.4, so this value is set to 0 in the filtered matrix. This way, only strong alarm propagation relationships are retained in the filtered time series propagation matrix.

[0044] Based on the filtered time-series propagation matrix, a depth-first search method is used to search for alarm propagation paths. Assuming the first alarm is device A, the search begins with device A as the starting point for possible alarm propagation paths. The depth-first search process follows a path as deeply as possible until it reaches a point where it cannot proceed further, then backtracks and explores other possible paths. During the search, starting from device A, it is possible to reach devices B (with a propagation probability of 0.75) and C (with a propagation probability of 0.6). From device B, it is possible to reach device D (with a propagation probability of 0.8), from device C, it is possible to reach device E (with a propagation probability of 0.7), and from device D, it is possible to reach device E (with a propagation probability of 0.5). This depth-first search method identifies multiple possible alarm propagation paths, including "A→B→D→E" and "A→C→E." To calculate the weight of each path, the propagation probabilities along the path are multiplied to form the path weight. For example, the weight of the path "A→B→D→E" is 0.75×0.8×0.5=0.3, and the weight of the path "A→C→E" is 0.6×0.7=0.42.

[0045] Set the path weight threshold to 0.35 and retain paths with weights greater than the threshold as the final diffusion paths. In this example, the path "A→B→D→E" has a weight of 0.3, which is less than the threshold of 0.35, and is therefore filtered out. However, the path "A→C→E" has a weight of 0.42, which is greater than the threshold of 0.35, and is therefore retained as a valid diffusion path.

[0046] To improve the accuracy of the alarm propagation matrix, you can use a time decay factor to weight historical alarm data. Newer alarm data is given a higher weight, while older alarm data is given a lower weight. For example, if the time decay factor is set to 0.9, the weight is multiplied by 0.9 for each time unit (such as a day) forward. This way, alarm data from a week ago will be weighted approximately 0.5 times as much as current alarm data, more accurately reflecting the alarm propagation patterns of the current network status.

[0047] In practical applications, dynamically adjusted probability and weight thresholds can be set. Thresholds are automatically adjusted based on the scale and complexity of the Industrial Internet environment. For example, in an environment with a large number of devices and frequent alarms, the threshold can be appropriately increased to reduce the number of paths and focus on the most important alarm propagation paths. In an environment with a small number of devices and fewer alarms, the threshold can be appropriately lowered to capture more possible alarm propagation relationships.

[0048] To address the sparsity of the time-series propagation matrix, matrix compression storage technology can be used. In large-scale industrial internet environments, the number of devices may reach thousands, making the time-series propagation matrix extremely large. Since most devices lack direct alarm propagation relationships, the majority of elements in the matrix are zero. By storing only non-zero elements and their location information, storage requirements can be significantly reduced, improving algorithm efficiency. To address false positives and false negatives, a confidence scoring mechanism can be introduced. For each alarm, a confidence score is calculated based on factors such as its source, type, and historical accuracy. When constructing the time-series propagation matrix, the alarm confidence is taken into account to reduce the impact of low-confidence alarms on the propagation probability calculation. For example, the alarm confidence can be used as a weighting factor and multiplied by the original count value. When constructing the diffusion path, the physical connection between devices can be considered. If two devices do not have a direct or indirect connection in the physical network, even if statistically indicated, an alarm propagation relationship between them may be a coincidence caused by other factors. By incorporating network topology information, such false propagation relationships can be eliminated, improving the accuracy of the diffusion path.

[0049] In this embodiment, by constructing an initial time-series propagation matrix and filtering low-probability propagation relationships, the system can focus on important alarm links and reduce analysis noise. The depth-first search combined with the path weight calculation method enables the system to find the most likely alarm diffusion path and improve the efficiency of security incident tracing. The introduction of the time decay factor and confidence scoring mechanism enhances the adaptability of the model to real-time network status and reduces the impact of false alarms. Combined with the analysis of the physical connection relationship of the equipment, false propagation relationships can be eliminated and the accuracy of the diffusion path can be improved. The dynamic threshold adjustment mechanism enables the system to adapt to industrial network environments of different scales and complexities. Matrix compression storage technology effectively solves the problem of computing efficiency in large-scale environments. The risk assessment function supports the security team to prioritize high-risk paths, improve security situation awareness and response efficiency, reduce the scope of security incidents, and ensure the stable operation of the industrial Internet system.

[0050] In an optional embodiment, an attack scenario graph is constructed based on the overlap and coverage of multiple diffusion paths, and the propagation characteristics and number of affected devices of the alarm data are extracted from the attack scenario graph to generate a situation assessment dataset, including: Constructing a path connection relationship table based on the diffusion path, wherein the path connection relationship table records the device node sequence passed by each diffusion path, and counting the number of repeated device nodes in the path connection relationship table to obtain the path overlap degree; The device node sequence of each diffusion path is traversed to obtain the total number of nodes accessible by the path, and the ratio of the total number of accessible nodes to the total number of network nodes is used as the path coverage; Dividing diffusion paths with the same repeated nodes into the same group according to the degree of path overlap, calculating the coverage of the paths in each group, and selecting the path with the largest coverage to construct an attack scenario graph; Identify the propagation feature nodes in the attack scenario graph, extract the propagation features based on the connection relationship between the propagation feature nodes, count the number of device nodes reachable by the propagation feature nodes as the number of affected devices, and generate a situation assessment data set.

[0051] In this embodiment, when constructing a path connection relationship table based on the diffusion path, it is necessary to record the sequence of device nodes passed by each diffusion path obtained in the above steps. In actual applications, it is assumed that multiple diffusion paths have been obtained through the above method, such as path 1: device A→device B→device D→device F; path 2: device A→device C→device E→device G; path 3: device A→device B→device E→device H; path 4: device B→device D→device F→device I. For these paths, a path connection relationship table is constructed, which contains the path ID and the corresponding device node sequence. For example, the record of path 1 is (1, [A, B, D, F]), the record of path 2 is (2, [A, C, E, G]), the record of path 3 is (3, [A, B, E, H]), and the record of path 4 is (4, [B, D, F, I]).

[0052] To determine the degree of path overlap by counting the number of repeated device nodes in the path connectivity table, it's necessary to traverse each path in the path connectivity table and calculate the number of times each device node appears in different paths. By comparing the shared device nodes across different paths, the degree of path overlap can be determined. In the above example, device A appears in paths 1, 2, and 3, with a repetition count of 3; device B appears in paths 1, 3, and 4, with a repetition count of 3; device D appears in paths 1 and 4, with a repetition count of 2; device E appears in paths 2 and 3, with a repetition count of 2; device F appears in paths 1 and 4, with a repetition count of 2; and devices C, G, H, and I each appear in only one path, with a repetition count of 1.

[0053] When traversing the device node sequence of each diffusion path to obtain the total number of nodes accessible by the path, it is necessary to calculate the number of device nodes that each path can directly or indirectly access. In practical applications, a graph traversal algorithm can be used to explore all reachable nodes, starting from each node in the path. Assume that there are 15 device nodes in an industrial Internet environment. After traversal analysis, path 1 can access 6 nodes (including 4 nodes on the path and 2 other nodes reachable from these nodes); path 2 can access 5 nodes; path 3 can access 7 nodes; and path 4 can access 8 nodes.

[0054] When using the ratio of the total number of accessible nodes to the total number of network nodes as the path coverage, it's necessary to calculate the proportion of the network that each path covers. In the above example, with a total number of 15 network nodes, the coverage of path 1 is 6 / 15 = 0.4; the coverage of path 2 is 5 / 15 = 0.33; the coverage of path 3 is 7 / 15 = 0.47; and the coverage of path 4 is 8 / 15 = 0.53. A larger coverage indicates a wider network impact, and thus a higher potential security risk.

[0055] When grouping diffusion paths with the same repeated nodes into the same group based on the degree of path overlap, it is necessary to identify similarities between the paths. This is achieved by grouping them based on the key nodes that recur within the paths. In the above example, paths 1 and 3 can be grouped together because they both contain devices A and B; paths 1 and 4 can be grouped together because they both contain devices B, D, and F; and paths 2 and 3 can be grouped together because they both contain devices A and E. Note that a path can belong to multiple groups, depending on its overlap with other paths.

[0056] When calculating the coverage of the paths in each group and selecting the path with the largest coverage to construct the attack scenario graph, it is necessary to compare the coverage of the paths within each group and select the path with the largest coverage as the representative of that group. In the above example, for the group containing paths 1 and 3, path 3's coverage (0.47) is greater than path 1's (0.4), so path 3 is selected as the representative of this group. For the group containing paths 1 and 4, path 4's coverage (0.53) is greater than path 1's (0.4), so path 4 is selected as the representative of this group. For the group containing paths 2 and 3, path 3's coverage (0.47) is greater than path 2's (0.33), so path 3 is selected as the representative of this group. Ultimately, paths 3 and 4 are selected to construct the attack scenario graph.

[0057] When constructing an attack scenario graph, the selected paths are combined to form a directed graph. In this graph, nodes represent devices, and edges represent alert propagation relationships. In the above example, path 3 (A→B→E→H) and path 4 (B→D→F→I) are combined to form an attack scenario graph that includes devices A, B, D, E, F, H, and I. This graph reflects the potential attack paths and the range of affected devices.

[0058] To identify propagation signature nodes in an attack scenario graph, it's necessary to analyze the nodes' positions and connections within the graph to identify key nodes. Propagation signature nodes typically have high connectivity or are located at key positions along the path. In the example above, device B is a propagation signature node because it connects two distinct paths and serves as a hub for the attack's spread. Other propagation signature nodes may include those with high in-degree or high out-degree, such as device A (which serves as the starting point for multiple paths) and device D (which connects to multiple downstream nodes).

[0059] When extracting propagation features based on the connections between nodes, it's necessary to analyze the link patterns between these nodes. Propagation features can include star structures (one node connecting to multiple other nodes), chain structures (nodes connecting sequentially), and ring structures (forming a closed loop). In the example above, a star-shaped propagation structure centered on device B can be identified, which connects to devices A, D, and E, forming a diverging propagation pattern. Another characteristic is a chain-shaped propagation structure from device D to device F and then to device I. These propagation features reflect the attack's spread pattern within the network.

[0060] To calculate the number of device nodes reachable by a propagation feature node as the number of affected devices, you need to calculate the total number of device nodes reachable from the propagation feature node. In the example above, starting from device B, you can reach devices D, E, F, H, and I, for a total of five device nodes. Starting from device A, you can reach device B and all downstream nodes, for a total of six device nodes. These data reflect the impact range of different propagation feature nodes.

[0061] When generating a situation assessment dataset, organize the data obtained in the previous steps into a structured dataset. This dataset includes information such as path overlap, path coverage, the topology of the attack scenario graph, propagation signature nodes and their characteristics, and the number of affected devices. For example, a dataset can be created containing the following fields: path ID, path node sequence, overlapping node list, coverage, attack scenario graph ID, propagation signature type, propagation signature node list, and number of affected devices. In the above example, a record in the dataset might be: {Path ID: 3, Path node sequence: [A, B, E, H], Overlapping node list: [A, B, E], Coverage: 0.47, Attack scenario graph ID: 1, Propagation signature type: "Star", Propagation signature node list: [B], Number of affected devices: 5}.

[0062] To improve the accuracy of the attack scenario graph, a node weighting mechanism can be introduced. Each device node is assigned a weight based on factors such as its importance in the network, the severity of its alerts, and its historical attack frequency. When constructing the attack scenario graph, paths containing high-weighted nodes are prioritized. For example, if device B is a core production control device and has a high weight, paths containing device B will be prioritized when constructing the attack scenario graph.

[0063] In practice, dynamic coverage thresholds can be set to automatically adjust based on network size and security requirements. For example, in large-scale industrial Internet environments, a lower coverage threshold (such as 0.2) can be set to capture more potential attack scenarios; in small networks, a higher threshold (such as 0.5) can be set to focus on the most important attack scenarios.

[0064] In this embodiment, by constructing a path connectivity table and calculating the degree of path overlap, the solution can effectively identify key propagation nodes and high-risk propagation paths, avoiding the one-sided analysis of isolated alarms in traditional methods. The introduction of a path coverage metric enables the system to quantitatively assess the impact range of potential attacks, prioritizing attack paths with high coverage. Grouping diffusion paths with the same repeated nodes into the same group reduces redundant analysis and improves computational efficiency. The attack scenario graph constructed based on the maximum coverage path comprehensively displays the propagation pattern and potential impact of alarms in the network, allowing security personnel to intuitively understand the attack chain. By identifying propagation feature nodes and extracting propagation features, the system can reveal typical attack patterns and behavioral characteristics, improving the accuracy of anomaly detection. Statistics on the number of affected devices provide a quantitative basis for determining security response priorities. The generated situation assessment dataset comprehensively reflects the security status of the Industrial Internet, providing high-quality training data for subsequent security situation prediction based on support vector regression, and enhancing the foresight and targeted nature of security protection.

[0065] Figure 3 Construct a performance heat map for the attack scenario graph of the embodiment of the present invention, such as Figure 3The figure shows the attack scenario graph construction performance of this technical solution under different network scales and path overlap conditions. It can be observed that when the network number is 100 nodes and the path overlap is 10%, the attack scenario graph construction takes only 8ms, demonstrating extremely high processing efficiency. Even in the most complex scenario (10,000 nodes and 60% overlap), the processing time is only 261ms. The data shows that processing time increases nonlinearly with the number of network nodes and path overlap, but the growth curve is relatively flat. Particularly noteworthy is that in a medium-sized network (1,000 nodes), even with a path overlap of 40%, the construction time is only 58ms. This is significantly superior to traditional graph traversal algorithms (such as the Kruskal minimum spanning tree algorithm), which typically require processing times of over 200ms. Furthermore, when the number of nodes increases to 5,000, even with a higher overlap (50%), this technical solution completes the construction process within 143ms, which is crucial for real-time situational awareness systems. The heat map clearly demonstrates the advantages of this technical solution based on overlapping path selection, proves its efficiency and scalability in large-scale network environments, and provides a solid performance foundation for network security situation assessment.

[0066] In an optional embodiment, the situation assessment dataset is input into a pre-trained support vector regression model to calculate the security situation index, determine the protection priority, and formulate a protection strategy based on the diffusion path in the attack scenario graph, including: Extract the alarm propagation feature sequence and the affected device number sequence from the situation assessment dataset, calculate the standardized feature based on the maximum and minimum values of the alarm propagation feature sequence, and calculate the normalized feature based on the ratio of the affected device number sequence to the total number of devices; The standardized features and the normalized features are combined to construct a feature matrix, the mutual information values between the features in the feature matrix are calculated to obtain feature redundancy, features are selected based on the feature redundancy and information gain is calculated, and training features are obtained when the information gain is less than a preset gain threshold; Dividing the training features into a training set and a validation set and performing cross-validation to obtain optimal kernel function parameters, inputting the training features and kernel function parameters into a support vector regression model, calculating an error value between a predicted result and a true result, updating the model parameters according to the error value, and obtaining a trained support vector regression model when the error value is less than a preset error threshold; Inputting the situation assessment data set into the trained support vector regression model to obtain a security situation index, and performing weighted smoothing processing on the security situation index within a preset time window to obtain a smoothed situation index; The difference in smoothed situation index between adjacent time points is calculated to obtain the situation change rate. The protection priority is determined based on the situation change rate. The node weight is calculated based on the time sequence of the nodes visited on the diffusion path in the attack scenario graph. The protection priority is adjusted to obtain a protection node sequence. A protection strategy is formulated based on the protection node sequence and the node access timing of the diffusion path.

[0067] In this embodiment, when extracting the alarm propagation feature sequence and the affected device number sequence from the situation assessment dataset, it is necessary to analyze the structured data in the situation assessment dataset. In practical applications, assume that the situation assessment dataset contains data from the past 30 days, where the alarm propagation feature sequence records the frequency of occurrence of different propagation features (such as star, chain, and ring), and the affected device number sequence records the number of devices affected by each security incident. For example, on day 1, the star propagation feature appeared 5 times, the chain propagation feature appeared 3 times, and the ring propagation feature appeared 1 time, affecting a total of 15 devices; on day 2, the star propagation feature appeared 4 times, the chain propagation feature appeared 5 times, and the ring propagation feature appeared 2 times, affecting a total of 18 devices, and so on.

[0068] When calculating standardized features based on the maximum and minimum values of the alarm propagation feature sequence, it is necessary to standardize the values of different propagation features to the same scale. In the above example, assume that within 30 days, the maximum value of the star propagation feature is 8, and the minimum value is 2; the maximum value of the chain propagation feature is 7, and the minimum value is 1; and the maximum value of the ring propagation feature is 4, and the minimum value is 0. For the data on day 1, the standardized star feature value is (5-2) / (8-2)=0.5, the chain feature value is (3-1) / (7-1)=0.33, and the ring feature value is (1-0) / (4-0)=0.25.

[0069] When calculating normalized features based on the ratio of the number of affected devices to the total number of devices, it's necessary to calculate the relative size of the impact of each security incident. Assuming there are 100 devices in an Industrial Internet environment, the ratio of affected devices on day 1 is 15 / 100 = 0.15, and on day 2 it's 18 / 100 = 0.18. This normalization process makes data comparable across network environments of varying sizes. When combining standardized and normalized features to construct a feature matrix, the processed features must be organized into a structured matrix. In the above example, the feature vector for day 1 is [0.5, 0.33, 0.25, 0.15], the feature vector for day 2 is [0.33, 0.67, 0.5, 0.18], and so on. These feature vectors form the feature matrix, with each row representing a day's data and each column representing a feature.

[0070] When calculating the mutual information between features in the feature matrix to determine feature redundancy, it is necessary to analyze the correlation between different features. Mutual information measures the degree of interdependence between two variables. Higher mutual information values indicate more shared information between the two features and higher redundancy. In the above example, assume that the calculated mutual information between the star feature and the chain feature is 0.6, the mutual information between the star feature and the ring feature is 0.3, and the mutual information between the chain feature and the ring feature is 0.4. The mutual information between the star feature and the affected device ratio is 0.7, the mutual information between the chain feature and the affected device ratio is 0.5, and the mutual information between the ring feature and the affected device ratio is 0.2.

[0071] When selecting features based on feature redundancy and calculating information gain, redundant features must be removed, while features with high information content must be retained. Assuming a mutual information threshold of 0.5, the redundancy between star features and chain features, and between star features and the proportion of affected devices, is high, necessitating further calculation of the information gain for each feature. Assuming the calculated information gain for the star feature is 0.4, the information gain for the chain feature is 0.3, the information gain for the ring feature is 0.2, and the information gain for the proportion of affected devices is 0.5. When the information gain is less than the preset gain threshold, a training feature is obtained. Assuming the preset gain threshold is 0.25, the ring feature's information gain of 0.2 is less than the threshold and is therefore removed. The final selected training features include the star feature, chain feature, and the proportion of affected devices.

[0072] When dividing the training features into training and validation sets and performing cross-validation to determine the optimal kernel function parameters, a 5-fold cross-validation method can be used. 30 days of data are randomly divided into five parts, with four parts used as the training set and one part used as the validation set, for a total of five validation cycles. Suppose that linear, polynomial, and radial basis kernel functions are tried. After cross-validation, it is found that the radial basis kernel parameter gamma = 0.1 achieves optimal model performance. The training features and kernel function parameters are input into the support vector regression model. When calculating the error between the predicted and true results, the actual security situation index is set as the label. Assuming the security situation index for the past 30 days is known, the mean squared error is used as the evaluation metric. Model parameters are continuously adjusted during training. When the error reaches 0.05, which is less than the preset error threshold of 0.1, the trained support vector regression model is obtained.

[0073] When inputting the situation assessment dataset into the trained support vector regression model to obtain the security situation index, the most recent situation assessment data is used as input. Assuming the standardized features for the most recent day are [0.6, 0.4, 0.35], the model predicts a security situation index of 0.75. Exponential smoothing can be used to weightedly smooth the security situation index within a preset time window to obtain a smoothed situation index. Assuming a smoothing coefficient of 0.3 and a preset time window of 7 days, the smoothed situation index for the most recent day is 0.3 multiplied by 0.75 plus 0.7 multiplied by the smoothed situation index for the previous day. This continues in this order, yielding a sequence of smoothed situation indices within the time window.

[0074] To calculate the situation change rate by calculating the difference between the smoothed situation indices at adjacent time points, subtract the smoothed situation index at the previous time point from the current time point's smoothed situation index. A positive difference indicates a deterioration in the security situation; a negative difference indicates an improvement. For example, a calculated situation change rate of 0.05 indicates a slight deterioration in the security situation. When determining protection priority based on the situation change rate, a higher situation change rate indicates a more rapid deterioration in the security situation and a higher protection priority. For example, if the change rate threshold is set to 0.1, and the current situation change rate of 0.05 is below the threshold, the protection priority is medium.

[0075] When calculating node weights based on the chronological order of node access along the diffusion path in the attack scenario graph, it's necessary to analyze the node's position and role in the attack propagation process. Suppose the attack scenario graph shows device A as the attack's starting point, passing through devices B and C to reach device D. Device A has the earliest access time and a weight of 1.0; device B is next, with a weight of 0.8; device C is next, with a weight of 0.6; and device D is the latest, with a weight of 0.4. When adjusting the protection priority to determine the protection node sequence, it's necessary to comprehensively consider the situation change rate and node weights. The node weight is multiplied by the protection priority to obtain the adjusted priority. Since the current protection priority is medium (assuming it's 0.5), the adjusted priority for device A is 1.0 x 0.5 = 0.5, for device B 0.8 x 0.5 = 0.4, for device C 0.6 x 0.5 = 0.3, and for device D 0.4 x 0.5 = 0.2. Sorting the adjusted priorities from highest to lowest yields the protection node sequence: device A, device B, device C, device D.

[0076] When developing a protection strategy based on the protected node sequence and the node access timing along the diffusion path, appropriate protection measures should be implemented based on the characteristics and locations of different nodes. For originating device A, intrusion detection and access control can be strengthened; for intermediate nodes B and C, traffic monitoring and abnormal behavior analysis can be implemented; and for endpoint device D, data protection and backup and recovery mechanisms can be strengthened. This targeted protection strategy can effectively block the attack diffusion path and reduce security risks.

[0077] In this embodiment, feature standardization and normalization address the difficulty of directly comparing different types of data, improving model adaptability. Inter-feature mutual information calculation and redundancy analysis effectively remove redundant features, reducing computational complexity while retaining key information. A cross-validation mechanism ensures optimal kernel function parameter selection and enhances model generalization. Weighted smoothing within the time window eliminates the interference of short-term fluctuations, making the situation index more stable and reliable. Situation change rate calculation provides a quantitative basis for determining protection priorities. Combined with the node timing weight adjustment mechanism in the attack scenario graph, precise allocation of protection resources is achieved. Protection strategies based on node access timing are highly targeted and can intercept attacks at key stages of propagation, effectively blocking the attack's spread path. The overall solution forms a complete closed loop from data processing, feature selection, model training, situation assessment, and protection decision-making, significantly improving the accuracy, foresight, and operability of industrial Internet security situation awareness and providing scientific decision-making support for security operations teams.

[0078] In an optional implementation, the node centrality of each industrial Internet device is calculated based on the attack scenario graph, a protection decision tree is constructed in combination with the protection strategy, and the optimal protection node combination is selected based on the node centrality and the protection decision tree. The protection instructions are generated including: The degree centrality is calculated by obtaining the alarm reception and propagation volume of each industrial Internet device node in the attack scenario graph. The betweenness centrality is obtained by calculating the number of alarm forwarding times of the node based on the attack scenario graph. The closeness centrality is obtained by calculating the shortest hop count between nodes. Extracting the alarm processing time sequence of the node in the attack scenario graph, calculating the alarm propagation delay of the node, adaptively weighting the degree centrality, betweenness centrality, and closeness centrality according to the propagation delay, and combining them to obtain the node centrality; Based on the node centrality, the industrial Internet devices are divided into importance levels, the protection measures are divided into different protection levels according to the protection requirements corresponding to the importance levels, a protection decision tree is constructed, and decision rules based on node centrality and propagation delay are set at the branch nodes of the protection decision tree; Selecting nodes to be protected at each protection level according to the decision rule, calculating the alarm processing efficiency of the selected nodes, calculating the protection priority based on the position of the nodes in the attack scenario graph, and matching the protection priority with the protection level to generate a candidate protection node combination; Calculate the alarm blocking effect of each candidate protection node combination, select the protection node combination with the best blocking effect and meeting the resource constraints, and generate protection instructions based on the protection node combination.

[0079] In this embodiment, when calculating the degree centrality by obtaining the alarm reception and propagation amount of each industrial Internet device node in the attack scenario graph, the number of connections of each node is counted. In actual applications, the attack scenario graph contains multiple industrial Internet device nodes, such as PLC controllers, SCADA servers, industrial gateways, smart sensors, etc. By analyzing the alarm flow records, it can be statistically learned that the PLC controller received alarms 8 times and propagated alarms 5 times, with a total number of connections of 13 and a degree centrality of 13; the SCADA server received alarms 10 times and propagated alarms 12 times, with a total number of connections of 22 and a degree centrality of 22; the industrial gateway received alarms 15 times and propagated alarms 8 times, with a total number of connections of 23 and a degree centrality of 23; the smart sensor received alarms 4 times and propagated alarms 2 times, with a total number of connections of 6 and a degree centrality of 6. The higher the degree centrality value, the more active the device is in the alarm propagation network, and it may be an important node for attack propagation.

[0080] When calculating the number of times a node forwards alarms based on the attack scenario graph to obtain betweenness centrality, we analyze the frequency of each node acting as a relay station. By traversing all alarm propagation paths in the attack scenario graph, we count the number of paths forwarded by each device. PLC controllers participated in seven forwarding paths, with a betweenness centrality of 7; SCADA servers participated in 18 forwarding paths, with a betweenness centrality of 18; industrial gateways participated in 15 forwarding paths, with a betweenness centrality of 15; and smart sensors participated in two forwarding paths, with a betweenness centrality of 2. Nodes with high betweenness centrality often act as "bridges" in the network, controlling the flow of information between different areas.

[0081] When calculating the shortest hop count between nodes to obtain proximity centrality, the average distance from each node to other nodes is analyzed. Closeness centrality is inversely proportional to the average hop count; the shorter the average distance from a node to other nodes, the higher its closeness centrality. Using a breadth-first search algorithm, the average hop count from a PLC controller to all other nodes was calculated to be 2.1, with a closeness centrality of 0.48. The average hop count for a SCADA server was 1.5, with a closeness centrality of 0.67. The average hop count for an industrial gateway was 1.3, with a closeness centrality of 0.77. The average hop count for a smart sensor was 2.8, with a closeness centrality of 0.36. Nodes with high closeness centrality can quickly reach other nodes in the network and play a crucial role in the dissemination of alarm information.

[0082] By extracting the alarm processing timing of nodes in the attack scenario graph and calculating the node's alarm propagation delay, we analyze the time required for alarm propagation between nodes. Through timestamp analysis, we calculated that the average alarm propagation delay for PLC controllers is 350 milliseconds; the average alarm propagation delay for SCADA servers is 180 milliseconds; the average alarm propagation delay for industrial gateways is 150 milliseconds; and the average alarm propagation delay for smart sensors is 420 milliseconds. Nodes with lower propagation delays process alarms more efficiently and are therefore more valuable in defense decisions.

[0083] The three centrality metrics are adaptively weighted based on propagation delay. When combined to obtain node centrality, a delay-inverse weighting mechanism is designed. The smaller the propagation delay, the greater the weight of the corresponding centrality metric. The base weights for degree centrality, betweenness centrality, and closeness centrality are set to 0.4, 0.3, and 0.3, respectively. Based on the propagation delay of each device, the weights for the PLC controller are adjusted to [0.35, 0.3, 0.35]; the weights for the SCADA server are [0.4, 0.35, 0.25]; the weights for the industrial gateway are [0.45, 0.3, 0.25]; and the weights for the smart sensor are [0.3, 0.35, 0.35]. Multiplying these adjusted weights by the corresponding centrality metric and summing them yields a node centrality of 6.92 for the PLC controller, 12.57 for the SCADA server, 13.69 for the industrial gateway, and 2.53 for the smart sensor.

[0084] When categorizing the importance of Industrial Internet devices based on node centrality, a grading standard is established. A node centrality greater than 10 is considered high importance, between 5 and 10 is medium importance, and less than 5 is considered low importance. Based on this, SCADA servers and industrial gateways are considered high importance devices, PLC controllers are considered medium importance devices, and smart sensors are considered low importance devices. This grading provides a foundation for subsequent protection decisions.

[0085] Protection measures are divided into different levels based on the protection requirements corresponding to the importance level. When constructing a protection decision tree, a three-tiered protection system is designed. High-importance devices utilize advanced protection measures, including deep packet inspection, behavioral anomaly analysis, and real-time monitoring; medium-importance devices utilize intermediate protection measures, including basic intrusion detection and access control; and low-importance devices utilize basic protection measures, including regular security scans and basic firewall rules. The root node of the protection decision tree represents the device importance determination; the second tier is the propagation delay determination; and the third tier is the selection of specific protection measures.

[0086] When setting decision rules based on node centrality and propagation delay for branch nodes in the protection decision tree, protection strategies are developed based on the characteristics of different nodes. For high-importance nodes with propagation delays less than 200 milliseconds (such as industrial gateways), real-time deep behavioral analysis is prioritized. For high-importance nodes with propagation delays greater than or equal to 200 milliseconds (such as SCADA servers), preventive access control is prioritized. For medium-importance nodes (such as PLC controllers), basic intrusion detection systems are used. For low-importance nodes (such as smart sensors), periodic security scans are used.

[0087] Based on decision-making rules, nodes to be protected were selected at each protection level. When calculating the alarm processing efficiency of the selected nodes, the historical data was analyzed to determine the percentage of successful alarm processing for each device. The alarm processing efficiency for industrial gateways was 92%, for SCADA servers 88%, for PLC controllers 85%, and for smart sensors 75%. The protection priority was calculated based on the node's position in the attack scenario graph, taking into account the node's upstream and downstream connections. Industrial gateways connect to multiple critical subnets and have a protection priority of 0.95; SCADA servers control multiple industrial processes and have a protection priority of 0.9; PLC controllers are responsible for key equipment operations and have a protection priority of 0.8; and smart sensors, which function relatively independently, have a protection priority of 0.6.

[0088] When matching protection priorities with protection levels to generate candidate protection node combinations, the node's importance level and protection priority are comprehensively considered. The generated candidate combinations include: Combination 1 (industrial gateway and SCADA server), Combination 2 (industrial gateway and PLC controller), Combination 3 (SCADA server and PLC controller), and Combination 4 (industrial gateway, SCADA server, and smart sensor).

[0089] When calculating the alarm blocking effectiveness of each candidate protection node combination, we evaluated the proportion of potential attack paths each combination could block. Combination 1 blocked 88% of attack paths, combination 2 blocked 75% of attack paths, combination 3 blocked 70% of attack paths, and combination 4 blocked 90% of attack paths. Considering resource constraints, assuming the system limits simultaneous protection to a maximum of two nodes, we selected the protection node combination with the best blocking effectiveness and meeting resource constraints: combination 1 (industrial gateway and SCADA server).

[0090] When generating protection instructions based on the combination of protection nodes, specific protection configurations are developed for the selected devices. Deep packet inspection instructions are generated for industrial gateways, rules are configured to identify abnormal traffic patterns, and real-time behavior analysis modules are deployed. Access control enhancement instructions are generated for SCADA servers, fine-grained permission management is configured, and abnormal operation audits are implemented. These protection instructions are distributed to the corresponding devices through the security management platform, achieving precise protection for the industrial internet environment and effectively improving network security. The execution results of the protection instructions are recorded and analyzed by a support vector regression model, providing data support for subsequent security assessments and forming a closed-loop optimization mechanism.

[0091] In this embodiment, through the comprehensive analysis of degree centrality, betweenness centrality and closeness centrality, the system can comprehensively identify key nodes in the network, and is not limited to the importance assessment of a single dimension in traditional methods. The introduction of an adaptive weighting mechanism for alarm propagation delay makes the node importance assessment more consistent with the actual network operation characteristics and improves the accuracy of the assessment. The equipment importance classification based on node centrality provides a scientific basis for the rational allocation of protection resources and avoids waste of resources. The construction of a protection decision tree combines expert experience with data analysis to make the protection strategy more targeted and adaptable. The decision rules based on node characteristics and location achieve precise matching of protection measures and improve protection efficiency. The generation and evaluation mechanism of candidate protection node combinations ensures the optimal protection effect under resource constraints and balances protection costs and security benefits.

[0092] A second aspect of an embodiment of the present invention provides an industrial Internet security situation analysis system based on support vector regression, the system comprising: The first unit is used to collect alarm data from industrial Internet devices. Based on the time series properties of the alarm data and the network topology between industrial Internet devices, it uses a time window sliding method to perform correlation analysis on the alarm data and identify homologous alarm events with causal relationships. The second unit is used to establish a time-series propagation matrix based on homologous alarm events, calculate the probability of alarm propagation between industrial Internet devices, construct diffusion paths, and build an attack scenario graph based on the overlap and coverage of multiple diffusion paths. In the attack scenario graph, the propagation characteristics of the alarm data and the number of affected devices are extracted to generate a situation assessment dataset. The third unit is used to input the situation assessment dataset into the pre-trained support vector regression model, calculate the security situation index, determine the protection priority, and formulate the protection strategy based on the diffusion path in the attack scenario graph; The fourth unit is used to calculate the node centrality of each industrial Internet device based on the attack scenario graph, build a protection decision tree in combination with the protection strategy, select the optimal protection node combination based on the node centrality and the protection decision tree, generate protection instructions, and send the protection instructions to the affected industrial Internet devices to implement active protection.

[0093] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0094] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0095] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. The industrial Internet security situation analysis method based on support vector regression is characterized by: include: Collect alarm data from industrial Internet devices and, based on the time series properties of the alarm data and the network topology between industrial Internet devices, perform correlation analysis on the alarm data using a time window sliding method to identify homologous alarm events with causal relationships. Based on homologous alarm events, a time-series propagation matrix is established to calculate the alarm propagation probability between industrial Internet devices and construct a diffusion path. An attack scenario graph is constructed based on the overlap and coverage of multiple diffusion paths. The propagation characteristics of the alarm data and the number of affected devices are extracted from the attack scenario graph to generate a situation assessment dataset. Input the situation assessment dataset into the pre-trained support vector regression model to calculate the security situation index, determine the protection priority, and formulate a protection strategy based on the diffusion path in the attack scenario graph; Based on the attack scenario graph, the node centrality of each industrial Internet device is calculated, and a protection decision tree is constructed in combination with the protection strategy. The optimal protection node combination is selected based on the node centrality and the protection decision tree, and protection instructions are generated. The protection instructions are sent to the affected industrial Internet devices to implement active protection.

2. The method according to claim 1, characterized in that Collect alarm data from industrial Internet devices. Based on the time series properties of the alarm data and the network topology between industrial Internet devices, use a time window sliding method to perform correlation analysis on the alarm data to identify homologous alarm events with causal relationships, including: Collect alarm data from industrial Internet devices, set a corresponding reference time window according to the alarm level of the alarm data, and construct a network connection diagram based on the network topology relationship between industrial Internet devices, the network connection diagram including device nodes and communication links; An attention-based density clustering algorithm is used to calculate the alarm density of the alarm data, and a reference time window is adaptively adjusted according to the alarm density and the number of shortest path hops between device nodes in a network connection graph to obtain a time window. Within the time window, calculating the time series similarity based on the time series attributes of the alarm data, calculating the path similarity based on the shortest path between the device nodes in the network connection graph, and calculating the feature similarity based on the attribute features of the alarm data; Calculating a propagation weight coefficient of the alarm data, performing a weighted combination of the time series similarity, path similarity, and feature similarity according to the propagation weight coefficient to obtain an alarm propagation probability, and constructing an alarm propagation graph, wherein nodes are alarm data and edges are alarm propagation relationships with alarm propagation probabilities; Based on the recursive search algorithm, the alarm propagation graph is dynamically verified for paths, the alarm propagation relationships that do not meet the propagation timing constraints are identified and eliminated, and multiple connected subgraphs are obtained. The connected subgraphs are then verified for timing consistency to identify homologous alarm events with causal relationships.

3. The method according to claim 2, characterized in that Dynamic path verification is performed on the alarm propagation graph based on a recursive search algorithm to identify and eliminate alarm propagation relationships that do not meet propagation timing constraints, thereby obtaining multiple connected subgraphs. Temporal consistency verification is performed on the connected subgraphs to identify homologous alarm events with causal relationships, including: Calculating a propagation delay attenuation coefficient according to the number of network hops between the alarm data, and determining a propagation timing constraint between adjacent alarm data based on the propagation delay attenuation coefficient; Constructing an alarm propagation rule based on the alarm type of the alarm data, preprocessing the alarm propagation graph according to the alarm propagation rule, deleting the alarm propagation relationship that does not meet the alarm propagation rule, and obtaining a preprocessed alarm propagation graph; Perform a depth-first search on the source alarm data in the pre-processed alarm propagation graph. During the search, determine whether the time intervals between adjacent alarm data satisfy the propagation timing constraints, and mark the alarm data that violates the propagation timing constraints as invalid alarm data. Skipping the path search starting from the invalid alarm data, updating the reachability mark of the alarm data, and saving the alarm propagation relationship that meets the propagation timing constraint as a valid alarm propagation relationship; According to the connectivity of the effective alarm propagation relationship, the preprocessed alarm propagation graph is decomposed into multiple connected sub-graphs, and the propagation probability of the alarm propagation relationship in the connected sub-graph and the timing consistency index of the alarm timestamp are calculated. Based on the timing consistency index, the connected sub-graph is verified for timing consistency to identify homologous alarm events with causal relationships.

4. The method according to claim 1, wherein Based on the same-source alarm events, a time series propagation matrix is established to calculate the alarm propagation probability between industrial Internet devices. The diffusion path is constructed, including: Construct an initial time series propagation matrix based on homologous alarm events. Use the industrial Internet devices in the homologous alarm events as row vectors and column vectors. Count the number of alarm propagations between each pair of devices within a specified time window. For any two industrial Internet devices, calculate the alarm propagation probability based on the ratio of the number of target device alarms generated after the source device generates an alarm to the total number of source device alarms. Fill the alarm propagation probability into the corresponding matrix element position to generate a time series propagation matrix. The alarm propagation probability less than the preset probability threshold is set to zero to obtain a filtered time series propagation matrix. Based on the filtered time series propagation matrix, a depth-first search method is used to search for the alarm propagation path starting from the first alarm device. The alarm propagation probability on each path is multiplied as the path weight, and a weight threshold is set for the path weight. The paths greater than the weight threshold are retained as diffusion paths.

5. The method according to claim 1, wherein An attack scenario graph is constructed based on the overlap and coverage of multiple diffusion paths. The propagation characteristics of the alarm data and the number of affected devices are extracted from the attack scenario graph to generate a situation assessment dataset including: Constructing a path connection relationship table based on the diffusion path, wherein the path connection relationship table records the device node sequence passed by each diffusion path, and counting the number of repeated device nodes in the path connection relationship table to obtain the path overlap degree; The device node sequence of each diffusion path is traversed to obtain the total number of nodes accessible by the path, and the ratio of the total number of accessible nodes to the total number of network nodes is used as the path coverage; Dividing diffusion paths with the same repeated nodes into the same group according to the degree of path overlap, calculating the coverage of the paths in each group, and selecting the path with the largest coverage to construct an attack scenario graph; Identify the propagation feature nodes in the attack scenario graph, extract the propagation features based on the connection relationship between the propagation feature nodes, count the number of device nodes reachable by the propagation feature nodes as the number of affected devices, and generate a situation assessment data set.

6. The method according to claim 1, characterized in that The situation assessment dataset is input into the pre-trained support vector regression model to calculate the security situation index, determine the protection priority, and formulate a protection strategy based on the diffusion path in the attack scenario graph, including: Extract the alarm propagation feature sequence and the affected device number sequence from the situation assessment dataset, calculate the standardized feature based on the maximum and minimum values of the alarm propagation feature sequence, and calculate the normalized feature based on the ratio of the affected device number sequence to the total number of devices; The standardized features and the normalized features are combined to construct a feature matrix, the mutual information values between the features in the feature matrix are calculated to obtain feature redundancy, features are selected based on the feature redundancy and information gain is calculated, and training features are obtained when the information gain is less than a preset gain threshold; Dividing the training features into a training set and a validation set and performing cross-validation to obtain optimal kernel function parameters, inputting the training features and kernel function parameters into a support vector regression model, calculating an error value between a predicted result and a true result, updating the model parameters according to the error value, and obtaining a trained support vector regression model when the error value is less than a preset error threshold; Inputting the situation assessment data set into the trained support vector regression model to obtain a security situation index, and performing weighted smoothing processing on the security situation index within a preset time window to obtain a smoothed situation index; The difference in smoothed situation index between adjacent time points is calculated to obtain the situation change rate. The protection priority is determined based on the situation change rate. The node weight is calculated based on the time sequence of the nodes visited on the diffusion path in the attack scenario graph. The protection priority is adjusted to obtain a protection node sequence. A protection strategy is formulated based on the protection node sequence and the node access timing of the diffusion path.

7. The method according to claim 1, characterized in that Based on the attack scenario graph, the node centrality of each industrial Internet device is calculated. A protection decision tree is constructed in combination with the protection strategy. The optimal protection node combination is selected based on the node centrality and the protection decision tree. The generated protection instructions include: The degree centrality is calculated by obtaining the alarm reception and propagation volume of each industrial Internet device node in the attack scenario graph. The betweenness centrality is obtained by calculating the number of alarm forwarding times of the node based on the attack scenario graph. The closeness centrality is obtained by calculating the shortest hop count between nodes. Extracting the alarm processing time sequence of the node in the attack scenario graph, calculating the alarm propagation delay of the node, adaptively weighting the degree centrality, betweenness centrality, and closeness centrality according to the propagation delay, and combining them to obtain the node centrality; Based on the node centrality, the industrial Internet devices are divided into importance levels, the protection measures are divided into different protection levels according to the protection requirements corresponding to the importance levels, a protection decision tree is constructed, and decision rules based on node centrality and propagation delay are set at the branch nodes of the protection decision tree; Selecting nodes to be protected at each protection level according to the decision rule, calculating the alarm processing efficiency of the selected nodes, calculating the protection priority based on the position of the nodes in the attack scenario graph, and matching the protection priority with the protection level to generate a candidate protection node combination; Calculate the alarm blocking effect of each candidate protection node combination, select the protection node combination with the best blocking effect and meeting the resource constraints, and generate protection instructions based on the protection node combination.

8. An industrial Internet security situation analysis system based on support vector regression, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to collect alarm data from industrial Internet devices. Based on the time series properties of the alarm data and the network topology between industrial Internet devices, it uses a time window sliding method to perform correlation analysis on the alarm data and identify homologous alarm events with causal relationships. The second unit is used to establish a time-series propagation matrix based on homologous alarm events, calculate the probability of alarm propagation between industrial Internet devices, construct diffusion paths, and build an attack scenario graph based on the overlap and coverage of multiple diffusion paths. In the attack scenario graph, the propagation characteristics of the alarm data and the number of affected devices are extracted to generate a situation assessment dataset. The third unit is used to input the situation assessment dataset into the pre-trained support vector regression model, calculate the security situation index, determine the protection priority, and formulate the protection strategy based on the diffusion path in the attack scenario graph; The fourth unit is used to calculate the node centrality of each industrial Internet device based on the attack scenario graph, build a protection decision tree in combination with the protection strategy, select the optimal protection node combination based on the node centrality and the protection decision tree, generate protection instructions, and send the protection instructions to the affected industrial Internet devices to implement active protection.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • A method for locating abnormal root causes of cloud data centers based on statistical analysis

    CN109254865A

  • Network security situation prediction method based on historical alarm event times

    CN117319074A

  • Fault determination method, electronic equipment, storage medium and computer product

    CN118802466A

  • Transmission network fault positioning method, system and device and storage medium

    CN119520248A

  • Operation and maintenance alarm processing method and system based on knowledge graph enhanced large model

    CN119988154A

Cited By

  • Machine room air conditioner alarm cooperative processing method and system for equipment group control

    CN121262811A

  • Device group control-oriented machine room air conditioner alarm cooperative processing method and system

    CN121262811B

  • Industrial network security situation prediction method and system based on generative large model

    CN121388958A

  • Network fault diagnosis method based on dynamic weighted graph

    CN121418276A

  • Multi-layer depth active protection method and system for safety platform

    CN122339850A