An industrial internet security capability evaluation framework
By constructing business asset traffic triples and time-series causal graphs and dynamically adjusting weight coefficients, the problems of data fragmentation and model staticization in industrial internet security assessment are solved, enabling efficient and interpretable security assessment of industrial internet systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 北京中关村实验室
- Filing Date
- 2026-04-22
- Publication Date
- 2026-07-03
AI Technical Summary
Existing industrial internet security assessment methods suffer from problems such as data semantic fragmentation, temporal misalignment, static assessment models that are insensitive to process changes, neglect of cross-node and cross-stage coupling effects, insufficient observability of closed-loop handling, and weak causal relationship mining. These issues lead to disjointed assessment results, lack of quantitative recovery capabilities, and inoperable results.
Collect multimodal industrial characteristic data, construct business asset traffic triples, build a security capability coupling assessment model through time-series causal graphs, dynamically adjust weight coefficients, generate dynamic security capability assessment results, quantify the nonlinear impact of protection, detection and response capabilities, and optimize the assessment model by combining causal discovery algorithms.
It enables the alignment of multi-source data on the process timeline, quantifies observation coverage and depth, reveals cross-node linkages and time delays, and generates interpretable and practical evaluation conclusions, thereby improving the objectivity and interpretability of the evaluation results and enabling dynamic adaptation to process changes.
Smart Images

Figure CN122339795A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of industrial internet security, and specifically relates to an industrial internet security capability assessment framework. Background Technology
[0002] In industrial internet scenarios, production control networks and information networks are deeply integrated. Control commands exhibit strong periodicity, strong temporality, and high determinism. Network structures include multi-layered partitions, protocol gateways, and remote maintenance channels. Devices are also generally constrained by resource limitations and change window restrictions. In this environment, common security assessments in the industry still primarily rely on compliance comparisons, vulnerability scanning, or summarizing logs from individual devices, outputting a linearly weighted "security index." These methods struggle to depict the linkage mechanism between business and security in real-world operation, primarily due to the following deep-seated problems:
[0003] Data semantic fragmentation and temporal misalignment are prominent issues. Asset fingerprints are difficult to map to specific processes and roles, and control traffic is mostly based on proprietary protocols (such as S7, EtherNet / IP, Modbus, and OPC UA). Relying solely on five-tuples or port statistics cannot reconstruct instruction semantics and temporal dependencies. Timestamps from different systems often deviate due to NTP / PTP asynchrony, and buffering and jitter occur during data transmission from edge nodes, making it difficult to align maintenance changes, alarms, and control events on a unified process timeline.
[0004] Observation coverage and depth cannot be quantified. Mirror ports on production switches are prone to congestion or rate limiting; SPAN / ERSPAN links suffer from packet loss and sampling bias; critical paths are not always covered by mirroring or probes. Protocol parsing depth is limited by probe type and decoding capabilities, easily resulting in the loss of critical information such as objects, function codes, and register areas, causing anomaly identification to remain at the connection level. Although boundary policies exist, cross-domain access can form covert paths through engineering workstations, jump servers, VPN tunnels, or protocol conversion. Existing assessments lack verifiable quantification regarding whether policies substantially block critical paths.
[0005] The assessment models are static and insensitive to process changes. Adjustments to production cycle time, production line changes, and equipment replacements can cause conceptual drift in instruction sequences and flow patterns; fixed thresholds and historical weights cannot adapt to changes in the scenario, resulting in both false positives and false negatives, and causing scoring to lag behind risk. Most methods lack dynamic baselines based on process stages, failing to reflect the scope of compliance behavior for the same asset at different stages.
[0006] The coupling effects across nodes and stages are ignored. Existing assessments typically simply add up the indicators of different domains and functional units, failing to characterize the linkage, amplification, or inhibition effects brought about by business data flow and control dependencies. For example, anomalies in upstream units can be transmitted to downstream units through the instruction chain, and local policy or monitoring adjustments may unexpectedly change overall robustness; linear weighting cannot reveal the nonlinear relationships of "who affects whom, the strength of the impact, and the time lag in which it manifests."
[0007] The observability of the closed-loop process is insufficient. The fields of cross-system work orders, alarms and device logs are heterogeneous, and the timestamp benchmarks are inconsistent, making it difficult to objectively quantify the time spent in each stage from discovery to location to recovery. Some key stages rely on manual and offline approval, which leads to missing data being ignored or arbitrarily assigned values, thereby distorting the timeliness of the assessment results.
[0008] The causal relationship mining is weak. Most practices substitute rules or correlations for causality, failing to systematically estimate the lag relationship and uncertainty of "change-alarm" from history, making it difficult to explain the triggering factors behind alarm storms, and even more difficult to feed back the causal backtracking results to the evaluation parameters to form an adaptive mechanism.
[0009] The combination of these problems directly leads to three consequences: First, the assessment is disconnected from the actual process, failing to accurately reflect phased risks; second, there is a lack of quantification regarding recovery capabilities after disturbances, making it impossible to answer the question of "how long it takes to recover to a steady state under a specific disturbance"; and third, the results are not actionable, making it difficult to pinpoint weaknesses and prioritize optimization. The industry urgently needs a system that can align multi-source data along the process timeline, extract standard behavioral relationships between business, instructions, and flow, quantify observation coverage and depth, and reveal cross-node linkages and time lags through temporal and causal perspectives, thereby obtaining assessment conclusions that are synchronized with production operations, interpretable, and actionable. Summary of the Invention
[0010] To address the aforementioned problems in existing technologies, namely data fragmentation, static models, lack of cross-node coupling relationships, and difficulty in quantifying resilience, this invention provides an industrial internet security capability assessment framework, comprising the following steps:
[0011] Collect multimodal industrial characteristic data, including asset fingerprint data, industrial control network flow data, security equipment alarm data, and operation and maintenance logs;
[0012] The process role in the fingerprint data of the asset is fused with the instruction sequence in the flow data of the industrial control network at the feature level to construct a business asset traffic triplet and form a business security baseline.
[0013] Assets in the industrial internet system are identified as multiple evaluation nodes. Protection capabilities are determined based on the configuration compliance and boundary protection capabilities of each evaluation node. Detection capabilities are determined based on the probe deployment and traffic coverage of each evaluation node. Response capabilities are determined based on the closed-loop processing time of each evaluation node.
[0014] A security capability coupling assessment model is constructed based on a time-series causal graph. The protection capability, detection capability, and response capability of each assessment node are defined as interacting nonlinear system nodes. The nonlinear influence weight relationship of each security capability among different nonlinear system nodes is characterized by a collaborative influence matrix. The steady-state recovery time of the multiple assessment nodes under virtual disturbance is solved according to the business security baseline and the collaborative influence matrix to obtain the resilience coefficient.
[0015] Align the time window with the process cycle in the industrial scenario, determine the change event based on the operation and maintenance log, analyze the causal lag time between the change event and the alarm event in the historical data through the causal discovery algorithm, and dynamically adjust the weight coefficients in the collaborative influence matrix based on the causal backtracking results.
[0016] The dynamic safety capability assessment results are generated and output based on the adjusted weighting coefficients and the resilience coefficients.
[0017] Furthermore, the process role in the fingerprint data of the asset is fused with the instruction sequence in the flow data of the industrial control network at the feature level to construct a business asset traffic triplet, forming a business security baseline. The method is as follows:
[0018] The fingerprint data of the assets is analyzed to extract the process stage affiliation and process execution role of each industrial asset, and the flow data of the industrial control network is analyzed to extract the instruction sequence reflecting the control logic and its temporal dependencies.
[0019] The process stage is aligned with the instruction sequence through a process execution time window. Within the time window, a mapping relationship is established between the process execution role and the control instructions in the instruction sequence, forming a triplet structure that includes a business asset identifier, the standard operation feature corresponding to the process execution role, and the traffic interaction object associated with the standard operation feature.
[0020] Using the triplet structure as a unit, the matching patterns of business assets and traffic behavior are accumulated and recorded in multiple consecutive process cycles. By statistically analyzing the recurrence frequency and fluctuation range of the matching patterns, a business security baseline that is dynamically updated with the process cycle of the industrial scenario is generated.
[0021] The business security baseline is used to characterize the standard compliance relationship between business assets and traffic behavior at each process stage.
[0022] Furthermore, protection capabilities are determined based on the configuration compliance and boundary protection capabilities of each assessment node; detection capabilities are determined based on the probe deployment and traffic coverage of each assessment node; and response capabilities are determined based on the handling closed-loop time of each assessment node. The method is as follows:
[0023] For each evaluation node, the security configuration parameters and boundary access control policies of the industrial control system under the evaluation node are extracted. The protection capability of the evaluation node is determined by the degree of matching between the security configuration parameters and the compliance baseline and the degree of isolation of cross-security domain access by the boundary access control policies.
[0024] The coverage and protocol parsing depth of the probes deployed in the network topology of the evaluation node are analyzed, and the traffic coverage of the network interfaces under the jurisdiction of the evaluation node is collected. The detection capability of the evaluation node is determined by the correlation and fusion of the probe coverage, protocol parsing depth and traffic coverage.
[0025] The response capability of the assessment node is determined by analyzing the time consumed in each step of the security incident handling process of the assessment node and using the ratio of the preset baseline period to the sum of the times consumed in each step. The missing steps are accumulated using the preset maximum time consumed.
[0026] Furthermore, the protection capability is determined based on the configuration compliance and boundary protection capabilities of each assessment node, using the following method:
[0027] For each evaluation node, extract the security configuration parameter set and boundary access control policy set of the industrial control system under the jurisdiction of the evaluation node;
[0028] The security configuration parameter set is compared with the preset compliance baseline library to mark deviation items, and the boundary access control policy set is mapped into an access control matrix;
[0029] By traversing the access path entries in the access control matrix, unauthorized access paths across security domains are identified to calculate policy coverage completeness;
[0030] The reverse normalized value of the deviation term is used to characterize the configuration compliance level, and the policy coverage integrity is used to characterize the boundary isolation strength. The configuration compliance level and the boundary isolation strength are multiplied, and the result of the multiplication is used as the protection capability of the evaluation node.
[0031] Furthermore, the detection capability is determined based on the probe deployment and traffic coverage of each evaluation node, using the following method:
[0032] For each evaluation node, the deployment location and probe type of the security probes deployed in its network topology are analyzed. The probe coverage is determined based on the matching relationship between the deployment location and the path that the critical traffic must pass through. The protocol parsing depth is determined based on the protocol parsing capability corresponding to the probe type.
[0033] The traffic mirroring configuration of the network interfaces under the jurisdiction of the evaluation node is collected, and the traffic coverage is obtained by comparing the coverage of the mirrored ports with the interaction path of the full traffic.
[0034] The effective detection coverage is represented by the product of the probe coverage area and the protocol parsing depth, and the product of the effective detection coverage and the traffic coverage rate is used as the detection capability of the evaluation node.
[0035] Furthermore, response capability is determined based on the closed-loop processing time of each assessment node, using the following method:
[0036] For each assessment node, the closed-loop cycle of its security incident handling process from the triggering of the detection event to the completion of the handling action is analyzed, and the closed-loop cycle is broken down into detection confirmation time, analysis and location time, and blocking and recovery time.
[0037] Collect event timestamps of the security devices and operation and maintenance systems under the jurisdiction of the evaluation node, extract the detection confirmation time, the analysis and location time and the blocking and recovery time respectively by the difference of event timestamps, and use the sum of the detection confirmation time, the analysis and location time and the blocking and recovery time as the total closed-loop time of the handling.
[0038] The response capability of the evaluation node is characterized by the ratio of a preset baseline period to the total closed-loop processing time.
[0039] If any processing step is missing in the evaluation node, the time of that processing step is set to the preset maximum threshold time and added to the accumulation.
[0040] Furthermore, the nonlinear influence weight relationship of each security capability among nodes of different nonlinear systems is characterized by a collaborative influence matrix. The method is as follows:
[0041] Using multiple evaluation nodes as row and column indices, and protection capability, detection capability, and response capability as capability components of each evaluation node, a collaborative influence matrix is constructed. Each element in the collaborative influence matrix corresponds to the weighting coefficient between the source capability component of the row index evaluation node and the target capability component of the column index evaluation node.
[0042] Based on the time sequence of dependence of each stage of the security incident attack chain on protection capabilities, detection capabilities, and response capabilities, initial values are assigned to the weighting coefficients of protection capabilities on detection capabilities, detection capabilities on response capabilities, and response capabilities on protection capabilities within the same evaluation node.
[0043] Based on the frequency and direction of business data flow interactions between different evaluation nodes, initial values are assigned to the weight coefficients of the same capability components between different evaluation nodes; these initial values are then filled into the corresponding element positions of the collaborative influence matrix to form an initial collaborative influence matrix.
[0044] Furthermore, based on the business security baseline and the collaborative impact matrix, the steady-state recovery time of the multiple evaluation nodes under virtual disturbances is calculated to obtain the resilience coefficient. The method is as follows:
[0045] Virtual disturbance events are applied to each evaluation node in the time-series cause-effect graph. The virtual disturbance events trigger the initial state value of the protection capability, detection capability, or response capability of the specified evaluation node to deviate from the standard state value defined by the business security baseline.
[0046] Using the influence relationships recorded in the collaborative influence matrix as iterative constraints, the state values of each capability component of each evaluation node are updated round by round within the discrete time step until the state values of each capability component of all evaluation nodes converge to a deviation from the business security baseline of less than a preset threshold.
[0047] The total number of time steps from the application of the virtual disturbance event to the convergence of the state value is counted, and the reciprocal or normalized value of the total number of time steps is used as the resilience coefficient.
[0048] Furthermore, the causal lag time between change events and alarm events in historical data is analyzed using a causal discovery algorithm. Based on the causal backtracking results, the weight coefficients in the collaborative influence matrix are dynamically adjusted. The method is as follows:
[0049] Collect operation and maintenance logs and security device alarm logs within a historical time window, extract change event sequences from the operation and maintenance logs, and extract alarm event sequences from the alarm logs;
[0050] Using the sequence of change events as the dependent variable sequence and the sequence of alarm events as the resultant variable sequence, the causal lag time between each change event and each alarm event is calculated using a causal discovery algorithm.
[0051] Using the change event and alarm event pairs with causal lag time less than a preset threshold as valid causal pairs, the number of valid causal pairs and the average causal lag time within the scope of each evaluation node are counted to generate causal backtracking results.
[0052] The weight coefficients in the collaborative influence matrix are dynamically adjusted based on the causal backtracking results. The dynamic adjustment takes the number of effective causal pairs and the average causal lag time as inputs, and maps the inputs to the correction amount of the weight coefficients through a preset mapping rule. The weight coefficients of the corresponding evaluation nodes in the collaborative influence matrix are updated with the correction amount.
[0053] Furthermore, the dynamic safety capability assessment result is generated and output based on the adjusted weighting coefficients and the toughness coefficients, and the method is as follows:
[0054] A weight vector is constructed using the weight coefficients of each evaluation node in the adjusted synergistic influence matrix, and a resilience vector is constructed using the resilience coefficients of each evaluation node. The weight vector and the resilience vector are then fused element by element to generate a comprehensive evaluation value for each evaluation node.
[0055] Taking the process stage in an industrial scenario as the dimension, the comprehensive evaluation values of each evaluation node under the same process stage are aggregated to obtain the stage evaluation value of that process stage.
[0056] The safety level of a process stage is determined by the degree of deviation between the stage evaluation value and the preset benchmark value of the process stage.
[0057] The safety level of each process stage is associated with the corresponding assessment node identifier and output to the display interface.
[0058] The beneficial effects of this invention are:
[0059] This solution constructs a triplet of "business asset - process role - instruction sequence" by feature-level fusion of asset fingerprints, industrial control network traffic, alarms, and operation and maintenance logs. This triplet is then rolled over with the process cycle to form a business security baseline, bridging the semantic gap between data from different sources. Standard compliance relationships are established using process stages as alignment windows, reducing deviations caused by differences and asynchronicity in data collection. This ensures comparability of data from different systems and across different periods, thereby mitigating inconsistencies in assessments caused by data fragmentation.
[0060] In terms of capability quantification, this solution uses the product of configuration compliance and boundary isolation strength to represent protection capability, the joint metric of probe coverage, protocol parsing depth, and traffic coverage to represent detection capability, and the ratio of handling loop closure time to the baseline period to represent response capability. Missing components are supplemented using threshold time. These indicators are clearly sourced, standardized, auditable, and reproducible, providing stable input for subsequent models, avoiding reliance on subjective scoring, and thus improving the objectivity and interpretability of the evaluation results.
[0061] To address the shortcomings of traditional assessments that neglect cross-node collaboration, this solution uses a time-series causal graph framework to construct a synergistic influence matrix covering the "protection-detection-response" sequence relationship within a node and the relationship between similar capabilities between nodes. This matrix, combined with the frequency and direction of business data flows, explicitly represents the non-linear influence weights between capabilities, enabling the model to characterize the cross-domain propagation of attack chains and the synergistic effects of capabilities. This avoids evaluating each node in isolation and improves the ability to characterize and locate cross-node coupling relationships.
[0062] In terms of resilience measurement, this scheme iteratively updates the capability status based on the collaborative influence matrix under virtual perturbation, statistically calculates the time step required to converge to within the business security baseline threshold, and normalizes it to obtain the resilience coefficient. This process quantifies "recovery speed" into a comparable value, enabling horizontal comparison of recovery capabilities across different nodes and process stages. It compensates for the shortcomings of traditional static risk scoring in reflecting dynamic recovery capabilities, providing a basis for recovery path optimization.
[0063] To mitigate the lag and mismatch caused by static modeling, this solution aligns the time window with the process cycle and combines a causal discovery algorithm to analyze the causal lag relationship between change events and alarm events, dynamically adjusting the weights of the collaborative influence matrix. Through causal backtracking of maintenance changes and alarms, the model can adaptively adjust the impact intensity after changes in equipment configuration, topology, and load. This ensures that the evaluation results remain stable with the process rhythm while also being sensitive to risk drift, reducing manual recalibration workload and misjudgments caused by environmental changes.
[0064] Ultimately, this solution integrates dynamic weights and resilience coefficients to output phased assessments and ratings, enabling the assessment results to be aggregated and presented at the business stage level, facilitating the identification of weak links and critical paths. These results can guide probe addition, boundary strategy optimization, and process streamlining, and also facilitate the formation of a governance closed loop based on objective indicators in multi-departmental collaborations. Technically, this provides targeted improvements to issues such as data fragmentation, static models, missing cross-node coupling relationships, and difficulty in quantifying resilience. Attached Figure Description
[0065] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0066] Figure 1 This is a schematic diagram of the execution flow of an industrial internet security capability assessment framework according to the present invention. Detailed Implementation
[0067] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0068] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0069] The first embodiment of the present invention provides an industrial internet security capability assessment framework, including the following steps:
[0070] Step S10: Collect multimodal industrial feature data, which includes asset fingerprint data, industrial control network flow data, security device alarm data, and operation and maintenance logs.
[0071] Step S20: Perform feature-level fusion of the process role in the fingerprint data of the asset and the instruction sequence in the flow data of the industrial control network to construct a business asset traffic triplet and form a business security baseline.
[0072] Step S30: Identify assets in the industrial internet system as multiple evaluation nodes, determine the protection capability based on the configuration compliance and boundary protection capability of each evaluation node, determine the detection capability based on the probe deployment and traffic coverage of each evaluation node, and determine the response capability based on the closed-loop processing time of each evaluation node.
[0073] Step S40: Construct a security capability coupling assessment model based on a time-series causal graph. Define the protection capability, detection capability, and response capability of each assessment node as interacting nonlinear system nodes. Characterize the nonlinear influence weight relationship of each security capability among different nonlinear system nodes through a collaborative influence matrix. Solve the steady-state recovery time of the multiple assessment nodes under virtual disturbances based on the business security baseline and the collaborative influence matrix to obtain the resilience coefficient.
[0074] Step S50: Align the time window with the process cycle of the industrial scenario, determine the change event according to the operation and maintenance log, analyze the causal lag time between the change event and the alarm event in the historical data through the causal discovery algorithm, and dynamically adjust the weight coefficients in the collaborative influence matrix according to the causal backtracking results.
[0075] Step S60: Generate and output dynamic safety capability assessment results based on the adjusted weighting coefficients and resilience coefficients.
[0076] To more clearly explain the industrial internet security capability assessment framework of this invention, the following will be combined with... Figure 1The steps in the embodiments of the present invention are described in detail below:
[0077] Step S10: Collect multimodal industrial feature data, which includes asset fingerprint data, industrial control network flow data, security device alarm data, and operation and maintenance logs.
[0078] In this implementation, step S10 employs a combination of layered data acquisition and local preprocessing to acquire asset fingerprints, industrial control network flow data, security device alarm data, and operation and maintenance logs without affecting production real-time performance. The overall strategy is to prioritize a passive approach in the production domain, exercise caution and proactive measures within change windows, minimize extraction and anonymization at the edge, and implement time synchronization and quality control throughout the data acquisition process. Read-only TAPs or mirrored ports are deployed on critical links of the production switch, and security and operation and maintenance data are aggregated in the management area through log brokers and adapters. All transmission channels use two-way authentication and encryption.
[0079] Asset fingerprint data is obtained through a fusion of passive identification and offline engineering file parsing. Passive identification captures and parses protocol interactions such as PROFINET DCP, EtherNet / IP ListIdentity / ForwardOpen, Modbus / TCP 0x2B / 0x0E, OPC UA handshake and read / write, and S7comm handshake and Job / Ack in the production network. It also combines common management messages such as LLDP, ARP, and SNMP to identify the manufacturer, model, firmware, slot and module, protocol stack capabilities, and role candidates. For network and security devices supporting read-only interfaces, SNMPv3 RO is used to pull topology information such as interface, VLAN, ARP, and routing during idle periods with rate limiting. Engineering project files are exported from the engineering station within the operations and maintenance window, and the device list, variable table, task cycle, recipe, and logical segment are parsed offline. The multi-source results are matched for consistency and confidence accumulation using MAC, IP, device serial number and engineering logical identifier as candidate keys. The output is a unified asset ID and fingerprint record, which includes at least a unique identifier, manufacturer / model / firmware, hardware configuration, network address / VLAN / security domain, exposed services and ports, protocol capabilities, process unit and stage attribution candidates and criticality level.
[0080] Industrial control network flow data is sent to a data acquisition network card with hardware timestamp capability via a mirror link, and packet-level time is marked under PTP (IEEE 1588) or PPS time synchronization. The acquisition process captures data with zero copy and performs DPI parsing on industrial protocols to construct a structured element of "session-request-response-object". This extracts quintuples, direction, session duration, request / response time, function code / service, object identifiers (such as register address, OPC UA node ID, S7 variable name), offset and length, return code, RTT, periodic average and jitter, and abnormal code ratio. To reduce sensitivity and bandwidth pressure, load values are retained at the edge using one-way hashes or digests, storing only necessary features. Under high load, only industrial ports are parsed based on a whitelist; other protocols are replaced with flow element statistical digests. When the parsing queue is overloaded, priority is given to ensuring critical process links and control sessions, triggering a controlled degradation strategy.
[0081] Security device alarm data is aggregated through a unified access agent from OT / IT perimeter firewalls, IDS / IPS, industrial protocol anomaly detection, endpoint security, bastion hosts, and security orchestration platforms. Standard formats such as Syslog, CEF, LEEF, or JSON over HTTPS / gRPC are prioritized. Non-standard logs are converted to a unified event model via an adapter, including event ID, type, source device, source / destination asset identifier, five-tuple, industrial protocol and object (if applicable), rule / signature ID, severity and trustworthiness, first / last time, policy version, and handling recommendations. The agent performs deduplication and merging, using device serial numbers and adjacent window clustering to suppress storm-like duplicate alarms; severity is normalized according to a unified mapping table. All alarms enter a reliable buffer queue at the edge, employing breakpoint resumption and two-way confirmation. In case of link failure, alarms are locally cached sequentially and replayed chronologically after recovery, ensuring no loss or duplication.
[0082] Operation logs are obtained through interfaces with ITSM / CMMS / EAM, version configuration repositories, and network / security device configuration auditing interfaces, covering change orders, policy / rule releases, configuration modifications, patch upgrades, account permission adjustments, and remote maintenance sessions. Simultaneously, session auditing is integrated with engineer workstations, bastion hosts, and remote access gateways. Systems with APIs use incremental fetching; legacy systems are accessed via read-only views or security log collection proxies. Network / security device configuration snapshots are periodically generated and differentially analyzed to create structured change entries. All operation events are mapped to a unified change model with fields including change ID, type, object (asset ID, CI, policy package), operator and approval chain, planned and actual window, scope of impact, previous and subsequent version / configuration hashes, status and rollback identifiers, associated sessions, and command summaries, facilitating subsequent causal backtracking.
[0083] To support cross-source alignment and causal analysis, time synchronization employs a dual-track mechanism of "device self-reported time + correction". The production domain collector uses PTP or a local precision clock as a reference, simultaneously storing both self-reported and received times in logs and events. The management area aligns with the central clock via NTP, periodically estimating deviations and drift to establish a correction model. In addition to protocol-level time synchronization, cross-calibration is performed using inherent periodic signals such as PLC scan cycles, periodic control interactions, historical database writes, and report timestamps to identify time zone errors and drift. The original and corrected times are stored along with correction coefficients, and data quality alarms are triggered when deviations exceed thresholds.
[0084] The data acquisition chain adopts an edge-containerized microservice deployment, including modules for data capture and DPI, fingerprint fusion, event access, time synchronization and correction, buffering and playback, and encrypted transmission. Resource contention is avoided through memory queues and backpressure control. Data undergoes anonymization and compression at the edge, and the transmission channel uses two-way certificates and encryption. The system continuously monitors image packet loss rate, parsing queue length, CPU / memory / disk usage, and clock deviation. When thresholds are reached, it automatically downgrades parsing or adjusts the image source and reports an alarm. Before entering the center, four types of data undergo schema verification and version labeling. Missing / abnormal records are tagged with quality control labels and isolated. A unified data dictionary and identification system are used globally. Asset IDs are generated jointly from MAC addresses, serial numbers, and engineering identifiers. Network security domains and process stage codes are consistent with the production planning system. All records carry the source, acquisition version, parsing version, anonymized version, and time-corrected version to ensure traceability.
[0085] In terms of deployment and performance, image mirroring prioritizes coverage of PLC-I / O, PLC-HMI / SCADA, and OT / IT boundary links. For batch processing or discrete manufacturing scenarios, coverage of engineer workstations, historical databases, and MES interfaces is added. Fingerprint collection follows a "passive priority, supported by engineering documents" approach, with alarm and maintenance access implemented in stages according to device maturity. Before going live, gray-scale load testing is conducted during non-production periods, aiming for a critical link image packet loss rate of less than 0.01%, DPI parsing latency at the millisecond level, and log access latency at the second level. The additional load introduced by data collection should not exceed 1% of the link bandwidth and should not increase the controller's CPU load. Through these implementation methods, multimodal data required for subsequent modeling can be obtained in a low-intrusion, highly reliable, and auditable manner, providing sufficient and reliable input for triple construction, baseline generation, capability quantification, coupled modeling, and resilience assessment.
[0086] Step S20: Perform feature-level fusion of the process role in the fingerprint data of the asset and the instruction sequence in the flow data of the industrial control network to construct a business asset traffic triplet and form a business security baseline.
[0087] In this embodiment, step S20 specifically includes:
[0088] Step S21: Analyze the fingerprint data of the assets to extract the process stage affiliation and process execution role of each industrial asset, and analyze the flow data of the industrial control network to extract the instruction sequence reflecting the control logic and its temporal dependencies.
[0089] Step S22: Align the process stage with the instruction sequence through the process execution time window, establish a mapping relationship between the process execution role and the control instructions in the instruction sequence within the time window, and form a triplet structure containing business asset identifier, standard operation feature corresponding to the process execution role, and traffic interaction object associated with the standard operation feature;
[0090] Step S23: Using the triplet structure as a unit, accumulate and record the matching patterns of business assets and traffic behavior in multiple consecutive process cycles. By statistically analyzing the recurrence frequency and fluctuation range of the matching patterns, generate a business security baseline that is dynamically updated with the process cycle of the industrial scenario.
[0091] The business security baseline is used to characterize the standard compliance relationship between business assets and traffic behavior at each process stage.
[0092] In one optional implementation, step S20 is executed by instructions running on a data processing platform, which receives raw data streams from an asset management system, an industrial control network acquisition device, and a security maintenance system. To ensure consistency of time bases across different systems, the processing platform first normalizes the timestamps of each data source: PTP is preferentially used for external time synchronization between edge acquisition probes and switching devices; if NTP and PTP are used interchangeably, the time difference of the same physical event (such as batch start signal, cycle bit flip, recipe distribution) in different logs is calculated within a sliding window, and robust regression is used to correct clock offsets and drifts between sources. Samples that fail to align are removed or marked as low confidence to eliminate the impact of time misalignment of heterogeneous data on subsequent alignment.
[0093] In step S21, the processing platform parses the asset fingerprint data and industrial control network flow data respectively. The asset fingerprint data includes fields such as device identifier, manufacturer model, firmware version, workstation number, security domain / network partition, access topology, and upstream system identifier. Based on a pre-set process resource dictionary and process-workstation-device mapping relationship, the platform maps each asset to a specific process stage and determines its process execution role by combining the device type and deployment location. For example, based on the functional templates of PLC / RTU / HMI / engineering station / driver / I / O module, role labels such as control unit, monitoring unit, parameter issuing unit, and execution unit are obtained. For assets lacking clear information, the system uses a semi-supervised classifier to infer roles based on network partition, VLAN, access control list, and interaction centrality with known assets, and adds a confidence score to the inference results. Industrial control network flow data is collected via bypass mirror ports or TAPs. The platform performs deep decoding of industrial protocols such as S7, EtherNet / IP, Modbus, OPC UA, and PROFINET based on protocol identification, extracting instruction semantics that reflect control logic, such as objects, function codes, service codes, register addresses, classes / instances / attributes, variable node IDs, subscription / publishing relationships, and round-trip latency and retries. Simultaneously, it retains the 5-tuple and session ID as indexes for instruction sequence reconstruction. To characterize timing dependencies, the platform performs session segmentation and sequence reconstruction on message flows within single and cross-links, capturing common instruction sequence fragments using n-gram or partial sequence matching methods, recording the directed dependencies between adjacent instructions and their time interval distribution, and reducing noise for packet loss or sampling deficiencies using maximum likelihood imputation or historical template-based completion strategies.
[0094] In step S22, the platform uses the process execution time window as the alignment framework. The time window is determined by the cycle time parameters of the process stage, and the starting point is triggered by anchor events, such as PLC cycle bit flipping, batch start signal, recipe number change, and robot arm homing completion. If multiple anchor points exist, the platform uses the most stable anchor point with the smallest variance as the primary reference and performs relative correction on other anchor points. For the network instruction sequence within each time window, the platform assigns the message to the sending / receiving asset based on the network interface information in the session endpoint and asset fingerprint; then, combined with the asset's process execution role in that process stage, it establishes a mapping relationship between roles and control instructions. To avoid confusion between occasional management flows and control flows, the platform introduces a whitelist / blacklist strategy and function code filtering during mapping, and constrains the assignment of instructions to roles based on the principle of minimum time lag. This forms a triplet structure, which includes: a business asset identifier (a unique device ID and its process stage context); standard operating characteristics (the set of instructions for the asset in this process stage and role, the frequency and order constraints of the instructions, the time interval distribution of adjacent instructions, the allowable parameter range and register address range, round-trip latency quantiles, retry thresholds, etc.); and traffic interaction objects (the set of peer assets, the peer application endpoint or object identifier, cross-domain direction, and necessary path labels). Standard operating characteristics are extracted from the denoised sequence and represented using statistical summaries, such as the probability of occurrence of each instruction, the support and confidence of the sequence segment, and the quantile interval of the time interval. When the protocol is not fully identified, the platform degenerates into a weak semantic template based on endpoints, ports, and periodic features, and labels the template's confidence level. The triplet is stored with "asset ID-process stage-role" as the primary key, supporting versioning management to accommodate template drift caused by recipe and equipment replacement.
[0095] In step S23, the platform accumulates and records matching patterns of business assets and traffic behavior over multiple consecutive process cycles, using triples as units. A matching pattern is defined as a matching event between the set of instructions and sequence fragments of an asset appearing under a specific role and the characteristics of historical standard operations within a process window. The platform maintains statistics using a combination of sliding window statistics and exponentially weighted updates: for each instruction and sequence fragment, it calculates the frequency of occurrence, occurrence rate, quantile interval of the time interval, and coefficient of variation over the most recent M cycles; for interactive objects, it statistically analyzes the frequency of occurrence and directional stability of the peer asset. To suppress the impact of abnormal cycles, the platform performs outlier detection before updating, for example, by removing extreme samples based on Hampel filtering or median absolute deviation, and then smooths the historical distribution using a forgetting factor. The platform incorporates behaviors that occur more frequently than a threshold but with fluctuations below the threshold into the business safety baseline. The frequency threshold can be set as a fixed proportion or an adaptive quantile threshold based on the process cycle and shifts. The fluctuation range is measured by the coefficient of variation of latency and count or IQR. Simultaneously, low-frequency but process-essential rare operations (such as calibration and switching) are included in the extended baseline using a labeling mechanism to avoid misjudgments within the maintenance window. The business safety baseline is dynamically updated with the process cycle. When a concept drift is observed that lasts for multiple cycles (e.g., a new equipment replacement causing an overall shift in instruction timing or a service code replacement), the platform triggers a baseline version upgrade through distribution comparison and drift detection within the window. During the upgrade, both the old and new templates are tolerated in parallel to ensure evaluation continuity.
[0096] Through the above processing, the business security baseline represents the standard compliance relationship between business assets and traffic behavior at each process stage in a data-driven manner. It includes semantic constraints on "what to do" (instruction set and parameter range), temporal constraints on "when to do it" (sequence order and time interval statistics), and channel constraints on "who to interact with" (interaction object and cross-domain direction). During online operation, the platform calculates the matching degree between the observed data of newly entered windows and the baseline in real time, outputting deviation measures at the instruction layer, sequence layer, and interaction layer, respectively. The deviation is then correlated with template credibility for input and calibration of subsequent protection, detection, and response capability assessment models. If the matching degree of an asset continuously falls below a threshold at a certain stage, the asset-stage-role triplet is marked as an object requiring review, and deviation items are recorded for compliance comparison and operational backtracking. This ensures that the baseline stably reflects process patterns while remaining sensitive and interpretable to changes in real-world scenarios.
[0097] To facilitate implementation, a set of operable threshold examples are provided, with values adaptively selected based on process cycle time and data quality. Regarding time alignment, anchor-aligned residual offsets are more robustly controlled within ±10ms in fast-cycle scenarios of discrete manufacturing (typical PLC cycle 20–50ms), while this can be relaxed to ±1s in batch processing or long-cycle scenarios. When the residual offset of the same physical event in multi-source logs exceeds the above threshold and exceeds the historical average of that source by ±3σ, it is designated as a low-confidence sample and does not participate in baseline updates.
[0098] The process window W takes the adaptive interval [0.9T, 1.1T] of the current stage cycle time T. When mapping control commands to roles, the minimum time delay principle is adopted and the command is limited to fall within μ±3σ of the expected arrival time. The upper limit of σ for fast loop is recommended to be 5ms, and the upper limit of σ for slow loop is 0.5s. If the protocol semantics are missing, weak semantic templates are only included when the periodic support of the five-tuple reaches ≥0.8. When the baseline is included in the standard operating features, the sliding window size M can be 30 consecutive process cycles. The support threshold of sequence fragments or command semantics is set to ≥0.8 (i.e., it appears in at least 24 cycles in the most recent 30 cycles), the coefficient of variation (CV) threshold of the time interval is set to ≤0.25, the 95th percentile p95 of the round-trip delay must not exceed 1.3 times the historical baseline p95, and the average number of retries must not exceed the historical average +3. The stability of the peer of the interaction object is considered to be ≥0.9 as the passing grade. For infrequent but process-essential rare operations, if the support level of the shift / week dimension is ≥0.1 and there is a matching change record in the operation and maintenance log, they are included as "extended baselines". During statistics, a more lenient threshold of CV≤0.5 and p95≤1.5 times is used, and they are not included in the regular bias judgment. When scoring baseline matching online, the instruction coverage threshold can be defined as ≥0.9 (the proportion of the observed instruction set covering the standard set), and the sequence similarity threshold as ≥0.8 (based on normalized edit distance or longest common subsequence similarity). Temporal consistency requires that the proportion of newly observed adjacent instruction intervals falling within the baseline p5–p95 interval is ≥0.85. If any dimension falls below the threshold for two consecutive process cycles, the asset-phase-role triplet is triggered for review.
[0099] Concept drift detection can be triggered by a combination of distribution distance and statistical tests: if the KS distance between the most recent 10 cycles and the historical baseline in the time interval distribution is ≥0.3 and the KS test p-value is <0.01, or if the semantic support drops continuously from ≥0.8 to ≤0.6 for more than 5 cycles, then the baseline upgrade observation period begins; during the observation period, both the old and new templates are tolerated in parallel until the support of the new template recovers to ≥0.8 and CV≤0.25, at which point the switch is completed. The above values are a set of practical references. In actual deployment, the values can be automatically scaled according to the speed of the process cycle. For example, when T<100ms, the time-related threshold is tightened proportionally to 0.5 of the original value, and when T>10min, it is widened proportionally to 2.0 of the original value. At the same time, historical percentiles are used for adaptive adjustment (e.g., the support threshold is set to max{0.8,P70}, and the p95 relaxation coefficient is set to min{1.3, P90 / P50}) to ensure stability and interpretability under different production lines and load fluctuations.
[0100] Step S30: Identify assets in the industrial internet system as multiple evaluation nodes, determine the protection capability based on the configuration compliance and boundary protection capability of each evaluation node, determine the detection capability based on the probe deployment and traffic coverage of each evaluation node, and determine the response capability based on the closed-loop processing time of each evaluation node.
[0101] Step S31: For each evaluation node, extract the security configuration parameters and boundary access control policies of the industrial control system under the jurisdiction of the evaluation node, and determine the protection capability of the evaluation node by the degree of matching between the security configuration parameters and the compliance baseline and the degree of isolation of cross-security domain access by the boundary access control policies.
[0102] Step S32: Analyze the coverage and protocol parsing depth of the probes deployed in the network topology of the evaluation node, collect the traffic coverage of the network interface under the jurisdiction of the evaluation node, and determine the detection capability of the evaluation node by the correlation and fusion of the probe coverage, protocol parsing depth and traffic coverage.
[0103] Step S33: Analyze the time consumption of each handling step in the security incident handling process of the assessment node, and determine the response capability of the assessment node by the ratio of the preset baseline period to the sum of the time consumption of each handling step, wherein the missing handling steps are included in the sum of the preset maximum time consumption.
[0104] In one embodiment, step S30 is executed uniformly by the security assessment platform after the process timeline is aligned. The platform pulls data from multiple sources, including asset and configuration management, network and security device policies, bypass acquisition and counters, alarm and work order systems, and merges them according to the unique asset identifier and network identity to determine the assessment node set. Nodes are based on a single industrial asset as the basic unit, and the cluster is aggregated by service instance. Assets with multiple network ports are merged into the same node while retaining interface attributes, forming a node base table with metadata such as device type, security domain, and topology, which is used for subsequent capability calculations.
[0105] In one embodiment, step S31 extracts the security configuration parameter set and boundary access control policy set of each node, compares them item by item with the preset compliance baseline to obtain the configuration compliance level, maps the boundary policy to an access control graph to identify unauthorized cross-domain paths, calculates the boundary isolation strength, and multiplies the two to obtain the protection capability value of the node.
[0106] Step S32 analyzes the location and type of deployed probes based on the topology and collection list, evaluates the probe coverage and protocol parsing depth by combining the necessary paths of key business flows, and corrects it by the traffic coverage of mirror or TAP. The three are then fused to obtain the detection capability value.
[0107] Step S33 calculates the time spent on three stages: SIEM alarm, operation and maintenance work order and change / issue log extraction, detection and confirmation, analysis and location, and blocking and recovery. The sum of these three stages is the closed-loop processing time, which is then compared with the preset baseline period to obtain the response capability value. Missing stages are substituted with the preset maximum time.
[0108] The aforementioned capability values are updated on a rolling basis according to process stages and time windows. They can be combined with data quality and smoothing strategies for reliability correction and fluctuation suppression, so as to seamlessly connect with subsequent steps of coupled evaluation and resilience calculation.
[0109] Furthermore, the protection capability is determined based on the configuration compliance and boundary protection capabilities of each assessment node. The method is as follows:
[0110] Step S311: For each evaluation node, extract the security configuration parameter set and boundary access control policy set of the industrial control system under the jurisdiction of the evaluation node;
[0111] Step S312: Compare the security configuration parameter set with the preset compliance baseline library to mark deviation items, and map the boundary access control policy set into an access control matrix;
[0112] Step S313: By traversing the access path entries in the access control matrix, identify unauthorized access paths across security domains to calculate policy coverage completeness;
[0113] Step S314: The reverse normalized value of the deviation term is used to characterize the configuration compliance level, and the policy coverage integrity is used to characterize the boundary isolation strength. The configuration compliance level and the boundary isolation strength are multiplied, and the result of the multiplication is used as the protection capability of the evaluation node.
[0114] In one specific embodiment, the process begins with step S311, which involves the automated collection of data from each designated evaluation node. The system first establishes a communication session with the target device using lightweight acquisition agents deployed in key locations or by leveraging standard network management protocols commonly supported in industrial environments, such as Simple Network Management Protocol (SNMP) and Network Configuration Protocol (NetConf), as well as dedicated application programming interfaces (APIs) provided by various equipment manufacturers. Based on this, the system performs deep information extraction on industrial hosts, servers, network devices (such as switches and routers), and security devices (such as industrial firewalls) within the evaluation node's jurisdiction, forming two core datasets: a security configuration parameter set and a boundary access control policy set. For the extraction of the security configuration parameter set, the system not only obtains the operating system version, the list of installed patches and the status of missing critical patches, but also captures detailed user account policies, including minimum password length, complexity requirements, update cycles, the number of historical passwords remembered, and account lockout thresholds.
[0115] Simultaneously, by executing remote commands or parsing device status information, the system obtains a list of currently running background services, open network ports (TCP / UDP), and their associated applications. For endpoint security software (such as antivirus and host intrusion detection systems), it collects the version number of the virus definition database, the last update time, and the completion time and results of the most recent full system scan. All collected parameters are parsed and stored in a structured database for subsequent comparison. For extracting boundary access control policy sets, the system downloads and fully parses the access control lists (ACLs) or security policy rule bases on industrial firewalls or Layer 3 switches. Each rule is broken down into standardized atomic elements, including source IP address / address range, destination IP address / address range, protocol type (e.g., TCP, UDP, ICAMP), source port number / range, destination port number / range, and explicit allow (Permit / Allow) or deny (Deny / Reject) actions, thus laying the foundation for building a network reachability model.
[0116] After data collection is complete, the process proceeds to step S312, the parameter comparison and data modeling stage. The security configuration parameter set extracted in the previous step is automatically compared item by item with a pre-built compliance baseline library based on industry best practices. This compliance baseline library is a structured knowledge base whose content originates from in-depth interpretation and paradigmatic application of industrial control system security standards such as IEC 62443 and NIST SP 800-82, as well as the company's own security management specifications. The comparison engine checks each collected configuration parameter to ensure it meets the baseline requirements. For example, if the baseline requires a password length of at least 12 bits, but the actual configuration uses 8 bits, this is marked as a deviation. Each deviation is not only recorded but also categorized according to its potential risk level (e.g., high, medium, low), determined by the ease of exploitation and the potential business impact. Simultaneously, the system performs syntax parsing and logical normalization on the collected boundary access control policy set. Because different manufacturers use different syntaxes for their equipment policies, the system first converts these into an intermediate language representation and then maps them to a standardized Access Control Matrix (ACM). The rows and columns of this matrix represent pre-defined logical security domains or critical asset groups within the network (e.g., L1 control layer, L2 monitoring layer, L3.5 industrial DMZ, etc., based on the Purdue model). Cell M(i,j) in the matrix details the set of all communication rules allowed from source region / asset group i to destination region / asset group j, typically stored as a network quintuple (source IP, destination IP, protocol, source port, destination port). This matrix provides a structured and directly computable data foundation for subsequent network access path analysis.
[0117] Subsequently, step S313 is executed, performing an effectiveness analysis of the security policy based on the constructed access control matrix. The core of this step is to utilize graph theory algorithms, such as Depth-First Search (DFS) or Breadth-First Search (BFS), to systematically traverse all possible access path entries in the access control matrix. The network is modeled as a directed graph, where security domains or asset groups are nodes, and the allowed communication rules in the access control matrix are edges. The algorithm starts from a specified starting node (e.g., an enterprise information network area) and recursively or iteratively explores all reachable paths to other nodes. During this process, the system focuses on identifying and statistically analyzing unauthorized access paths that cross different security domains. Here, "unauthorized" does not refer to device-level configuration errors, but specifically to communication paths that, although explicitly allowed in the rules of a specific device (e.g., a firewall), violate higher-level macro-level security domain isolation principles defined by the security architecture design. A typical example is that the Purdue model strictly prohibits direct access from the L4 enterprise information layer to the L1 field control layer, but if a firewall configuration rule allows this path, the path is identified as an unauthorized access path. Ultimately, the policy coverage completeness metric was precisely quantified, and its calculation formula is: Policy Coverage Completeness = 1 - (Number of unauthorized access paths actually discovered / Total number of critical cross-domain paths that should be prohibited according to the security architecture design). The denominator is the total number of paths that must be isolated according to the predefined security policy, and the numerator is the number of illegal paths discovered through actual analysis.
[0118] Finally, in step S314, the protection capability of the evaluation node is comprehensively quantified. This calculation integrates the evaluation results of two dimensions: internal configuration security and external boundary control. First, taking the quantification of intrinsic security as an example, the configuration compliance level is calculated based on all deviation items marked in step S312. Each deviation item is assigned a preset risk weight according to its risk level (e.g., high risk = 0.8, medium risk = 0.4, low risk = 0.1). The total risk score is obtained by weighted summation, and then converted into a value between 0 and 1 by reverse normalization (e.g., compliance = 1 - (total risk score / maximum possible risk score)) to represent the configuration compliance level. Second, taking the quantification of external boundary protection as an example, the system directly uses the policy coverage integrity calculated in step S313 as the measure of boundary isolation strength. This invention innovatively adopts a method of multiplying the configuration compliance level and the boundary isolation strength, and uses the result as the final protection capability score of the evaluation node. The technical logic behind using product operations lies in its ability to accurately represent the interdependent and indispensable logical relationship between these two capability components; both are necessary conditions for effective protection. Mathematically, the product model ensures that when the evaluation value of either component (whether it's configuration compliance or boundary isolation strength) is extremely low (approaching 0), the product result—the final protection capability score—will also approach 0. This method fundamentally differs from the traditional linear weighted model, which may mask a serious deficiency in another capability component due to a high score in one (e.g., 0.91 + 0.1 × 0.1 = 0.91, yet the total score remains high). The product model (1 × 0.1 = 0.1) does not produce this compensatory effect. Therefore, this calculation method, through quantitative means, profoundly reflects the reality that the overall protection level is limited by its weakest link, thus generating a more objective, rigorous, and risk-sensitive assessment result.
[0119] The detection capability is determined based on the probe deployment and traffic coverage of each evaluation node, using the following method:
[0120] Step S321: For each evaluation node, analyze the deployment location and probe type of the security probes deployed in its network topology, determine the probe coverage based on the matching relationship between the deployment location and the critical traffic path, and determine the protocol parsing depth based on the protocol parsing capability corresponding to the probe type.
[0121] Step S322: Collect the traffic mirroring configuration of the network interface under the jurisdiction of the evaluation node, and obtain the traffic coverage by comparing the coverage of the mirrored port with the interaction path of the full service traffic.
[0122] Step S323: The effective detection coverage is characterized by the product of the probe coverage range and the protocol parsing depth, and the product of the effective detection coverage and the traffic coverage is used as the detection capability of the evaluation node.
[0123] In a specific embodiment, the process of quantifying the detection capabilities of each evaluation node is detailed below. This process begins with step S321, which involves in-depth analysis of the security probes deployed in the network topology of each evaluation node to determine their probe coverage and protocol parsing depth. The system first uses previously collected asset fingerprint data and network device configuration information to construct a detailed Layer 2 and Layer 3 network topology diagram, including physical connections and logical VLAN divisions. Subsequently, based on the generated business security baseline, the system identifies the critical traffic paths associated with the evaluation node. These paths do not refer to all network pathways in general, but specifically to specific communication links carrying core production control commands, critical parameter issuance, or important status reporting, such as the control and monitoring traffic path between PLCs and HMIs, or the data interaction path between SCADA servers and field RTUs. The system precisely maps the probe deployment locations, such as mirror ports on specific switches or TAP devices connected in series in the links, onto the topology diagram and compares each location to determine whether it can effectively intercept data packets on the aforementioned critical traffic paths. The probe coverage is quantified as a normalized value, calculated as the ratio of the number of paths effectively monitored by at least one probe in the critical traffic pathways associated with the evaluation node to the total number of critical paths. Simultaneously, the system queries a pre-built probe capability library to determine the protocol parsing depth based on the probe's model, software version, and other information. This capability library details the types of industrial protocols that different probes can parse, such as S7, Modbus / TCP, EtherNet / IP, and OPC UA, as well as the granularity of the parsing, such as whether it can only identify the protocol type or delve into semantic information like function codes, register addresses, object IDs, and variable names. Protocol parsing depth is also assigned a normalized score; for example, a probe supporting only 5-tuple analysis has a depth of 0.2, one supporting general industrial protocol identification has 0.6, and one capable of deep semantic parsing has 1.0.
[0124] Next, the system executes step S322, which aims to obtain an objective traffic coverage rate by collecting and analyzing the traffic mirroring configuration of the network interfaces under the evaluation node. This step aims to address the gap between theoretical probe coverage and actual data collection quality, especially considering potential congestion, packet loss, or incomplete configuration issues at the mirroring ports of production network switches. The system automatically retrieves detailed configurations of traffic mirroring sessions from the relevant network switches via SNMP, NetConf, or device command-line interface scripts. This includes all physical ports, VLANs, or port channels configured as mirroring sources, as well as the mirroring destination ports. The system compares these mirroring sources with the full range of service traffic interaction paths extracted from the service security baseline for the evaluation node—that is, all compliant and expected communication sessions. Traffic coverage is precisely calculated as the proportion of traffic whose paths pass through interfaces fully included in the mirroring source configuration. To further improve the accuracy of the evaluation, the system also checks for oversubscription in the mirroring sessions, i.e., the total bandwidth of all source ports exceeding the processing capacity of the destination port. Simultaneously, by combining the destination port output packet loss data read from the switch interface counter via SNMP, the calculated coverage rate is dynamically penalized and corrected. If significant packet loss is detected, the original coverage rate is reduced by the packet loss rate, thus accurately reflecting the integrity of the data actually received by the probe.
[0125] Finally, the method proceeds to step S323, where the quantitative indicators obtained in the first two steps are fused to calculate the final detection capability. In a preferred embodiment of the invention, the probe coverage area and the protocol parsing depth are first multiplied to obtain an intermediate indicator called "effective detection coverage." The logic of using multiplication is that it accurately reflects the complementary relationship between the "breadth" of coverage and the "depth" of parsing; the effectiveness of a probe deployment that covers all critical paths but can only perform shallow parsing, or a deployment that can perform deep parsing but only covers a few paths, will be limited, and the multiplication operation can appropriately reflect this "weakest link effect." Subsequently, this effective detection coverage is multiplied again with the corrected traffic coverage obtained in step S322, and the result is used as the final detection capability of the evaluation node. This final product operation is crucial because it ensures the accuracy of the evaluation results: even if the probe deployment location and type are theoretically perfect, resulting in an effective detection coverage of 1.0, if the actual traffic mirror link has serious problems, such as a traffic coverage of only 0.1, the final detection capability will be correspondingly reduced to an extremely low level of 0.1. This calculation method avoids the problem that can occur with traditional weighted average models, where a high score in one aspect can mask fatal flaws in other aspects, thus generating a highly sensitive and objectively rigorous quantitative evaluation result that is highly sensitive to visibility risks.
[0126] Response capability is determined based on the closed-loop processing time of each assessment node, using the following method:
[0127] Step S331: For each evaluation node, analyze the closed-loop cycle of its security incident handling process from the triggering of the detection event to the completion of the handling action, and break down the closed-loop cycle into detection confirmation time, analysis and location time and blocking and recovery time.
[0128] Step S332: Collect the event timestamps of the security devices and operation and maintenance systems under the jurisdiction of the evaluation node, extract the detection confirmation time, the analysis and location time and the blocking and recovery time respectively by the difference of the event timestamps, and use the sum of the detection confirmation time, the analysis and location time and the blocking and recovery time as the total closed-loop time of the handling.
[0129] Step S333: The response capability of the evaluation node is characterized by the ratio of the preset baseline period to the total closed-loop processing time.
[0130] If any processing step is missing in the evaluation node, the time of that processing step is set to the preset maximum threshold time and added to the accumulation.
[0131] In a specific embodiment, the quantification of the response capability of each evaluation node begins in step S331 by standardizing and analyzing the security incident handling process associated with each evaluation node. This analysis, based on a predefined, industry-practice-compliant incident response lifecycle model, meticulously breaks down the entire closed-loop cycle from initial incident detection to final handling into three core, quantifiable stages:
[0132] The timeframes are: detection and confirmation time, analysis and location time, and blocking and recovery time. Detection and confirmation time covers the entire process from the moment a security device (such as an intrusion detection system or industrial firewall) generates an initial alarm to the moment that alarm is confirmed as a genuine security event requiring follow-up. Analysis and location time follows immediately, recording the total time from the moment the event is confirmed until, through correlation analysis of asset information, network traffic, and endpoint logs, the affected assessment node is accurately located, the attack path is identified, and an effective remediation plan is developed. Finally, blocking and recovery time measures the time span from the approval and implementation of the remediation plan until the threat is cut off by applying blocking policies on network boundary devices or performing isolation actions on endpoints, and finally verifying that services have returned to the normal operating state defined by the business security baseline.
[0133] In step S322, event timestamps recorded by various security devices and operation and maintenance systems related to the evaluation node are collected through an automated interface to accurately calculate the actual time consumption of the three handling stages mentioned above. To ensure the accuracy of the calculation, all collected timestamps are calibrated and normalized using a unified clock source. Specifically, the application programming interface extracts the state transition time points associated with a specific event ID from the security information and event management platform, the security orchestration automation and response system, and the IT service management ticket system. For example, the detection confirmation time is calculated by obtaining the difference between the time of the first alarm occurrence and the time of ticket creation or status change to "confirmed"; the analysis and location time is obtained by calculating the difference between the timestamps of the ticket's "confirmed" status and the "analysis completed" or "pending handling" status; and the blocking and recovery time is calculated by using the difference between the time of the handling instruction issuance and the timestamp of the firewall policy effective log or the endpoint security software isolation success log. The sum of these three independently calculated time periods constitutes the total closed-loop handling time of the event.
[0134] In step S333, the response capability of the evaluation node is ultimately characterized by the ratio of a preset baseline period to the total closed-loop time calculated in the previous step. This baseline period is not a fixed value, but is dynamically set according to the business criticality of the evaluation node, the importance of the security domain it belongs to, and the requirements of the relevant service level agreement. For example, the baseline period for a critical PLC node that directly controls the production process may be set to a few minutes, while the baseline for an auxiliary data acquisition server may be relaxed to several hours. Through this ratio calculation, the shorter the total closed-loop time, the higher the response capability evaluation value, thus intuitively reflecting the response efficiency.
[0135] Furthermore, to address potential process omissions in actual operation, such as the lack of collectable digital timestamps due to reliance on manual operations, a robust penalty mechanism is employed. If any of the three steps in the handling process of an evaluation node is found to be missing during data collection, the time for that missing step will be automatically set to a preset maximum threshold time, and then included in the cumulative calculation of the total handling closure time. This maximum threshold time is typically set to several times the baseline period to ensure that any break in the process or gap in capability will lead to a significant decrease in the response capability score, thus objectively and severely reflecting the serious deficiencies in the completeness of the response process at that evaluation node.
[0136] To further illustrate this, let's consider a specific application scenario. Suppose a critical PLC node controlling a robotic arm on a production line is attacked by an aberrant Modbus write command. First, the PLC node's baseline cycle time is set to 15 minutes (900 seconds) due to its high criticality. During the incident handling process, at 10:00:05, the intrusion detection system generates an alarm; at 10:02:35, the security operations personnel mark the event status as "confirmed" in the work order system, at which point the calculated confirmation time is 150 seconds.
[0137] Subsequently, after analysis, at 10:12:05, the source of the attack was determined to be an infected engineering workstation, and an isolation plan was formulated. The work order status was updated to "Pending Handling," thus calculating the analysis and location time as 570 seconds. The handling instruction was issued through the automation platform at 10:12:20, and the firewall recorded the blocking policy taking effect at 10:12:50. Finally, at 10:14:50, monitoring data showed that the robotic arm had resumed normal production rhythm, thus determining the blocking recovery time as 150 seconds. Adding these three times together, the total closed-loop handling time is 150 + 570 + 150 = 870 seconds.
[0138] Ultimately, the response capability evaluation value for this PLC node is the ratio of the baseline period to the total handling closure time, i.e., 900 / 870, approximately equal to 1.034. This evaluation value, higher than 1, indicates that the efficiency of this response exceeded the preset standard. Conversely, if in another incident, the analysis and location stage, relying on offline communication, failed to leave a "pending handling" timestamp in the work order system, resulting in missing data collection, the penalty mechanism would be triggered. Assuming the preset maximum threshold time is set to three times the baseline period, i.e., 2700 seconds, then the missing analysis and location time would be counted as 2700 seconds. The total handling closure time would be recalculated as 150 + 2700 + 150 = 3000 seconds. The final response capability evaluation value would then become 900 / 3000 = 0.3. The evaluation value plummeted from 1.034 to 0.3, a significant change that accurately quantifies and exposes serious deficiencies in the handling process regarding record completeness and standardized execution.
[0139] Step S40: Construct a security capability coupling assessment model based on a time-series causal graph. Define the protection capability, detection capability, and response capability of each assessment node as interacting nonlinear system nodes. Characterize the nonlinear influence weight relationship of each security capability among different nonlinear system nodes through a collaborative influence matrix. Solve the steady-state recovery time of the multiple assessment nodes under virtual disturbances based on the business security baseline and the collaborative influence matrix to obtain the resilience coefficient.
[0140] In this embodiment, the nonlinear influence weight relationship of each security capability among different nonlinear system nodes is characterized by a collaborative influence matrix. The method is as follows:
[0141] Step S41: Using multiple evaluation nodes as row and column indices, and protection capability, detection capability, and response capability as capability components of each evaluation node, construct a collaborative influence matrix. Each element in the collaborative influence matrix corresponds to the weighting coefficient between the source capability component of the row index evaluation node and the target capability component of the column index evaluation node.
[0142] Step S42: Based on the time sequence of dependence of each stage in the security incident attack chain on protection capability, detection capability, and response capability, assign initial values to the weight coefficients of protection capability on detection capability, detection capability on response capability, and response capability on protection capability within the same evaluation node.
[0143] Step S43: Based on the frequency and direction of business data flow interaction between different evaluation nodes, assign initial values to the weight coefficients of the same capability components between different evaluation nodes; fill the initial values into the corresponding element positions of the collaborative influence matrix to form an initial collaborative influence matrix.
[0144] Based on the business security baseline and the collaborative impact matrix, the steady-state recovery time of the multiple evaluation nodes under virtual disturbances is calculated to obtain the resilience coefficient. The method is as follows:
[0145] Step S44: Apply virtual disturbance events to each evaluation node in the time-series cause-effect graph. The virtual disturbance events trigger the initial state value of the protection capability, detection capability, or response capability of the specified evaluation node to deviate from the standard state value defined by the business security baseline.
[0146] Step S45: Using the influence relationship recorded in the collaborative influence matrix as an iterative constraint, update the state value of each capability component of each evaluation node round by round within the discrete time step until the state value of each capability component of all evaluation nodes converges to a deviation from the business security baseline of less than a preset threshold.
[0147] Step S46: Count the total number of time steps from the application of the virtual disturbance event to the convergence of the state value, and use the reciprocal or normalized value of the total number of time steps as the resilience coefficient.
[0148] In step S40 of this embodiment of the invention, a security capability coupling evaluation model based on a time-series causal graph is constructed. This model defines the protection capability, detection capability and response capability of each evaluation node in the industrial Internet system as interacting nonlinear system nodes. The model accurately describes the nonlinear influence weight relationship between different nodes and between different capability components through a collaborative influence matrix. Finally, based on the business security baseline and the collaborative influence matrix, the steady-state recovery time of each evaluation node under virtual disturbance is solved to obtain the resilience coefficient.
[0149] In practical implementation, step S40 is executed by the security capability coupling assessment module based on a unified process timeline alignment. This module first obtains the basic quantitative values of the protection capability, detection capability, and response capability of each assessment node from step S30, and obtains the dynamically updated business security baseline with the process cycle from step S20, serving as input for subsequent modeling and calculation. To construct the synergistic influence matrix, the module uses a pre-defined set of assessment nodes as row and column indices, and uses protection capability, detection capability, and response capability as the three capability components of each assessment node, constructing a synergistic influence matrix with dimensions of (number of nodes × 3) × (number of nodes × 3). Each element in this matrix... Indicates the source node The Class capability components ( ) for target node The Class capability components ( The influence weighting coefficient is a non-negative real number used to describe the strength and direction of the influence of the state change of the former on the state change of the latter.
[0150] When initializing the collaborative impact matrix, the module assigns initial values to the weights of different capability components within the same evaluation node according to the temporal dependencies of the security event attack chain. Specifically, based on the classic attack chain model, the attack first attempts to breach the protection measures, so changes in protection capabilities directly affect the difficulty of subsequent detection capabilities in discovering threats; that is, protection capabilities have a positive impact on detection capabilities. After the detection capability discovers a threat, the activation of the response capability depends on the input of the detection result; therefore, detection capabilities have a positive impact on response capabilities. After the response capability performs its actions, through policy updates, patch repairs, isolation recovery, etc., it will have a feedback effect on the protection capability, improving or restoring it; therefore, response capabilities have a positive impact on protection capabilities. Based on the above logic, the module assigns initial weight coefficients to the three directed edges "protection → detection", "detection → response", and "response → protection" within each evaluation node. The initial values can be set based on historical experience or expert knowledge; for example, in the absence of prior information, they can all be set to 0.5, representing a moderate positive impact.
[0151] Regarding cross-node impact, the module initializes the weighting coefficients of the same capability components between different nodes based on the frequency and direction of business data flow interactions between them. The module first extracts a standardized business data flow interaction matrix between each assessment node from the business security baseline. This matrix records the node's performance under normal operating conditions. To the node The frequency and direction of sent control commands, status data, or process parameters. For two nodes with high-frequency business data flow interaction, the module assumes a coupling relationship between their security capabilities. In particular, changes in the upstream node's protection or detection capabilities can affect the input data quality or threat exposure of the downstream node, thus impacting the status of the same capability component in the downstream node. For example, if a node... To the node Frequent issuance of control commands will cause the node to... If the protection capabilities are flawed, malicious commands may be injected, thereby increasing the risk to nodes. The detection pressure, therefore the node The protective capabilities of the nodes The detection capability is also affected. To simplify the model and highlight the main coupling relationships, this embodiment prioritizes establishing influence relationships for the same capability components between different evaluation nodes, namely "protection → protection", "detection → detection", and "response → response". The values are normalized based on the frequency of business data flow interactions; the higher the interaction frequency, the larger the initial weight. For example, the normalized interaction frequency can be directly used as the initial weight of the corresponding influence edge. If there is no direct business data flow interaction between nodes, the corresponding weight is set to 0, indicating no direct influence. After completing the above initialization, the module stores all the filled weight coefficients into the collaborative influence matrix to form the initial collaborative influence matrix.
[0152] After obtaining the initial collaborative impact matrix, the module enters the resilience coefficient calculation phase. The resilience coefficient is used to quantify the ability of an evaluation node to recover to the standard state defined by the business security baseline after being subjected to a disturbance. Its core is to obtain the steady-state recovery time by simulating virtual disturbance events and iteratively calculating the dynamic evolution of the capability components. Specifically, the module first applies virtual disturbance events to one or more evaluation nodes based on the current state of each node in the time-series causal graph. These virtual disturbance events can be predefined typical threat scenarios, such as implementing a momentary degradation of the protection capability of a critical PLC node, causing its configuration compliance level to drop from the baseline value to a certain low value, or temporarily reducing the detection capability of a node to zero due to probe failure. The virtual disturbance is applied by selecting a target node. A certain component of ability , its state value Standard state values defined from the business security baseline Forced adjustment to disturbance state value This disturbance value usually deviates significantly from the baseline value, for example, by taking 20% of the baseline value or setting it directly to 0, in order to simulate severe failure scenarios.
[0153] After applying a virtual perturbation, the module uses the influence relationships recorded in the collaborative influence matrix as iterative constraints to update the state values of each capability component of all evaluation nodes round by round within a discrete time step. Each discrete time step corresponds to a process cycle or an integer multiple of a process cycle to ensure that the update rhythm is consistent with the actual operating rhythm of the industrial site. In each iteration, for any node... Arbitrary ability components The state value of its next time step From the current state value This is determined together with all the influence terms pointing to this component. A specific update model can employ nonlinear dynamic equations, such as:
[0154] ;
[0155] in, It is a non-linear activation function (such as the Sigmoid function) used to constrain the state value within a reasonable range (e.g., [0,1]). The self-holding coefficient represents the inertia or memory effect of the capability component itself. These are the corresponding weight coefficients in the synergistic influence matrix; As the influence function, it can usually be a linear function or a nonlinear function with a threshold, used to map the state value of the source capability component into a driving force on the target capability component; This is an external recovery item used to simulate the proactive recovery effect brought about by operational intervention or system self-healing. During the iteration process, the module continuously monitors the state values of each capability component of all evaluation nodes and calculates the deviation between them and the standard state values defined by the business security baseline. When the deviation of all capability components of all nodes from the baseline is less than a preset threshold (e.g., the maximum allowable deviation is 5% of the baseline value), the system is considered to have converged to a steady state.
[0156] The module starts timing from the application of the virtual disturbance event and records the total number of time steps taken until the state value converges to within the deviation threshold. The total number of time steps reflects the time required for the system to recover from instability under a specific disturbance; the shorter the recovery time, the stronger the system's resilience. To obtain a dimensionless and easily comparable resilience coefficient... The module normalizes the reciprocal of the total number of time steps, for example, by defining... ,in The preset baseline recovery time step can be an industry standard requirement or a historical average recovery time. A toughness coefficient greater than or equal to 1 indicates a recovery rate better than the baseline, while a coefficient less than 1 indicates a recovery rate less than 1. Another common normalization method is to map the toughness coefficient to the [0,1] interval, for example... ,in This is a time constant, which can be set according to the process characteristics. Regardless of the method used, the toughness coefficient is ultimately output in numerical form and used in the subsequent step S60 to fuse with the dynamic weights to generate a comprehensive evaluation result.
[0157] To make the above process more concrete and easier to implement in engineering, this embodiment provides a simplified example. Assume there are two evaluation nodes, A and B, where node A is a PLC and node B is an HMI. Each node has three capability components: Protection (P), Detection (D), and Response (R). In the initial collaborative influence matrix, the weights of "Protection → Detection" within node A are 0.5, "Detection → Response" are 0.5, and "Response → Protection" are 0.5; the same applies to node B. Since node A periodically reports process data to node B, resulting in a high frequency of interaction, the weights of "Protection → Protection" from node A to node B are set to 0.3, "Detection → Detection" from node A to node B are 0.3, and "Response → Response" from node A to node B are 0.2. Other cross-node weights are 0. In the business security baseline, the standard state value of each capability component of nodes A and B is 0.8. Now, a virtual perturbation is applied to the protection capability of node A, forcibly reducing its state value from 0.8 to 0.2. In subsequent iterations, the decrease in node A's protection capability gradually reduces its detection capability through internal "protection → detection" edges, and negatively impacts node B's protection capability through cross-node "protection → protection" edges, triggering a series of chain reactions. Simultaneously, the module simulates the intervention and repair process by setting self-holding coefficients and external recovery terms, causing each capability component to gradually revert to its baseline value. After multiple iterations, iteration stops when the deviation of all capability components from the baseline value is less than 0.05, and the number of iterations is recorded as 5 time steps. If the preset baseline recovery time step is 4, then the resilience coefficient... This indicates that the node's recovery capability under disturbance is slightly lower than the benchmark, and its robustness in protection needs to be monitored. Through the above methods, step S40 transforms the originally static and isolated security capability indicators into resilience quantification indicators that reflect nonlinear coupling and dynamic recovery processes, laying a solid model foundation for subsequent dynamic evaluation and adaptive adjustment.
[0158] Step S50: Align the time window with the process cycle of the industrial scenario, determine the change event according to the operation and maintenance log, analyze the causal lag time between the change event and the alarm event in the historical data through the causal discovery algorithm, and dynamically adjust the weight coefficients in the collaborative influence matrix according to the causal backtracking results.
[0159] In this embodiment, the causal lag time between change events and alarm events in historical data is analyzed using a causal discovery algorithm. The weight coefficients in the collaborative influence matrix are dynamically adjusted based on the causal backtracking results. The method is as follows:
[0160] Step S51: Collect operation logs and alarm logs of security devices within the historical time window, extract change event sequences from the operation logs, and extract alarm event sequences from the alarm logs;
[0161] Step S52: Using the sequence of change events as the dependent variable sequence and the sequence of alarm events as the resultant variable sequence, calculate the causal lag time between each change event and each alarm event using a causal discovery algorithm.
[0162] Step S53: Using the change event and alarm event pairs with causal lag time less than a preset threshold as valid causal pairs, count the number of valid causal pairs and the average causal lag time within the scope of each evaluation node, and generate causal backtracking results.
[0163] Step S54: Dynamically adjust the weight coefficients in the collaborative influence matrix according to the causal backtracking results. The dynamic adjustment takes the number of effective causal pairs and the average causal lag time as inputs, and maps the inputs to the correction amount of the weight coefficients through a preset mapping rule. The weight coefficients of the corresponding evaluation nodes in the collaborative influence matrix are updated with the correction amount.
[0164] In the specific implementation of step S50, the safety capability coupling assessment model introduces a causal discovery mechanism, utilizing the temporal correlation between historical operation and maintenance changes and alarm data to dynamically adjust the weight coefficients in the collaborative influence matrix. This enables the assessment model to adapt to process changes and risk drift in industrial scenarios, improving the real-time performance and accuracy of the assessment. This process is executed by the causal backtracking and weight adaptation module on the basis of unified process timeline alignment, employing a rolling update strategy of "daily batch calculation + small incremental steps within stages" to ensure that weight adjustments are both stable and sensitive.
[0165] First, the module collects operation and maintenance logs and safety equipment alarm logs covering multiple complete process cycles within historical time windows. A window length of 4 to 12 weeks is typically recommended to cover typical scenarios such as shift fluctuations, weekends, planned maintenance, and batch switching. To ensure the time-series comparability of data from different sources, the module employs a dual-track time synchronization mechanism combining edge and center methods before collection. PTP or NTP is used to normalize the timestamps of each system to a unified time base. Records with fixed offsets or drifts are corrected, and the original and corrected timestamps, as well as the correction factor, are retained. Records that cannot be reliably aligned to the process timeline are labeled as low-confidence or directly removed. From the operation and maintenance logs, the module extracts standardized change event sequences by parsing change orders, configuration distribution records, patch upgrade logs, and policy change audits. Each change event includes at least the change ID, change type (e.g., policy adjustment, signature update, script optimization, patch installation, topology change), target (mapped to a specific assessment node or asset group), operator, planned and actual effective time window, version summary or configuration hash, scope of impact, and rollback flag. From the alarm logs, the module aggregates alarm records from firewalls, IDS / IPS, industrial protocol anomaly detection, and endpoint security systems. After deduplication and severity normalization, a unified alarm event sequence is formed. Each alarm event includes at least the event ID, source device, network 5-tuple, industrial protocol and object, rule ID, severity, credibility, first and last timestamps, and handling summary.
[0166] Subsequently, the module maps the extracted change events and alarm events to a predefined set of evaluation nodes based on asset identifiers and network topology. It further associates these events with capability dimensions (protection, detection, and response). For example, firewall policy changes primarily affect protection capabilities, IDS signature updates primarily affect detection capabilities, and SOAR playbook optimizations primarily affect response capabilities. For composite changes affecting multiple capabilities simultaneously, the impact is distributed according to a preset ratio, and the mapping reliability is recorded. After completing the event mapping, the module uses the industrial process cycle as the basic unit for time alignment. Each process cycle (such as a production takt time, a batch, or a recipe cycle) is treated as a discrete time step. Change events and alarm events are resampled according to the process cycle in which their occurrence time falls, forming two time series for each evaluation node on each capability dimension: a change trigger sequence (usually using 0 / 1 markers or counts to indicate whether a change occurred and its intensity within that time step) and an alarm result sequence (using severity-weighted intensity or alarm counts). To control for confounding factors, the module also introduces a series of covariates, including process load, shift identifier, weekend and holiday markers, maintenance window, probe coverage, clock synchronization anomaly marker, network packet loss rate, and service security baseline deviation, in order to adjust for causal estimation.
[0167] In the causal lag time estimation phase, the module adopts a strategy of prioritizing intra-node operations and then proceeding cross-node operations as needed, and enhances robustness through cross-validation using multiple causal discovery algorithms. For intra-node influence relationships, the module primarily employs a multivariate process model, treating change events as external stimuli that influence the conditional strength of alarm events, and assuming that this influence decays over time. The model form is as follows: ,in For alarm events at any time Conditional strength, Baseline intensity, To change the time of the event, To influence the kernel function, an exponential decay form is typically used. The initial influence coefficient is obtained through maximum likelihood estimation. and attenuation coefficient This allows us to determine the impact intensity and lag time distribution of each change on subsequent alarms. Indicates time difference.
[0168] To enhance interpretability, the module also employs a generalized linear model with lags, regressing the alarm result sequence against the change event sequence and covariates. The model form is as follows: ,in The alarm intensity at the current time step. Lagging The intensity of change at each time step The lag effect coefficient to be estimated is... As covariates, pass significance tests (e.g.) (Values less than 0.05) identify which lag terms have statistically significant causal effects and extract the causal lag time from them. For the intercept term, For the lagging step size, For the maximum lag step size, For covariate index, Covariate coefficients This is the random error term.
[0169] Regarding cross-node impacts, the module only operates on links with strong business interactions, stable directions, and acceptable data quality, employing instrumental variables or synthetic control methods to mitigate confounding biases caused by changes in common upstream nodes. The upper limit of the causal lag time is dynamically set based on the process cycle time, typically 3 to 5 process cycles; long-term effects exceeding this limit are considered unsuitable for online weight correction.
[0170] After obtaining the causal lag time between each change event and each alarm event, the module selects change-alarm event pairs whose causal lag time is less than a preset threshold (e.g., 3 process cycles) and whose statistical significance meets the standard, as valid causal pairs. For each evaluation node and its corresponding capability dimension combination, the module counts the number of valid causal pairs. And calculate the average causal lag time. (Based on process cycles), the average lag time can be weighted according to the intensity of causal impact to form a comprehensive causal backtracking result. For example, if a node triggers multiple alarms due to policy change events in the detection capability dimension, and the average lag time is 1.5 process cycles, then the backtracking result reflects that the change in the node's protection capability has a significant and rapid impact on the detection capability.
[0171] Based on the causal backtracking results, the module dynamically adjusts the weight coefficients in the synergistic influence matrix. The adjustment process employs a preset mapping rule, mapping the number of effective causal pairs and the average causal lag time to weight correction amounts. Specifically, within the same evaluation node, if there are effective causal pairs between changes in the "protection → detection" direction (such as strategy adjustments) and detection-related alarms, and these pairs are numerous and have short average lag times, it indicates that the impact of protection capabilities on detection capabilities is underestimated, and the corresponding weight coefficients should be increased. Conversely, if alarms do not increase significantly after the change, or even decrease, it may indicate that the original weights are too high and should be appropriately reduced. The mapping rule can use a piecewise linear function or a saturation function, for example, defining a correction amount. ,in All are preset hyperparameters. This represents the maximum single correction magnitude. This is the average causal lag time. For cross-node impacts, if a change in an upstream node triggers an alarm in a downstream node, the weight coefficient of the corresponding upstream node's capability on the same capability of the downstream node is increased. During weight adjustment, the module sets an upper limit (e.g., 0.65) for the incident edges of each target capability component. When the weight of an edge approaches the upper limit, the adjustment step size is automatically reduced to prevent excessive weight concentration. Simultaneously, a dominant coupling strength check is performed on the entire matrix. If the overall coupling is too high, the weights are reduced overall while maintaining relative proportions to ensure model stability. To avoid drastic weight fluctuations, the module uses exponential forgetting to smoothly fuse old and new weights, i.e. ,in The suggested weights are calculated based on the causal backtracking results. This is a smoothing factor (usually set to 0.7-0.9), which automatically increases when data quality is insufficient or sample size is too small. To reduce the update step size, For the new weighting coefficients, These are the old weighting coefficients.
[0172] At the engineering implementation level, the module adopts a rolling mechanism of "daily batch calculation + small incremental steps within a phase". Every natural day or after completing a certain number of process cycles (e.g., every 5 batches), the module updates the causal statistics incrementally. When new valid evidence is added or the cumulative drift exceeds a preset threshold, a weight update is triggered. All update operations are included in version management, recording the weight values before and after the update, causal evidence summaries, the time windows involved, a list of change and alarm IDs, mapping parameters, and stability check results. Auditing and one-click rollback are supported, and a visual interface is provided to display the traceability link of "causal evidence → weight adjustment", which helps security operations personnel understand the model behavior and assist in decision-making. When deployed across multiple production lines or scenarios, the module adaptively adjusts parameters according to the process cycle time and data quality of each production line: for fast-cycle production lines (e.g., cycle time less than 100ms), a stricter lag upper limit (e.g., 2 process cycles) and a smaller correction step size are adopted; for slow-cycle or batch processing scenarios, the time window and evidence accumulation period are appropriately relaxed to ensure the effectiveness of causal discovery and the stability of weight adjustment.
[0173] Through the above implementation method, step S50 realizes data-driven adaptive adjustment of the weights of the collaborative influence matrix, enabling the evaluation model to continuously learn the causal relationship between changes and alarms from actual operating data, dynamically reflecting the real nonlinear coupling relationship between security capabilities in the industrial internet system, thus providing solid model support for step S60 to generate accurate, real-time and interpretable dynamic security capability evaluation results.
[0174] Step S60: Generate and output the dynamic safety capability assessment result based on the adjusted weighting coefficient and the toughness coefficient.
[0175] Step S61: Construct a weight vector using the weight coefficients of each evaluation node in the adjusted collaborative influence matrix, construct a resilience vector using the resilience coefficients of each evaluation node, and fuse the weight vector and the resilience vector element by element to generate a comprehensive evaluation value for each evaluation node.
[0176] Step S62: Using the process stage of the industrial scenario as the dimension, aggregate the comprehensive evaluation values of each evaluation node under the same process stage to obtain the stage evaluation value of the process stage.
[0177] Step S63: Determine the safety level of the process stage based on the degree of deviation between the stage evaluation value and the preset benchmark value of the process stage;
[0178] Step S64: Associate the safety level of each process stage with the corresponding assessment node identifier and output the information to the display interface.
[0179] In the specific implementation of step S60, the security assessment platform will integrate and calculate the dynamically adjusted collaborative influence matrix and resilience coefficient to generate a dynamic security capability assessment result for industrial internet systems. The result will be aggregated and classified according to the process stage, and finally output through a visual interface to provide security operation and maintenance personnel, production management personnel and decision-makers with an intuitive and operable understanding of the security situation.
[0180] First, in step S61, the platform constructs weight vectors and resilience vectors for each evaluation node based on the adjusted collaborative influence matrix, and obtains the comprehensive evaluation value of each node through element-by-element fusion. Specifically, for each evaluation node, the platform extracts the weight information related to that node from the collaborative influence matrix to form a weight vector. Since the collaborative influence matrix records the influence weights of all source capability components on the target capability component, the platform needs to compress the multidimensional influence weights into a single weight coefficient that can reflect the overall importance of the node. In this embodiment, the platform calculates the weight value of each node using eigenvector centrality or the sum of weighted out-degrees. Its weight Defined as the normalized sum of the weights of all capability components pointing to this node, calculated using the following formula: This weight reflects the strength of a node's influence in a security capability coupled network, that is, the criticality of the node in the overall security architecture. Indicates the source node index. Indicates the first 1 node Represents the indexes of all target nodes. Indicates the source node capability component index. This represents the index of the target node's capability components.
[0181] At the same time, the platform obtains the resilience coefficient of each evaluation node from step S40. This coefficient has been calculated in the previous step using the steady-state recovery time under virtual perturbation and normalized to... The range is defined as an interval or higher, with larger values indicating stronger recovery capabilities. Subsequently, the platform element-wise merges the weight vector and resilience vector to generate a comprehensive evaluation value for each evaluation node. The fusion method can be selected from weighted product or weighted summation according to actual needs, for example, product fusion can be used. This method can amplify the advantages of nodes with high weights and strong resilience. However, when the resilience coefficient is low, even nodes with high weights will have their overall evaluation value significantly lowered, reflecting the risk amplification effect of "insufficient capability of critical nodes." A weighted summation method can also be used. ,in The preset balance coefficient can be dynamically adjusted according to the evaluation objectives. The platform stores the calculated comprehensive evaluation value, along with a timestamp and process cycle identifier, for subsequent stage aggregation and trend analysis.
[0182] In step S62, the platform aggregates the comprehensive evaluation values of each evaluation node within the same process stage, using the process stage of the industrial scenario as the dimension, to obtain the stage evaluation value for that process stage. Industrial scenarios are typically divided into multiple process stages according to the production flow, such as raw material preparation, reaction control, finished product separation, and packaging. Each process stage contains a specific set of assets and evaluation nodes. The platform pre-marks the process stage to which each evaluation node belongs in the asset fingerprint data and solidifies this mapping relationship in the business security baseline of step S20. During aggregation, the platform calculates the stage evaluation value based on the comprehensive evaluation values of each node within the current evaluation period, using methods such as arithmetic average, weighted average, or minimum value aggregation. The weighted average method can weight nodes according to their business criticality within the process stage; for example, PLC nodes directly involved in the control loop have a higher weight than HMI nodes that provide auxiliary monitoring. The minimum value aggregation method reflects the "weakest link" effect, using the evaluation value of the weakest node within the stage as the stage evaluation value, suitable for continuous production processes with extremely high safety redundancy requirements. The platform then calculates the stage evaluation value... Record this information and compare it with the preset baseline value for this process stage. The preset baseline value is obtained by statistically analyzing the stage evaluation values during historical normal operation periods. It can be set using a sliding window quantile (such as the median or 75th quantile) or the theoretical optimal value based on the business safety baseline, and serves as a reference point for judging the degree of deviation of the current safety status.
[0183] In step S63, the platform determines the safety level of the process stage based on the degree of deviation between the stage assessment value and the preset benchmark value of the process stage. The degree of deviation can be expressed as an absolute difference or a relative deviation rate, for example, defining a relative deviation rate. ,in For the first Preset baseline values for each process stage This represents the stage evaluation value for the k-th process stage, when... When the phase assessment value is not lower than the benchmark, the safety status is good; when A negative value indicates a security capability below the baseline, with a larger negative value indicating a higher security risk. The platform divides security levels into multiple tiers based on the degree of deviation, for example, setting it to "Excellent" (…). ),"good"( ),"Notice"( ),"Danger"( The platform offers four risk levels, with customizable thresholds based on industry standards or enterprise safety requirements. To enhance the robustness of the risk level classification, a confidence interval concept is introduced when calculating the degree of deviation. When the fluctuation range of the stage assessment value is large or the data quality is low, the risk level boundaries are appropriately relaxed or the risk level results with confidence markers are output. Furthermore, the platform can combine the degree of deviation with historical trends in the process stage. If the degree of deviation continues to worsen over multiple consecutive process cycles, the warning level is automatically raised to indicate potential systemic risks.
[0184] In step S64, the platform associates the safety level of each process stage with the corresponding assessment node identifier and outputs it to the display interface. The display interface adopts the form of a visual dashboard, typically including a process flow diagram overlaid with a heatmap, a time trend curve, a node details panel, and an alarm and suggestion module. On the process flow diagram, the platform displays the safety level of each process stage using color coding, for example, green represents "excellent," yellow represents "good," orange represents "caution," and red represents "danger," enabling production managers to clearly identify safety vulnerabilities in the current production process. When a user clicks on a process stage, the interface automatically expands to show the comprehensive assessment value, resilience coefficient, weight coefficient, and contribution percentage of each assessment node under that stage, and displays the key influence edge of that node in the collaborative influence matrix, helping to pinpoint which specific node or type of capability deficiency caused the stage assessment value to decline. The time trend curve module displays the comprehensive assessment value change curves of each process stage and key node over the past several process cycles, combined with the annotation of change events and alarm events, facilitating the analysis of the causes of safety situation changes. The alarm and suggestion module automatically generates actionable suggestions based on the security level. For example, when the assessment value continuously declines in a certain stage, it suggests prioritizing the investigation of nodes with high weight but low resilience coefficients within that stage, and provides information on possible change records or recent alarm patterns to guide operations personnel in taking targeted optimization measures, such as adding probes, adjusting boundary policies, and optimizing handling processes. All output content supports multi-dimensional filtering and export by time range, process stage, node type, etc., facilitating the generation of assessment reports and the conduct of governance reviews.
[0185] Through the above implementation method, step S60 organically integrates the dynamically adjusted weight coefficient and resilience coefficient to form a hierarchical evaluation result from node to process stage, and outputs it through a visual interface, realizing a complete closed loop of security evaluation from quantitative calculation to business decision-making. This makes the evaluation result not only scientific and accurate, but also intuitive and operable, truly serving the continuous security optimization and resilience improvement of the industrial internet system.
[0186] The terms “first”, “second”, etc., are used to distinguish similar objects, not to describe or indicate a specific order or sequence.
[0187] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent in such process, method, article, or apparatus / device.
[0188] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. An industrial internet security capability assessment framework, characterized in that, Includes the following steps: Collect multimodal industrial characteristic data, including asset fingerprint data, industrial control network flow data, security equipment alarm data, and operation and maintenance logs; The process role in the fingerprint data of the asset is fused with the instruction sequence in the flow data of the industrial control network at the feature level to construct a business asset traffic triplet and form a business security baseline. Assets in the industrial internet system are identified as multiple evaluation nodes. Protection capabilities are determined based on the configuration compliance and boundary protection capabilities of each evaluation node. Detection capabilities are determined based on the probe deployment and traffic coverage of each evaluation node. Response capabilities are determined based on the closed-loop processing time of each evaluation node. A security capability coupling assessment model is constructed based on a time-series causal graph. The protection capability, detection capability, and response capability of each assessment node are defined as interacting nonlinear system nodes. The nonlinear influence weight relationship of each security capability among different nonlinear system nodes is characterized by a collaborative influence matrix. The steady-state recovery time of the multiple assessment nodes under virtual disturbance is solved according to the business security baseline and the collaborative influence matrix to obtain the resilience coefficient. Align the time window with the process cycle in the industrial scenario, determine the change event based on the operation and maintenance log, analyze the causal lag time between the change event and the alarm event in the historical data through the causal discovery algorithm, and dynamically adjust the weight coefficients in the collaborative influence matrix based on the causal backtracking results. The dynamic safety capability assessment results are generated and output based on the adjusted weighting coefficients and the resilience coefficients. 2.The industrial internet security capability evaluation framework of claim 1, wherein, The process role in the fingerprint data of the asset is fused with the instruction sequence in the flow data of the industrial control network at the feature level to construct a business asset traffic triplet, forming a business security baseline. The method is as follows: The fingerprint data of the assets is analyzed to extract the process stage affiliation and process execution role of each industrial asset, and the flow data of the industrial control network is analyzed to extract the instruction sequence reflecting the control logic and its temporal dependencies. The process stage is aligned with the instruction sequence through a process execution time window. Within the time window, a mapping relationship is established between the process execution role and the control instructions in the instruction sequence, forming a triplet structure that includes a business asset identifier, the standard operation feature corresponding to the process execution role, and the traffic interaction object associated with the standard operation feature. Using the triplet structure as a unit, the matching patterns of business assets and traffic behavior are accumulated and recorded in multiple consecutive process cycles. By statistically analyzing the recurrence frequency and fluctuation range of the matching patterns, a business security baseline that is dynamically updated with the process cycle of the industrial scenario is generated. The business security baseline is used to characterize the standard compliance relationship between business assets and traffic behavior at each process stage. 3.The industrial internet security capability evaluation framework of claim 1, wherein, Protection capabilities are determined based on the configuration compliance and boundary protection capabilities of each assessment node; detection capabilities are determined based on the probe deployment and traffic coverage of each assessment node; and response capabilities are determined based on the closed-loop processing time of each assessment node. The method is as follows: For each evaluation node, the security configuration parameters and boundary access control policies of the industrial control system under the evaluation node are extracted. The protection capability of the evaluation node is determined by the degree of matching between the security configuration parameters and the compliance baseline and the degree of isolation of cross-security domain access by the boundary access control policies. The coverage and protocol parsing depth of the probes deployed in the network topology of the evaluation node are analyzed, and the traffic coverage of the network interfaces under the jurisdiction of the evaluation node is collected. The detection capability of the evaluation node is determined by the correlation and fusion of the probe coverage, protocol parsing depth and traffic coverage. The response capability of the assessment node is determined by analyzing the time consumed in each step of the security incident handling process of the assessment node and using the ratio of the preset baseline period to the sum of the times consumed in each step. The missing steps are accumulated using the preset maximum time consumed.
4. The industrial internet security capability assessment framework according to claim 3, characterized in that, The protection capability is determined based on the configuration compliance and boundary protection capabilities of each assessment node, using the following method: For each evaluation node, extract the security configuration parameter set and boundary access control policy set of the industrial control system under the jurisdiction of the evaluation node; The security configuration parameter set is compared with the preset compliance baseline library to mark deviation items, and the boundary access control policy set is mapped into an access control matrix; By traversing the access path entries in the access control matrix, unauthorized access paths across security domains are identified to calculate policy coverage completeness; The reverse normalized value of the deviation term is used to characterize the configuration compliance level, and the policy coverage integrity is used to characterize the boundary isolation strength. The configuration compliance level and the boundary isolation strength are multiplied, and the result of the multiplication is used as the protection capability of the evaluation node.
5. The industrial internet security capability assessment framework according to claim 3, characterized in that, The detection capability is determined based on the probe deployment and traffic coverage of each evaluation node, using the following method: For each evaluation node, the deployment location and probe type of the security probes deployed in its network topology are analyzed. The probe coverage is determined based on the matching relationship between the deployment location and the path that the critical traffic must pass through. The protocol parsing depth is determined based on the protocol parsing capability corresponding to the probe type. The traffic mirroring configuration of the network interfaces under the jurisdiction of the evaluation node is collected, and the traffic coverage is obtained by comparing the coverage of the mirrored ports with the interaction path of the full traffic. The effective detection coverage is represented by the product of the probe coverage area and the protocol parsing depth, and the product of the effective detection coverage and the traffic coverage rate is used as the detection capability of the evaluation node.
6. The industrial internet security capability assessment framework according to claim 3, characterized in that, Response capability is determined based on the closed-loop processing time of each assessment node, using the following method: For each assessment node, the closed-loop cycle of its security incident handling process from the triggering of the detection event to the completion of the handling action is analyzed, and the closed-loop cycle is broken down into detection confirmation time, analysis and location time, and blocking and recovery time. Collect event timestamps of the security devices and operation and maintenance systems under the jurisdiction of the evaluation node, extract the detection confirmation time, the analysis and location time and the blocking and recovery time respectively by the difference of event timestamps, and use the sum of the detection confirmation time, the analysis and location time and the blocking and recovery time as the total closed-loop time of the handling. The response capability of the evaluation node is characterized by the ratio of a preset baseline period to the total closed-loop processing time. If any processing step is missing in the evaluation node, the time of that processing step is set to the preset maximum threshold time and added to the accumulation.
7. The industrial internet security capability assessment framework according to claim 1, characterized in that, The nonlinear influence weight relationship of each security capability among nodes in different nonlinear systems is characterized by a collaborative influence matrix. The method is as follows: Using multiple evaluation nodes as row and column indices, and protection capability, detection capability, and response capability as capability components of each evaluation node, a collaborative influence matrix is constructed. Each element in the collaborative influence matrix corresponds to the weighting coefficient between the source capability component of the row index evaluation node and the target capability component of the column index evaluation node. Based on the time sequence of dependence of each stage of the security incident attack chain on protection capabilities, detection capabilities, and response capabilities, initial values are assigned to the weighting coefficients of protection capabilities on detection capabilities, detection capabilities on response capabilities, and response capabilities on protection capabilities within the same evaluation node. Based on the frequency and direction of business data flow interactions between different evaluation nodes, initial values are assigned to the weight coefficients of the same capability components between different evaluation nodes; these initial values are then filled into the corresponding element positions of the collaborative influence matrix to form an initial collaborative influence matrix.
8. The industrial internet security capability assessment framework according to claim 1, characterized in that, Based on the business security baseline and the collaborative impact matrix, the steady-state recovery time of the multiple evaluation nodes under virtual disturbances is calculated to obtain the resilience coefficient. The method is as follows: Virtual disturbance events are applied to each evaluation node in the time-series cause-effect graph. The virtual disturbance events trigger the initial state value of the protection capability, detection capability, or response capability of the specified evaluation node to deviate from the standard state value defined by the business security baseline. Using the influence relationships recorded in the collaborative influence matrix as iterative constraints, the state values of each capability component of each evaluation node are updated round by round within the discrete time step until the state values of each capability component of all evaluation nodes converge to a deviation from the business security baseline of less than a preset threshold. The total number of time steps from the application of the virtual disturbance event to the convergence of the state value is counted, and the reciprocal or normalized value of the total number of time steps is used as the resilience coefficient.
9. The industrial internet security capability assessment framework according to claim 1, characterized in that, The causal lag time between change events and alarm events in historical data is analyzed using a causal discovery algorithm. Based on the causal backtracking results, the weight coefficients in the collaborative influence matrix are dynamically adjusted. The method is as follows: Collect operation and maintenance logs and security device alarm logs within a historical time window, extract change event sequences from the operation and maintenance logs, and extract alarm event sequences from the alarm logs; Using the sequence of change events as the dependent variable sequence and the sequence of alarm events as the resultant variable sequence, the causal lag time between each change event and each alarm event is calculated using a causal discovery algorithm. Using the change event and alarm event pairs with causal lag time less than a preset threshold as valid causal pairs, the number of valid causal pairs and the average causal lag time within the scope of each evaluation node are counted to generate causal backtracking results. The weight coefficients in the collaborative influence matrix are dynamically adjusted based on the causal backtracking results. The dynamic adjustment takes the number of effective causal pairs and the average causal lag time as inputs, and maps the inputs to the correction amount of the weight coefficients through a preset mapping rule. The weight coefficients of the corresponding evaluation nodes in the collaborative influence matrix are updated with the correction amount.
10. The industrial internet security capability assessment framework according to claim 1, characterized in that, The dynamic safety capability assessment result is generated and output based on the adjusted weighting coefficients and the toughness coefficients, and the method is as follows: A weight vector is constructed using the weight coefficients of each evaluation node in the adjusted synergistic influence matrix, and a resilience vector is constructed using the resilience coefficients of each evaluation node. The weight vector and the resilience vector are then fused element by element to generate a comprehensive evaluation value for each evaluation node. Taking the process stage in an industrial scenario as the dimension, the comprehensive evaluation values of each evaluation node under the same process stage are aggregated to obtain the stage evaluation value of that process stage. The safety level of a process stage is determined by the degree of deviation between the stage evaluation value and the preset benchmark value of the process stage. The safety level of each process stage is associated with the corresponding assessment node identifier and output to the display interface.